{
  "id": 434353,
  "title": "Silver - LB 0.770 - Two Lines of Code!",
  "url": "/competitions/asl-fingerspelling/discussion/434353",
  "author_name": "Chris Deotte",
  "post_date": "2023-08-24T23:59:51.818000",
  "votes": 98,
  "comment_count": 45,
  "views": 0,
  "content": "<p>Thanks Kaggle and Google for hosting a fun ASL competition.</p>\n<p>I joined a few days ago, so I didn't have time to build my own model. Instead, I began with the best public notebook <a href=\"https://www.kaggle.com/code/saidineshpola/aslfr-ctc-based-on-prev-comp-1st-place?scriptVersionId=139048726\" target=\"_blank\">here</a> (version 17) and attempted to improve it. After making a few changes, I boosted the LB from 0.700 to an amazing 0.770! and obtained Silver medal !!</p>\n<h1>Change Two Lines of Code</h1>\n<p>The best public notebook is Rohith Ingilela's awesome public notebook <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">here</a> which was improved by Saidineshpola <a href=\"https://www.kaggle.com/code/saidineshpola/aslfr-ctc-based-on-prev-comp-1st-place?scriptVersionId=139048726\" target=\"_blank\">here</a>. Version 17 of Saidineshpola's notebook achieves CV = 0.689 (scroll to bottom of version 17) and LB = 0.697.</p>\n<p>That notebook uses TF Records made by Rohith Ingilela <a href=\"https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset-tfrecords-mean-std\" target=\"_blank\">here</a>. An easy trick to boost the performance of the public notebook is create TF Records which keep frames where hands are missing. The TF Records made by linked notebook removes all frames without hands. Instead we can use <code>\"output two\"</code> below and keep 50% of the frames with missing hands below to boost CV and LB. Updated notebook published <a href=\"https://www.kaggle.com/code/cdeotte/2-lines-of-code-change-lb-0-760\" target=\"_blank\">here</a>.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2023/asl_preprocess.png\" alt=\"\"></p>\n<p>Here is preprocess code to use when making TF Records and during inference:</p>\n<pre><code>hand = tf.concat(, axis=)\nhand = tf.where(tf.math.is, , hand)\nmask = tf.math.not, )\nalternating_tensor = tf.math.equal( tf.cumsum(\n    tf.ones ))%,  )\nmask = tf.math.logical\n</code></pre>\n<h1>CTC (Connectionist Temporal Classification) Loss:</h1>\n<p>In this competition, I learned about CTC loss. This loss is amazing! It allows us to create a model with variable length input and predict variable length output in one step. So we can do things like seq2seq without waiting for a sequential decoder to decode each step. Instead we predict the entire output at once super fast! Giving the model some of the frames with hands missing helps identify duplicates and transistions:</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2023/asl_ctc.png\" alt=\"\"></p>\n<h1>Time Augmentation LB +0.004!</h1>\n<p>The public notebook trains for 50 epochs, I found that training for more epochs continues to boost CV and LB score! My final submission trains for 200 epochs. Furthermore, we can add augmentation (i.e. regularization) which helps the model train longer, prevent overfitting, and generalize better.</p>\n<p>Below we see the histogram of train data Number of Frames divided by Character Length of Target Phrase. From this plot, we see that the ratio varies a lot. Some videos have a different recorded frame rate than others and some participants sign faster than others.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2023/frame_hist.png\" alt=\"\"></p>\n<p>What this means is that the model has a hard time transfer learning one frame rate to learn about another frame rate. We can help the model by using time augmentation. For each input sequence, we can randomly shrink frame length by 50% or enlarge 150%. This will add lots of new train data and help the model learn about different frame rates. Here is augmentation code:</p>\n<pre><code> tf.random.uniform(shape=(), minval=, maxval=)&lt;.:\n     = tf.math.round( tf.random.uniform(\n        =(), minval = tf.cast(tf.shape(lip)[],tf.float32) / ., \n         = tf.cast(tf.shape(lip)[],tf.float32) * .) )\n     x in</code></pre>\n<h1>Post Processing LB +0.004!</h1>\n<p>Sometimes the model doesn't make a good prediction. The shortest train data is length 3. When the model predicts length 2 or less, we know it is a bad prediction. Therefore we can replace bad predictions with the best constant length prediction. Anokas found the best constant length prediction <a href=\"https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb\" target=\"_blank\">here</a>. We add this to our TF Lite mode with the following code:</p>\n<pre><code>x = tf(tf(x) &lt; , lambda: tf(\n    , tf.int64), lambda: tf(x))\n</code></pre>\n<h1>Solution Code</h1>\n<p>I published a Kaggle notebook <a href=\"https://www.kaggle.com/code/cdeotte/2-lines-of-code-change-lb-0-760\" target=\"_blank\">here</a> demonstrating the above 3 changes. The first four bullet points below achieve <code>LB = 0.763</code>, then time augmentation boost <code>+0.004</code> and PP boost <code>+0.004</code>. Note that the second bullet point (about batch size) just makes things faster but doesn't change the CV nor LB score. The other bullet points are the key:</p>\n<ul>\n<li>Change 2 lines of code to keep 50% missing hand frames</li>\n<li>Change batch size from 32 to 128 and learning rate from 1e-3 to 4e-3</li>\n<li>Change FRAME_LEN from 128 to 216</li>\n<li>Increase epochs 50 to 200</li>\n<li>Add time augmenation</li>\n<li>Add post process for preds less than 3 chars</li>\n</ul>",
  "messages": [
    {
      "id": 2407193,
      "postDate": "2023-08-24T23:59:51.820Z",
      "content": "<p>Thanks Kaggle and Google for hosting a fun ASL competition.</p>\n<p>I joined a few days ago, so I didn't have time to build my own model. Instead, I began with the best public notebook <a href=\"https://www.kaggle.com/code/saidineshpola/aslfr-ctc-based-on-prev-comp-1st-place?scriptVersionId=139048726\" target=\"_blank\">here</a> (version 17) and attempted to improve it. After making a few changes, I boosted the LB from 0.700 to an amazing 0.770! and obtained Silver medal !!</p>\n<h1>Change Two Lines of Code</h1>\n<p>The best public notebook is Rohith Ingilela's awesome public notebook <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">here</a> which was improved by Saidineshpola <a href=\"https://www.kaggle.com/code/saidineshpola/aslfr-ctc-based-on-prev-comp-1st-place?scriptVersionId=139048726\" target=\"_blank\">here</a>. Version 17 of Saidineshpola's notebook achieves CV = 0.689 (scroll to bottom of version 17) and LB = 0.697.</p>\n<p>That notebook uses TF Records made by Rohith Ingilela <a href=\"https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset-tfrecords-mean-std\" target=\"_blank\">here</a>. An easy trick to boost the performance of the public notebook is create TF Records which keep frames where hands are missing. The TF Records made by linked notebook removes all frames without hands. Instead we can use <code>\"output two\"</code> below and keep 50% of the frames with missing hands below to boost CV and LB. Updated notebook published <a href=\"https://www.kaggle.com/code/cdeotte/2-lines-of-code-change-lb-0-760\" target=\"_blank\">here</a>.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2023/asl_preprocess.png\" alt=\"\"></p>\n<p>Here is preprocess code to use when making TF Records and during inference:</p>\n<pre><code>hand = tf.concat(, axis=)\nhand = tf.where(tf.math.is, , hand)\nmask = tf.math.not, )\nalternating_tensor = tf.math.equal( tf.cumsum(\n    tf.ones ))%,  )\nmask = tf.math.logical\n</code></pre>\n<h1>CTC (Connectionist Temporal Classification) Loss:</h1>\n<p>In this competition, I learned about CTC loss. This loss is amazing! It allows us to create a model with variable length input and predict variable length output in one step. So we can do things like seq2seq without waiting for a sequential decoder to decode each step. Instead we predict the entire output at once super fast! Giving the model some of the frames with hands missing helps identify duplicates and transistions:</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2023/asl_ctc.png\" alt=\"\"></p>\n<h1>Time Augmentation LB +0.004!</h1>\n<p>The public notebook trains for 50 epochs, I found that training for more epochs continues to boost CV and LB score! My final submission trains for 200 epochs. Furthermore, we can add augmentation (i.e. regularization) which helps the model train longer, prevent overfitting, and generalize better.</p>\n<p>Below we see the histogram of train data Number of Frames divided by Character Length of Target Phrase. From this plot, we see that the ratio varies a lot. Some videos have a different recorded frame rate than others and some participants sign faster than others.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2023/frame_hist.png\" alt=\"\"></p>\n<p>What this means is that the model has a hard time transfer learning one frame rate to learn about another frame rate. We can help the model by using time augmentation. For each input sequence, we can randomly shrink frame length by 50% or enlarge 150%. This will add lots of new train data and help the model learn about different frame rates. Here is augmentation code:</p>\n<pre><code> tf.random.uniform(shape=(), minval=, maxval=)&lt;.:\n     = tf.math.round( tf.random.uniform(\n        =(), minval = tf.cast(tf.shape(lip)[],tf.float32) / ., \n         = tf.cast(tf.shape(lip)[],tf.float32) * .) )\n     x in</code></pre>\n<h1>Post Processing LB +0.004!</h1>\n<p>Sometimes the model doesn't make a good prediction. The shortest train data is length 3. When the model predicts length 2 or less, we know it is a bad prediction. Therefore we can replace bad predictions with the best constant length prediction. Anokas found the best constant length prediction <a href=\"https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb\" target=\"_blank\">here</a>. We add this to our TF Lite mode with the following code:</p>\n<pre><code>x = tf(tf(x) &lt; , lambda: tf(\n    , tf.int64), lambda: tf(x))\n</code></pre>\n<h1>Solution Code</h1>\n<p>I published a Kaggle notebook <a href=\"https://www.kaggle.com/code/cdeotte/2-lines-of-code-change-lb-0-760\" target=\"_blank\">here</a> demonstrating the above 3 changes. The first four bullet points below achieve <code>LB = 0.763</code>, then time augmentation boost <code>+0.004</code> and PP boost <code>+0.004</code>. Note that the second bullet point (about batch size) just makes things faster but doesn't change the CV nor LB score. The other bullet points are the key:</p>\n<ul>\n<li>Change 2 lines of code to keep 50% missing hand frames</li>\n<li>Change batch size from 32 to 128 and learning rate from 1e-3 to 4e-3</li>\n<li>Change FRAME_LEN from 128 to 216</li>\n<li>Increase epochs 50 to 200</li>\n<li>Add time augmenation</li>\n<li>Add post process for preds less than 3 chars</li>\n</ul>",
      "rawMarkdown": "Thanks Kaggle and Google for hosting a fun ASL competition.\n\nI joined a few days ago, so I didn't have time to build my own model. Instead, I began with the best public notebook [here][2] (version 17) and attempted to improve it. After making a few changes, I boosted the LB from 0.700 to an amazing 0.770! and obtained Silver medal !!\n\n# Change Two Lines of Code\nThe best public notebook is Rohith Ingilela's awesome public notebook [here][1] which was improved by Saidineshpola [here][2]. Version 17 of Saidineshpola's notebook achieves CV = 0.689 (scroll to bottom of version 17) and LB = 0.697.\n\nThat notebook uses TF Records made by Rohith Ingilela [here][3]. An easy trick to boost the performance of the public notebook is create TF Records which keep frames where hands are missing. The TF Records made by linked notebook removes all frames without hands. Instead we can use `\"output two\"` below and keep 50% of the frames with missing hands below to boost CV and LB. Updated notebook published [here][4].\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2023/asl_preprocess.png)\n\nHere is preprocess code to use when making TF Records and during inference:\n\n    hand = tf.concat([rhand, lhand], axis=1)\n    hand = tf.where(tf.math.is_nan(hand), 0.0, hand)\n    mask = tf.math.not_equal(tf.reduce_sum(hand, axis=[1, 2]), 0.0)\n    alternating_tensor = tf.math.equal( tf.cumsum(\n        tf.ones_like( tf.reduce_sum(hand, axis=[1, 2]) ))%2, 1.0 )\n    mask = tf.math.logical_or(mask, alternating_tensor)\n\n# CTC (Connectionist Temporal Classification) Loss:\nIn this competition, I learned about CTC loss. This loss is amazing! It allows us to create a model with variable length input and predict variable length output in one step. So we can do things like seq2seq without waiting for a sequential decoder to decode each step. Instead we predict the entire output at once super fast! Giving the model some of the frames with hands missing helps identify duplicates and transistions:\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2023/asl_ctc.png)\n\n# Time Augmentation LB +0.004!\nThe public notebook trains for 50 epochs, I found that training for more epochs continues to boost CV and LB score! My final submission trains for 200 epochs. Furthermore, we can add augmentation (i.e. regularization) which helps the model train longer, prevent overfitting, and generalize better.\n\nBelow we see the histogram of train data Number of Frames divided by Character Length of Target Phrase. From this plot, we see that the ratio varies a lot. Some videos have a different recorded frame rate than others and some participants sign faster than others.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2023/frame_hist.png)\n\nWhat this means is that the model has a hard time transfer learning one frame rate to learn about another frame rate. We can help the model by using time augmentation. For each input sequence, we can randomly shrink frame length by 50% or enlarge 150%. This will add lots of new train data and help the model learn about different frame rates. Here is augmentation code:\n\n    if tf.random.uniform(shape=(), minval=0, maxval=1)<0.2:\n        new_height = tf.math.round( tf.random.uniform(\n            shape=(), minval = tf.cast(tf.shape(lip)[0],tf.float32) / 2.0, \n            maxval = tf.cast(tf.shape(lip)[0],tf.float32) * 1.5) )\n        for x in [lip, rhand, lhand, rpose, lpose]:\n            x = tf.image.resize(x, (new_height, tf.shape(x)[1]) )\n\n# Post Processing LB +0.004!\nSometimes the model doesn't make a good prediction. The shortest train data is length 3. When the model predicts length 2 or less, we know it is a bad prediction. Therefore we can replace bad predictions with the best constant length prediction. Anokas found the best constant length prediction [here][5]. We add this to our TF Lite mode with the following code:\n\n    x = tf.cond(tf.shape(x)[0] < 3, lambda: tf.constant(\n        [17, 0, 32, 12, 36, 0, 12, 32, 49, 46, 36], tf.int64), lambda: tf.identity(x))\n\n# Solution Code\nI published a Kaggle notebook [here][4] demonstrating the above 3 changes. The first four bullet points below achieve `LB = 0.763`, then time augmentation boost `+0.004` and PP boost `+0.004`. Note that the second bullet point (about batch size) just makes things faster but doesn't change the CV nor LB score. The other bullet points are the key:\n\n* Change 2 lines of code to keep 50% missing hand frames\n* Change batch size from 32 to 128 and learning rate from 1e-3 to 4e-3\n* Change FRAME_LEN from 128 to 216\n* Increase epochs 50 to 200\n* Add time augmenation\n* Add post process for preds less than 3 chars\n\n[1]: https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\n[2]: https://www.kaggle.com/code/saidineshpola/aslfr-ctc-based-on-prev-comp-1st-place?scriptVersionId=139048726\n[3]: https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset-tfrecords-mean-std\n[4]: https://www.kaggle.com/code/cdeotte/2-lines-of-code-change-lb-0-760\n[5]: https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb",
      "votes": 98
    },
    {
      "id": 2407214,
      "postDate": "2023-08-25T00:34:17.757Z",
      "content": "<p>I used:</p>\n<pre><code> len() &lt;= :\n     =  + \n</code></pre>\n<p>it gives me +0.003 in CV and LB</p>",
      "rawMarkdown": "I used:\n```\nif len(pred) <= 4:\n    pred = pred + \" -aero\"\n```\nit gives me +0.003 in CV and LB",
      "votes": 3,
      "replies": [
        {
          "id": 2407219,
          "postDate": "2023-08-25T00:47:00.200Z",
          "content": "<p>Could you explain to me how do you come up with adding the string ‘-aero’ ? </p>",
          "rawMarkdown": "Could you explain to me how do you come up with adding the string ‘-aero’ ? ",
          "votes": 2,
          "replies": [
            {
              "id": 2407231,
              "postDate": "2023-08-25T01:13:37.313Z",
              "content": "<ol>\n<li>I checked my worst predictions in the validation set and found that shorter predictions are worse. </li>\n<li>I found that only few phrase's length is less or equal to 5. </li>\n<li>I picked the most common chars in training set, they are: \"a\", \"e\", \"r\", \"o\", \"-\", \" \".</li>\n<li>I test all the combinations of \"a\", \"e\", \"r\", \"o\", \"-\", \" \" based on validation dataset and \" -aero\" is the best.</li>\n</ol>",
              "rawMarkdown": "1. I checked my worst predictions in the validation set and found that shorter predictions are worse. \n2. I found that only few phrase's length is less or equal to 5. \n3. I picked the most common chars in training set, they are: \"a\", \"e\", \"r\", \"o\", \"-\", \" \".\n4. I test all the combinations of \"a\", \"e\", \"r\", \"o\", \"-\", \" \" based on validation dataset and \" -aero\" is the best.",
              "votes": 2
            },
            {
              "id": 2408401,
              "postDate": "2023-08-25T16:07:45.123Z",
              "content": "<p>Great trick <a href=\"https://www.kaggle.com/nightsh4de\" target=\"_blank\">@nightsh4de</a> Congratulations on 17th place solo Silver. You were so close to receiving another solo Gold, fantastic!</p>",
              "rawMarkdown": "Great trick @nightsh4de Congratulations on 17th place solo Silver. You were so close to receiving another solo Gold, fantastic!",
              "votes": 2
            },
            {
              "id": 2408960,
              "postDate": "2023-08-26T00:29:23.847Z",
              "content": "<p>Thanks, Chris. Your write-up is insightful as always😄</p>",
              "rawMarkdown": "Thanks, Chris. Your write-up is insightful as always😄",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2556020,
      "postDate": "2023-12-10T12:04:02.797Z",
      "content": "<p>This makes so much sense! Was reading up on CTC Loss and this cleared so many things, extremely grateful!</p>",
      "rawMarkdown": "This makes so much sense! Was reading up on CTC Loss and this cleared so many things, extremely grateful!",
      "votes": 1,
      "replies": [
        {
          "id": 2556158,
          "postDate": "2023-12-10T13:53:09.927Z",
          "content": "<p>Indeed, probably the best CTC explanation I have read as well. 👌</p>",
          "rawMarkdown": "Indeed, probably the best CTC explanation I have read as well. 👌"
        }
      ]
    },
    {
      "id": 2407199,
      "postDate": "2023-08-25T00:10:13.820Z",
      "content": "<p>good work!</p>\n<p>the reason why  original frame length 128 does not work is because it is \"too short\". you can check the error of the predicted phrase, many error error are truncation at the end. if you just change 128 to 256 (without dropping missing), you will see immediate improvement. </p>\n<p>dropping missing frame while keeping it at 128 is another way.</p>\n<p>it is important to consider the need (or the need not) to adjust kernel size in 1d conv when you change the temporal length or resolution</p>",
      "rawMarkdown": "good work!\n\nthe reason why  original frame length 128 does not work is because it is \"too short\". you can check the error of the predicted phrase, many error error are truncation at the end. if you just change 128 to 256 (without dropping missing), you will see immediate improvement. \n\ndropping missing frame while keeping it at 128 is another way.\n\nit is important to consider the need (or the need not) to adjust kernel size in 1d conv when you change the temporal length or resolution",
      "votes": 2,
      "replies": [
        {
          "id": 2407227,
          "postDate": "2023-08-25T01:04:41.470Z",
          "content": "<p>The original notebook resizes the frames to 128. So the original notebook \"sees\" all the frames. But the original notebook drops the frames without hands (before resize). If we just keep 50% frames without hands and resize to 128, we get a big boost in CV and LB. So in other words even with 128 in both the original and updated, just adding frames without hands (before resize) gives a big boost.</p>",
          "rawMarkdown": "The original notebook resizes the frames to 128. So the original notebook \"sees\" all the frames. But the original notebook drops the frames without hands (before resize). If we just keep 50% frames without hands and resize to 128, we get a big boost in CV and LB. So in other words even with 128 in both the original and updated, just adding frames without hands (before resize) gives a big boost.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2423241,
      "postDate": "2023-09-04T14:25:04.587Z",
      "content": "<p>I am going to work on it as well and as always your knowledge and information helps a lot. Thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "I am going to work on it as well and as always your knowledge and information helps a lot. Thanks for sharing @cdeotte ",
      "votes": 1,
      "replies": [
        {
          "id": 2429905,
          "postDate": "2023-09-08T22:34:40.827Z",
          "content": "<p>Thank you   </p>",
          "rawMarkdown": "Thank you   "
        }
      ]
    },
    {
      "id": 2422399,
      "postDate": "2023-09-04T02:37:46.707Z",
      "content": "<p>Your explanation is very clear, thank you for your work!</p>",
      "rawMarkdown": "Your explanation is very clear, thank you for your work!",
      "votes": 1
    },
    {
      "id": 2418460,
      "postDate": "2023-09-01T09:32:54.913Z",
      "content": "<p>Congratulations and thank you for sharing!</p>",
      "rawMarkdown": "Congratulations and thank you for sharing!",
      "votes": 1,
      "replies": [
        {
          "id": 2421468,
          "postDate": "2023-09-03T11:44:57.067Z",
          "content": "<p>Thanks Gazuwu</p>",
          "rawMarkdown": "Thanks Gazuwu"
        }
      ]
    },
    {
      "id": 2418200,
      "postDate": "2023-09-01T07:35:46.720Z",
      "content": "<p>I'm working on a similar project. This dicussion is going to help me a lot, thank you for sharing! </p>",
      "rawMarkdown": "I'm working on a similar project. This dicussion is going to help me a lot, thank you for sharing! ",
      "votes": 1
    },
    {
      "id": 2418020,
      "postDate": "2023-09-01T04:33:18.390Z",
      "content": "<p>man what a great way to get to silver! Love that we can help each other grow, good job!</p>",
      "rawMarkdown": "man what a great way to get to silver! Love that we can help each other grow, good job!",
      "votes": 1
    },
    {
      "id": 2416415,
      "postDate": "2023-08-31T03:16:17.807Z",
      "content": "<p>Thanks for sharing Chris - this was really informative!</p>",
      "rawMarkdown": "Thanks for sharing Chris - this was really informative!",
      "votes": 1
    },
    {
      "id": 2414443,
      "postDate": "2023-08-29T16:23:59.920Z",
      "content": "<p>Easy as hell, well done !!!</p>",
      "rawMarkdown": "Easy as hell, well done !!!",
      "votes": 1
    },
    {
      "id": 2413636,
      "postDate": "2023-08-29T04:24:59.600Z",
      "content": "<p>congrats and thanks for sharing the elegant approach!</p>",
      "rawMarkdown": "congrats and thanks for sharing the elegant approach!",
      "votes": 1,
      "replies": [
        {
          "id": 2414301,
          "postDate": "2023-08-29T14:39:17.093Z",
          "content": "<p>Thanks Abishek</p>",
          "rawMarkdown": "Thanks Abishek"
        }
      ]
    },
    {
      "id": 2413026,
      "postDate": "2023-08-28T16:30:40.130Z",
      "content": "<p>This is very nice effort sir!</p>",
      "rawMarkdown": "This is very nice effort sir!",
      "votes": 1
    },
    {
      "id": 2412748,
      "postDate": "2023-08-28T13:29:32.193Z",
      "content": "<p>I also considered whether I should remove some frames without hands, but I didn't proceed to validate this idea, which is quite unfortunate. I'm curious to know how much removing 50% of frames without hands would improve the score on the public test set compared to not removing these frames at all. Also, I'm interested in understanding how much faster the training process would be.</p>",
      "rawMarkdown": "I also considered whether I should remove some frames without hands, but I didn't proceed to validate this idea, which is quite unfortunate. I'm curious to know how much removing 50% of frames without hands would improve the score on the public test set compared to not removing these frames at all. Also, I'm interested in understanding how much faster the training process would be.",
      "votes": 1,
      "replies": [
        {
          "id": 2413525,
          "postDate": "2023-08-29T01:30:12.817Z",
          "content": "<p>Good question. I'm not sure how much exactly. But removing 50% was better than 0% for both CV and LB. And it made training faster.</p>",
          "rawMarkdown": "Good question. I'm not sure how much exactly. But removing 50% was better than 0% for both CV and LB. And it made training faster.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2409075,
      "postDate": "2023-08-26T04:00:14.197Z",
      "content": "<p>Great Sir !</p>",
      "rawMarkdown": "Great Sir !",
      "votes": 1
    },
    {
      "id": 2408658,
      "postDate": "2023-08-25T18:35:46.540Z",
      "content": "<p>Thanks for sharing !! It is really helpful and informative</p>",
      "rawMarkdown": "Thanks for sharing !! It is really helpful and informative",
      "votes": 1,
      "replies": [
        {
          "id": 2411046,
          "postDate": "2023-08-27T11:47:02.210Z",
          "content": "<p>Thank you Priyanshu1235</p>",
          "rawMarkdown": "Thank you Priyanshu1235"
        }
      ]
    },
    {
      "id": 2408306,
      "postDate": "2023-08-25T14:59:39.027Z",
      "content": "<p>Congratulations. Thanks for sharing your innovative approach. <br>\nVery interesting code improvements and insights. </p>",
      "rawMarkdown": "Congratulations. Thanks for sharing your innovative approach. \nVery interesting code improvements and insights. ",
      "votes": 1
    },
    {
      "id": 2407751,
      "postDate": "2023-08-25T08:35:28.990Z",
      "content": "<p>Thank you for the tips. This will help a lot in upcoming challenges. My thinking was that removing the frames with missing hands while using positional encodings wouldn't be good for the model.</p>",
      "rawMarkdown": "Thank you for the tips. This will help a lot in upcoming challenges. My thinking was that removing the frames with missing hands while using positional encodings wouldn't be good for the model.",
      "votes": 1
    },
    {
      "id": 2407348,
      "postDate": "2023-08-25T03:38:41.520Z",
      "content": "<p>Very helpful!</p>",
      "rawMarkdown": "Very helpful!",
      "votes": 1
    },
    {
      "id": 2407197,
      "postDate": "2023-08-25T00:06:30.543Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nCan you tell me what machine are you using to train your model? Do you have your own or you use cloud one?</p>",
      "rawMarkdown": "Congratulations @cdeotte \nCan you tell me what machine are you using to train your model? Do you have your own or you use cloud one?",
      "votes": 1,
      "replies": [
        {
          "id": 2407277,
          "postDate": "2023-08-25T02:18:24.630Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ankurlimbashia\" target=\"_blank\">@ankurlimbashia</a> As an Nvidia employee, I have access to Nvidia cloud GPUs. So in the past few days, I had a bunch of V100 GPUs running experiments in the cloud. </p>",
          "rawMarkdown": "Hi @ankurlimbashia As an Nvidia employee, I have access to Nvidia cloud GPUs. So in the past few days, I had a bunch of V100 GPUs running experiments in the cloud. ",
          "votes": 3,
          "replies": [
            {
              "id": 2407319,
              "postDate": "2023-08-25T03:00:53.763Z",
              "content": "<p>a bunch means ~five? 😁</p>",
              "rawMarkdown": "a bunch means ~five? 😁",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2407622,
      "postDate": "2023-08-25T06:52:31.690Z",
      "content": "<p>These are very clever tricks. Out of curiosity, how did you find out that using frames with missing hands was better? Was it by luck or you had an intuition behind it? Anyways, congratulations on the silver on such a short time.🥳</p>",
      "rawMarkdown": "These are very clever tricks. Out of curiosity, how did you find out that using frames with missing hands was better? Was it by luck or you had an intuition behind it? Anyways, congratulations on the silver on such a short time.🥳",
      "votes": 2,
      "replies": [
        {
          "id": 2408898,
          "postDate": "2023-08-25T22:07:57.020Z",
          "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> Congratulations on becoming Kaggle Discussion Grandmaster!</p>\n<p>The first thing i did when analyzing the public notebook was to train longer than 50 epochs. That boosted LB from 0.697 to like LB 0.700 with 100 epochs. Then i increased the number of <code>FRAME_LEN</code>. That boost CV and LB too.</p>\n<p>The next thing i did was analyze the preprocessing code. Lots of time we can improve computer vision, NLP, etc with better preprocess. I saw that the preprocess removed frames, so i tried training a model <strong>without</strong> removing any frames. And wow! The CV and LB boost a lot. Then i changed it to keep 50% frames and that boost CV and LB more and speed up training. In retrospect it makes sense with the CTC loss because the model needs to know how many spacers to add to the output. Providing the missing frames gives the model clues, i think. </p>\n<p>Next I tried augmentation and training for more epochs. This continued to boost CV and LB. Next i tried simple changes to the architecture but I didn't have much success with that.</p>\n<p>After reading other solutions, it seems like the way to boost more would be to apply even more aggressive augmentation and train longer.</p>",
          "rawMarkdown": "@yassinealouini Congratulations on becoming Kaggle Discussion Grandmaster!\n\nThe first thing i did when analyzing the public notebook was to train longer than 50 epochs. That boosted LB from 0.697 to like LB 0.700 with 100 epochs. Then i increased the number of `FRAME_LEN`. That boost CV and LB too.\n\nThe next thing i did was analyze the preprocessing code. Lots of time we can improve computer vision, NLP, etc with better preprocess. I saw that the preprocess removed frames, so i tried training a model **without** removing any frames. And wow! The CV and LB boost a lot. Then i changed it to keep 50% frames and that boost CV and LB more and speed up training. In retrospect it makes sense with the CTC loss because the model needs to know how many spacers to add to the output. Providing the missing frames gives the model clues, i think. \n\nNext I tried augmentation and training for more epochs. This continued to boost CV and LB. Next i tried simple changes to the architecture but I didn't have much success with that.\n\nAfter reading other solutions, it seems like the way to boost more would be to apply even more aggressive augmentation and train longer.",
          "votes": 2,
          "replies": [
            {
              "id": 2409314,
              "postDate": "2023-08-26T07:21:22.373Z",
              "content": "<p>If I understand correctly, it seems that space characters are spelled with missing hands or at least it helps the model? </p>",
              "rawMarkdown": "If I understand correctly, it seems that space characters are spelled with missing hands or at least it helps the model? ",
              "votes": 1
            },
            {
              "id": 2410099,
              "postDate": "2023-08-26T16:25:48.407Z",
              "content": "<p>That's my guess. Also if we remove missing hand frames, maybe certain consecutive hand activity blends together and fools the model into classifying the wrong letter. These are just guesses, i have not performed EDA nor ablation study to figure out why it helps. But using missing hand frames makes a big difference to the model for some reason.</p>",
              "rawMarkdown": "That's my guess. Also if we remove missing hand frames, maybe certain consecutive hand activity blends together and fools the model into classifying the wrong letter. These are just guesses, i have not performed EDA nor ablation study to figure out why it helps. But using missing hand frames makes a big difference to the model for some reason.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2407210,
      "postDate": "2023-08-25T00:22:51.710Z",
      "content": "<p>Congratulations and thanks for your solution</p>",
      "rawMarkdown": "Congratulations and thanks for your solution",
      "votes": 2
    },
    {
      "id": 2407445,
      "postDate": "2023-08-25T05:08:29.373Z",
      "content": "<p>Well, so cool for the missing hands dealing, I only found do not throw miss hands frames work much better, I think your solution might be much better for long input frames. For less input frames as non hand feature might also help I think we might not modify them. Not sure, maybe your method work best, I will give it a try, thanks for sharing!</p>",
      "rawMarkdown": "Well, so cool for the missing hands dealing, I only found do not throw miss hands frames work much better, I think your solution might be much better for long input frames. For less input frames as non hand feature might also help I think we might not modify them. Not sure, maybe your method work best, I will give it a try, thanks for sharing!",
      "replies": [
        {
          "id": 2408901,
          "postDate": "2023-08-25T22:12:58.397Z",
          "content": "<p>Congratulations <a href=\"https://www.kaggle.com/goldenlock\" target=\"_blank\">@goldenlock</a> Amazing job achieving solo cash gold 3rd place!</p>\n<p>I didn't think about keeping more non hand frames for the shorter videos (versus long videos), that might help. I did try both keeping 100% and keeping 50%. But i applied my logic to all videos and using 50% did better.</p>\n<p>Going from 0% non hand to 50% or 100% was a huge boost. I couldn't believe the CV when i first saw it. But i submitted and the LB jumped up just as much.</p>",
          "rawMarkdown": "Congratulations @goldenlock Amazing job achieving solo cash gold 3rd place!\n\nI didn't think about keeping more non hand frames for the shorter videos (versus long videos), that might help. I did try both keeping 100% and keeping 50%. But i applied my logic to all videos and using 50% did better.\n\nGoing from 0% non hand to 50% or 100% was a huge boost. I couldn't believe the CV when i first saw it. But i submitted and the LB jumped up just as much."
        }
      ]
    },
    {
      "id": 2408181,
      "postDate": "2023-08-25T13:48:36.047Z",
      "content": "<p>Congratulations on your Silver medal in the ASL competition! Your two-line code changes leading to a remarkable LB boost are impressive. Your combination of preprocessing, CTC loss, time augmentation, and post-processing displays a solid understanding of techniques. Sharing your journey with clear explanations is valuable for the community. Your success is inspiring, encouraging others to experiment and learn. Well done!</p>",
      "rawMarkdown": "Congratulations on your Silver medal in the ASL competition! Your two-line code changes leading to a remarkable LB boost are impressive. Your combination of preprocessing, CTC loss, time augmentation, and post-processing displays a solid understanding of techniques. Sharing your journey with clear explanations is valuable for the community. Your success is inspiring, encouraging others to experiment and learn. Well done!",
      "votes": 1,
      "replies": [
        {
          "id": 2408892,
          "postDate": "2023-08-25T21:57:22.383Z",
          "content": "<p>Thanks Ranamalla</p>",
          "rawMarkdown": "Thanks Ranamalla"
        }
      ]
    },
    {
      "id": 2407326,
      "postDate": "2023-08-25T03:11:04.873Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2412525,
      "postDate": "2023-08-28T10:25:43.043Z",
      "content": "<p>Nice information thanks </p>",
      "rawMarkdown": "Nice information thanks ",
      "votes": 1
    },
    {
      "id": 2412197,
      "postDate": "2023-08-28T06:27:23.170Z",
      "content": "<p>great. thanks for sharing</p>",
      "rawMarkdown": "great. thanks for sharing",
      "votes": 1
    },
    {
      "id": 2410224,
      "postDate": "2023-08-26T18:19:32.567Z",
      "content": "<p>Thank You for Sharing!!!</p>",
      "rawMarkdown": "Thank You for Sharing!!!\n",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2407214,
      "author_name": "Yu Wu",
      "author_url": "",
      "post_date": "2023-08-25T00:34:17.757000",
      "content": "<p>I used:</p>\n<pre><code> len() &lt;= :\n     =  + \n</code></pre>\n<p>it gives me +0.003 in CV and LB</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2407219,
          "author_name": "HW",
          "author_url": "",
          "post_date": "2023-08-25T00:47:00.200000",
          "content": "<p>Could you explain to me how do you come up with adding the string ‘-aero’ ? </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2407231,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-08-25T01:13:37.313000",
              "content": "<ol>\n<li>I checked my worst predictions in the validation set and found that shorter predictions are worse. </li>\n<li>I found that only few phrase's length is less or equal to 5. </li>\n<li>I picked the most common chars in training set, they are: \"a\", \"e\", \"r\", \"o\", \"-\", \" \".</li>\n<li>I test all the combinations of \"a\", \"e\", \"r\", \"o\", \"-\", \" \" based on validation dataset and \" -aero\" is the best.</li>\n</ol>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2408401,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-08-25T16:07:45.123000",
              "content": "<p>Great trick <a href=\"https://www.kaggle.com/nightsh4de\" target=\"_blank\">@nightsh4de</a> Congratulations on 17th place solo Silver. You were so close to receiving another solo Gold, fantastic!</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2408960,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-08-26T00:29:23.847000",
              "content": "<p>Thanks, Chris. Your write-up is insightful as always😄</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2556020,
      "author_name": "Shlok Ingle",
      "author_url": "",
      "post_date": "2023-12-10T12:04:02.797000",
      "content": "<p>This makes so much sense! Was reading up on CTC Loss and this cleared so many things, extremely grateful!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2556158,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2023-12-10T13:53:09.927000",
          "content": "<p>Indeed, probably the best CTC explanation I have read as well. 👌</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2407199,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-25T00:10:13.820000",
      "content": "<p>good work!</p>\n<p>the reason why  original frame length 128 does not work is because it is \"too short\". you can check the error of the predicted phrase, many error error are truncation at the end. if you just change 128 to 256 (without dropping missing), you will see immediate improvement. </p>\n<p>dropping missing frame while keeping it at 128 is another way.</p>\n<p>it is important to consider the need (or the need not) to adjust kernel size in 1d conv when you change the temporal length or resolution</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2407227,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-08-25T01:04:41.470000",
          "content": "<p>The original notebook resizes the frames to 128. So the original notebook \"sees\" all the frames. But the original notebook drops the frames without hands (before resize). If we just keep 50% frames without hands and resize to 128, we get a big boost in CV and LB. So in other words even with 128 in both the original and updated, just adding frames without hands (before resize) gives a big boost.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2423241,
      "author_name": "Rishabh Jain",
      "author_url": "",
      "post_date": "2023-09-04T14:25:04.587000",
      "content": "<p>I am going to work on it as well and as always your knowledge and information helps a lot. Thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2429905,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-09-08T22:34:40.827000",
          "content": "<p>Thank you   </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2422399,
      "author_name": "Courtyard",
      "author_url": "",
      "post_date": "2023-09-04T02:37:46.707000",
      "content": "<p>Your explanation is very clear, thank you for your work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2418460,
      "author_name": "Gazuwu",
      "author_url": "",
      "post_date": "2023-09-01T09:32:54.913000",
      "content": "<p>Congratulations and thank you for sharing!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2421468,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-09-03T11:44:57.067000",
          "content": "<p>Thanks Gazuwu</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2418200,
      "author_name": "Abdullah Al Asif",
      "author_url": "",
      "post_date": "2023-09-01T07:35:46.720000",
      "content": "<p>I'm working on a similar project. This dicussion is going to help me a lot, thank you for sharing! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2418020,
      "author_name": "WiseK9",
      "author_url": "",
      "post_date": "2023-09-01T04:33:18.390000",
      "content": "<p>man what a great way to get to silver! Love that we can help each other grow, good job!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2416415,
      "author_name": "Marcus Chan",
      "author_url": "",
      "post_date": "2023-08-31T03:16:17.807000",
      "content": "<p>Thanks for sharing Chris - this was really informative!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2414443,
      "author_name": "Bot_developer11",
      "author_url": "",
      "post_date": "2023-08-29T16:23:59.920000",
      "content": "<p>Easy as hell, well done !!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2413636,
      "author_name": "Abishek Sudarshan",
      "author_url": "",
      "post_date": "2023-08-29T04:24:59.600000",
      "content": "<p>congrats and thanks for sharing the elegant approach!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2414301,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-08-29T14:39:17.093000",
          "content": "<p>Thanks Abishek</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2413026,
      "author_name": "Mystic Shadow",
      "author_url": "",
      "post_date": "2023-08-28T16:30:40.130000",
      "content": "<p>This is very nice effort sir!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2412748,
      "author_name": "Scenery SunFireInk",
      "author_url": "",
      "post_date": "2023-08-28T13:29:32.193000",
      "content": "<p>I also considered whether I should remove some frames without hands, but I didn't proceed to validate this idea, which is quite unfortunate. I'm curious to know how much removing 50% of frames without hands would improve the score on the public test set compared to not removing these frames at all. Also, I'm interested in understanding how much faster the training process would be.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2413525,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-08-29T01:30:12.817000",
          "content": "<p>Good question. I'm not sure how much exactly. But removing 50% was better than 0% for both CV and LB. And it made training faster.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2409075,
      "author_name": "Mystic Shadow",
      "author_url": "",
      "post_date": "2023-08-26T04:00:14.197000",
      "content": "<p>Great Sir !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2408658,
      "author_name": "Priyanshu1235",
      "author_url": "",
      "post_date": "2023-08-25T18:35:46.540000",
      "content": "<p>Thanks for sharing !! It is really helpful and informative</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2411046,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-08-27T11:47:02.210000",
          "content": "<p>Thank you Priyanshu1235</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2408306,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-25T14:59:39.027000",
      "content": "<p>Congratulations. Thanks for sharing your innovative approach. <br>\nVery interesting code improvements and insights. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2407751,
      "author_name": "Busang Mfaladi",
      "author_url": "",
      "post_date": "2023-08-25T08:35:28.990000",
      "content": "<p>Thank you for the tips. This will help a lot in upcoming challenges. My thinking was that removing the frames with missing hands while using positional encodings wouldn't be good for the model.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2407348,
      "author_name": "Muhammad Usman",
      "author_url": "",
      "post_date": "2023-08-25T03:38:41.520000",
      "content": "<p>Very helpful!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2407197,
      "author_name": "Ankur Limbashia",
      "author_url": "",
      "post_date": "2023-08-25T00:06:30.543000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nCan you tell me what machine are you using to train your model? Do you have your own or you use cloud one?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2407277,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-08-25T02:18:24.630000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ankurlimbashia\" target=\"_blank\">@ankurlimbashia</a> As an Nvidia employee, I have access to Nvidia cloud GPUs. So in the past few days, I had a bunch of V100 GPUs running experiments in the cloud. </p>",
          "votes": 3,
          "replies": [
            {
              "id": 2407319,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-08-25T03:00:53.763000",
              "content": "<p>a bunch means ~five? 😁</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2407622,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2023-08-25T06:52:31.690000",
      "content": "<p>These are very clever tricks. Out of curiosity, how did you find out that using frames with missing hands was better? Was it by luck or you had an intuition behind it? Anyways, congratulations on the silver on such a short time.🥳</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2408898,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-08-25T22:07:57.020000",
          "content": "<p><a href=\"https://www.kaggle.com/yassinealouini\" target=\"_blank\">@yassinealouini</a> Congratulations on becoming Kaggle Discussion Grandmaster!</p>\n<p>The first thing i did when analyzing the public notebook was to train longer than 50 epochs. That boosted LB from 0.697 to like LB 0.700 with 100 epochs. Then i increased the number of <code>FRAME_LEN</code>. That boost CV and LB too.</p>\n<p>The next thing i did was analyze the preprocessing code. Lots of time we can improve computer vision, NLP, etc with better preprocess. I saw that the preprocess removed frames, so i tried training a model <strong>without</strong> removing any frames. And wow! The CV and LB boost a lot. Then i changed it to keep 50% frames and that boost CV and LB more and speed up training. In retrospect it makes sense with the CTC loss because the model needs to know how many spacers to add to the output. Providing the missing frames gives the model clues, i think. </p>\n<p>Next I tried augmentation and training for more epochs. This continued to boost CV and LB. Next i tried simple changes to the architecture but I didn't have much success with that.</p>\n<p>After reading other solutions, it seems like the way to boost more would be to apply even more aggressive augmentation and train longer.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2409314,
              "author_name": "Yassine Alouini",
              "author_url": "",
              "post_date": "2023-08-26T07:21:22.373000",
              "content": "<p>If I understand correctly, it seems that space characters are spelled with missing hands or at least it helps the model? </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2410099,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-08-26T16:25:48.407000",
              "content": "<p>That's my guess. Also if we remove missing hand frames, maybe certain consecutive hand activity blends together and fools the model into classifying the wrong letter. These are just guesses, i have not performed EDA nor ablation study to figure out why it helps. But using missing hand frames makes a big difference to the model for some reason.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2407210,
      "author_name": "Chakkrit Termritthikun",
      "author_url": "",
      "post_date": "2023-08-25T00:22:51.710000",
      "content": "<p>Congratulations and thanks for your solution</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2407445,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2023-08-25T05:08:29.373000",
      "content": "<p>Well, so cool for the missing hands dealing, I only found do not throw miss hands frames work much better, I think your solution might be much better for long input frames. For less input frames as non hand feature might also help I think we might not modify them. Not sure, maybe your method work best, I will give it a try, thanks for sharing!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2408901,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-08-25T22:12:58.397000",
          "content": "<p>Congratulations <a href=\"https://www.kaggle.com/goldenlock\" target=\"_blank\">@goldenlock</a> Amazing job achieving solo cash gold 3rd place!</p>\n<p>I didn't think about keeping more non hand frames for the shorter videos (versus long videos), that might help. I did try both keeping 100% and keeping 50%. But i applied my logic to all videos and using 50% did better.</p>\n<p>Going from 0% non hand to 50% or 100% was a huge boost. I couldn't believe the CV when i first saw it. But i submitted and the LB jumped up just as much.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2408181,
      "author_name": "Ranamalla Nithin Reddy",
      "author_url": "",
      "post_date": "2023-08-25T13:48:36.047000",
      "content": "<p>Congratulations on your Silver medal in the ASL competition! Your two-line code changes leading to a remarkable LB boost are impressive. Your combination of preprocessing, CTC loss, time augmentation, and post-processing displays a solid understanding of techniques. Sharing your journey with clear explanations is valuable for the community. Your success is inspiring, encouraging others to experiment and learn. Well done!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2408892,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-08-25T21:57:22.383000",
          "content": "<p>Thanks Ranamalla</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2407326,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-25T03:11:04.873000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2412525,
      "author_name": "Abhay tiwari",
      "author_url": "",
      "post_date": "2023-08-28T10:25:43.043000",
      "content": "<p>Nice information thanks </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2412197,
      "author_name": "Mehmet ISIK",
      "author_url": "",
      "post_date": "2023-08-28T06:27:23.170000",
      "content": "<p>great. thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2410224,
      "author_name": "Ahmad Zainul",
      "author_url": "",
      "post_date": "2023-08-26T18:19:32.567000",
      "content": "<p>Thank You for Sharing!!!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2407193": "Thanks Kaggle and Google for hosting a fun ASL competition.\n\nI joined a few days ago, so I didn't have time to build my own model. Instead, I began with the best public notebook [here][2] (version 17) and attempted to improve it. After making a few changes, I boosted the LB from 0.700 to an amazing 0.770! and obtained Silver medal !!\n\n# Change Two Lines of Code\nThe best public notebook is Rohith Ingilela's awesome public notebook [here][1] which was improved by Saidineshpola [here][2]. Version 17 of Saidineshpola's notebook achieves CV = 0.689 (scroll to bottom of version 17) and LB = 0.697.\n\nThat notebook uses TF Records made by Rohith Ingilela [here][3]. An easy trick to boost the performance of the public notebook is create TF Records which keep frames where hands are missing. The TF Records made by linked notebook removes all frames without hands. Instead we can use `\"output two\"` below and keep 50% of the frames with missing hands below to boost CV and LB. Updated notebook published [here][4].\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2023/asl_preprocess.png)\n\nHere is preprocess code to use when making TF Records and during inference:\n\n    hand = tf.concat([rhand, lhand], axis=1)\n    hand = tf.where(tf.math.is_nan(hand), 0.0, hand)\n    mask = tf.math.not_equal(tf.reduce_sum(hand, axis=[1, 2]), 0.0)\n    alternating_tensor = tf.math.equal( tf.cumsum(\n        tf.ones_like( tf.reduce_sum(hand, axis=[1, 2]) ))%2, 1.0 )\n    mask = tf.math.logical_or(mask, alternating_tensor)\n\n# CTC (Connectionist Temporal Classification) Loss:\nIn this competition, I learned about CTC loss. This loss is amazing! It allows us to create a model with variable length input and predict variable length output in one step. So we can do things like seq2seq without waiting for a sequential decoder to decode each step. Instead we predict the entire output at once super fast! Giving the model some of the frames with hands missing helps identify duplicates and transistions:\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2023/asl_ctc.png)\n\n# Time Augmentation LB +0.004!\nThe public notebook trains for 50 epochs, I found that training for more epochs continues to boost CV and LB score! My final submission trains for 200 epochs. Furthermore, we can add augmentation (i.e. regularization) which helps the model train longer, prevent overfitting, and generalize better.\n\nBelow we see the histogram of train data Number of Frames divided by Character Length of Target Phrase. From this plot, we see that the ratio varies a lot. Some videos have a different recorded frame rate than others and some participants sign faster than others.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2023/frame_hist.png)\n\nWhat this means is that the model has a hard time transfer learning one frame rate to learn about another frame rate. We can help the model by using time augmentation. For each input sequence, we can randomly shrink frame length by 50% or enlarge 150%. This will add lots of new train data and help the model learn about different frame rates. Here is augmentation code:\n\n    if tf.random.uniform(shape=(), minval=0, maxval=1)<0.2:\n        new_height = tf.math.round( tf.random.uniform(\n            shape=(), minval = tf.cast(tf.shape(lip)[0],tf.float32) / 2.0, \n            maxval = tf.cast(tf.shape(lip)[0],tf.float32) * 1.5) )\n        for x in [lip, rhand, lhand, rpose, lpose]:\n            x = tf.image.resize(x, (new_height, tf.shape(x)[1]) )\n\n# Post Processing LB +0.004!\nSometimes the model doesn't make a good prediction. The shortest train data is length 3. When the model predicts length 2 or less, we know it is a bad prediction. Therefore we can replace bad predictions with the best constant length prediction. Anokas found the best constant length prediction [here][5]. We add this to our TF Lite mode with the following code:\n\n    x = tf.cond(tf.shape(x)[0] < 3, lambda: tf.constant(\n        [17, 0, 32, 12, 36, 0, 12, 32, 49, 46, 36], tf.int64), lambda: tf.identity(x))\n\n# Solution Code\nI published a Kaggle notebook [here][4] demonstrating the above 3 changes. The first four bullet points below achieve `LB = 0.763`, then time augmentation boost `+0.004` and PP boost `+0.004`. Note that the second bullet point (about batch size) just makes things faster but doesn't change the CV nor LB score. The other bullet points are the key:\n\n* Change 2 lines of code to keep 50% missing hand frames\n* Change batch size from 32 to 128 and learning rate from 1e-3 to 4e-3\n* Change FRAME_LEN from 128 to 216\n* Increase epochs 50 to 200\n* Add time augmenation\n* Add post process for preds less than 3 chars\n\n[1]: https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\n[2]: https://www.kaggle.com/code/saidineshpola/aslfr-ctc-based-on-prev-comp-1st-place?scriptVersionId=139048726\n[3]: https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset-tfrecords-mean-std\n[4]: https://www.kaggle.com/code/cdeotte/2-lines-of-code-change-lb-0-760\n[5]: https://www.kaggle.com/code/anokas/static-greedy-baseline-0-157-lb",
    "2407214": "I used:\n```\nif len(pred) <= 4:\n    pred = pred + \" -aero\"\n```\nit gives me +0.003 in CV and LB",
    "2556020": "This makes so much sense! Was reading up on CTC Loss and this cleared so many things, extremely grateful!",
    "2407199": "good work!\n\nthe reason why  original frame length 128 does not work is because it is \"too short\". you can check the error of the predicted phrase, many error error are truncation at the end. if you just change 128 to 256 (without dropping missing), you will see immediate improvement. \n\ndropping missing frame while keeping it at 128 is another way.\n\nit is important to consider the need (or the need not) to adjust kernel size in 1d conv when you change the temporal length or resolution",
    "2423241": "I am going to work on it as well and as always your knowledge and information helps a lot. Thanks for sharing @cdeotte ",
    "2422399": "Your explanation is very clear, thank you for your work!",
    "2418460": "Congratulations and thank you for sharing!",
    "2418200": "I'm working on a similar project. This dicussion is going to help me a lot, thank you for sharing! ",
    "2418020": "man what a great way to get to silver! Love that we can help each other grow, good job!",
    "2416415": "Thanks for sharing Chris - this was really informative!",
    "2414443": "Easy as hell, well done !!!",
    "2413636": "congrats and thanks for sharing the elegant approach!",
    "2413026": "This is very nice effort sir!",
    "2412748": "I also considered whether I should remove some frames without hands, but I didn't proceed to validate this idea, which is quite unfortunate. I'm curious to know how much removing 50% of frames without hands would improve the score on the public test set compared to not removing these frames at all. Also, I'm interested in understanding how much faster the training process would be.",
    "2409075": "Great Sir !",
    "2408658": "Thanks for sharing !! It is really helpful and informative",
    "2408306": "Congratulations. Thanks for sharing your innovative approach. \nVery interesting code improvements and insights. ",
    "2407751": "Thank you for the tips. This will help a lot in upcoming challenges. My thinking was that removing the frames with missing hands while using positional encodings wouldn't be good for the model.",
    "2407348": "Very helpful!",
    "2407197": "Congratulations @cdeotte \nCan you tell me what machine are you using to train your model? Do you have your own or you use cloud one?",
    "2407622": "These are very clever tricks. Out of curiosity, how did you find out that using frames with missing hands was better? Was it by luck or you had an intuition behind it? Anyways, congratulations on the silver on such a short time.🥳",
    "2407210": "Congratulations and thanks for your solution",
    "2407445": "Well, so cool for the missing hands dealing, I only found do not throw miss hands frames work much better, I think your solution might be much better for long input frames. For less input frames as non hand feature might also help I think we might not modify them. Not sure, maybe your method work best, I will give it a try, thanks for sharing!",
    "2408181": "Congratulations on your Silver medal in the ASL competition! Your two-line code changes leading to a remarkable LB boost are impressive. Your combination of preprocessing, CTC loss, time augmentation, and post-processing displays a solid understanding of techniques. Sharing your journey with clear explanations is valuable for the community. Your success is inspiring, encouraging others to experiment and learn. Well done!",
    "2407326": "",
    "2412525": "Nice information thanks ",
    "2412197": "great. thanks for sharing",
    "2410224": "Thank You for Sharing!!!\n"
  }
}