{
  "id": 434363,
  "title": "12th solution- augmentation and more epochs",
  "url": "/competitions/asl-fingerspelling/discussion/434363",
  "author_name": "greySnow",
  "post_date": "2023-08-25T00:59:34.494000",
  "votes": 24,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I will clean and probably publish my colab notebook in a few days, but here is the gist of things:<br>\nAt first, I took the 1st solution from the <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">previous competition</a>, which includes a lot of augmentations. I made a few changes and some adjustments, partly based on ideas from <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">Rohith's notebook</a> (mainly to use CTC loss and the size of the model). I trained on samples that had (No nan hands frames number)&gt; 2*phrase_len with ~300 epochs. This model got a 0.779 public LB score.<br>\nThen, I made my model much larger (~15M parameters) and trained for ~500 epochs. This change got me to 0.79 LB.<br>\nFinally, I added to the training set all the remaining samples (besides my 3000 samples validation set) on one condition: that my 0.779 LB model prediction on them has normalized Levenshtein distance &gt; 0.2. Then I trained for 1500 epochs, haha. Well, past ~1000 epochs, there was no improvement already. My submission was the ~1200 epoch or so with 0.794 LB.<br>\nI want to thank <a href=\"https://www.kaggle.com/irohith\" target=\"_blank\">@irohith</a> and <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a>; their notebooks were a big part of my success. Thank you!</p>",
  "messages": [
    {
      "id": 2407225,
      "postDate": "2023-08-25T00:59:34.493Z",
      "content": "<p>I will clean and probably publish my colab notebook in a few days, but here is the gist of things:<br>\nAt first, I took the 1st solution from the <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">previous competition</a>, which includes a lot of augmentations. I made a few changes and some adjustments, partly based on ideas from <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">Rohith's notebook</a> (mainly to use CTC loss and the size of the model). I trained on samples that had (No nan hands frames number)&gt; 2*phrase_len with ~300 epochs. This model got a 0.779 public LB score.<br>\nThen, I made my model much larger (~15M parameters) and trained for ~500 epochs. This change got me to 0.79 LB.<br>\nFinally, I added to the training set all the remaining samples (besides my 3000 samples validation set) on one condition: that my 0.779 LB model prediction on them has normalized Levenshtein distance &gt; 0.2. Then I trained for 1500 epochs, haha. Well, past ~1000 epochs, there was no improvement already. My submission was the ~1200 epoch or so with 0.794 LB.<br>\nI want to thank <a href=\"https://www.kaggle.com/irohith\" target=\"_blank\">@irohith</a> and <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a>; their notebooks were a big part of my success. Thank you!</p>",
      "rawMarkdown": "I will clean and probably publish my colab notebook in a few days, but here is the gist of things:\nAt first, I took the 1st solution from the [previous competition](https://www.kaggle.com/competitions/asl-signs/discussion/406684), which includes a lot of augmentations. I made a few changes and some adjustments, partly based on ideas from [Rohith's notebook](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place) (mainly to use CTC loss and the size of the model). I trained on samples that had (No nan hands frames number)> 2*phrase_len with ~300 epochs. This model got a 0.779 public LB score.\nThen, I made my model much larger (~15M parameters) and trained for ~500 epochs. This change got me to 0.79 LB.\nFinally, I added to the training set all the remaining samples (besides my 3000 samples validation set) on one condition: that my 0.779 LB model prediction on them has normalized Levenshtein distance > 0.2. Then I trained for 1500 epochs, haha. Well, past ~1000 epochs, there was no improvement already. My submission was the ~1200 epoch or so with 0.794 LB.\nI want to thank @irohith and @hoyso48; their notebooks were a big part of my success. Thank you!",
      "votes": 24
    },
    {
      "id": 2407235,
      "postDate": "2023-08-25T01:19:51.220Z",
      "content": "<p>Congrats on the solo gold! </p>\n<p>it seems I need more epochs (I only use 500 epochs) 😢</p>",
      "rawMarkdown": "Congrats on the solo gold! \n\nit seems I need more epochs (I only use 500 epochs) 😢",
      "votes": 1,
      "replies": [
        {
          "id": 2407256,
          "postDate": "2023-08-25T01:47:13.857Z",
          "content": "<p>Thank you! Sometimes more epochs is all you need 🤣 It was actually a bit strange; the validation CTC loss started to increase after several hundred epochs (i.e., overfitting), but the normalized Levenshtein distance (i.e., the metric) kept increasing, so I kept training. Maybe someone with more knowledge can explain this discrepancy.</p>",
          "rawMarkdown": "Thank you! Sometimes more epochs is all you need 🤣 It was actually a bit strange; the validation CTC loss started to increase after several hundred epochs (i.e., overfitting), but the normalized Levenshtein distance (i.e., the metric) kept increasing, so I kept training. Maybe someone with more knowledge can explain this discrepancy.",
          "replies": [
            {
              "id": 2407311,
              "postDate": "2023-08-25T02:57:35.133Z",
              "content": "<p>\"t was actually a bit strange; the validation CTC loss started to increase after several hundred epochs \"</p>\n<p>i think it is class imbalance.</p>\n<ol>\n<li><p>first, any machine learning algorithm takes care of majority class first. in this case it is actually the BLANK token id.<br>\n(another example is segmentation with the majority class is background)</p></li>\n<li><p>after the  majority class is correct (and \"balanced\"), the learning starts to looks at the other character. actually both train loss and validation loss are always decreasing, just that the metric don't reflect that (for our case here).</p></li>\n</ol>\n<p>same is observed for image segementation: bce loss and mou metric. the jumps in long epoch are small and \"rare/difficult\" object in segemntation.</p>",
              "rawMarkdown": "\"t was actually a bit strange; the validation CTC loss started to increase after several hundred epochs \"\n\ni think it is class imbalance.\n\n1. first, any machine learning algorithm takes care of majority class first. in this case it is actually the BLANK token id.\n(another example is segmentation with the majority class is background)\n\n2. after the  majority class is correct (and \"balanced\"), the learning starts to looks at the other character. actually both train loss and validation loss are always decreasing, just that the metric don't reflect that (for our case here).\n\nsame is observed for image segementation: bce loss and mou metric. the jumps in long epoch are small and \"rare/difficult\" object in segemntation.\n"
            },
            {
              "id": 2407334,
              "postDate": "2023-08-25T03:20:48.727Z",
              "content": "<p>It's the opposite here. You suggest that 'train loss and validation loss are always decreasing just that the metric don't reflect that\", but in my case it is the validation loss that start to increase instead of decreasing, which is a clear indication of overfitting, and the metric is the one that keep increasing (which is good in our case- we need higher metric score).<br>\nActually it was a bit more general- I would say that, the correlation between ctc loss and the Levenshtein metric was far from good. I had models with 10 ctc loss and 0.7 Levenshtein distance, and models with 20 ctc loss and 0.78 Levenshtein distance (not exact numbers but approximately)- on the exact same validation set.</p>",
              "rawMarkdown": "It's the opposite here. You suggest that 'train loss and validation loss are always decreasing just that the metric don't reflect that\", but in my case it is the validation loss that start to increase instead of decreasing, which is a clear indication of overfitting, and the metric is the one that keep increasing (which is good in our case- we need higher metric score).\nActually it was a bit more general- I would say that, the correlation between ctc loss and the Levenshtein metric was far from good. I had models with 10 ctc loss and 0.7 Levenshtein distance, and models with 20 ctc loss and 0.78 Levenshtein distance (not exact numbers but approximately)- on the exact same validation set."
            }
          ]
        }
      ]
    },
    {
      "id": 2408335,
      "postDate": "2023-08-25T15:19:54.193Z",
      "content": "<p>Congratulations. <br>\nWith respect to migrating from Kaggle notebook to colab notebook, can you share the details. I tried to move into colab but faced the issue of large data set.<br>\nWith the GPU limit on Kaggle notebooks, the number of experiments we could do was limited. </p>",
      "rawMarkdown": "Congratulations. \nWith respect to migrating from Kaggle notebook to colab notebook, can you share the details. I tried to move into colab but faced the issue of large data set.\nWith the GPU limit on Kaggle notebooks, the number of experiments we could do was limited. ",
      "replies": [
        {
          "id": 2408524,
          "postDate": "2023-08-25T17:16:36.527Z",
          "content": "<p>I created the tfrecords in kaggle. If you saw public notebooks, after taking only the necessary points it's ~6 GB, which is very manageable.</p>",
          "rawMarkdown": "I created the tfrecords in kaggle. If you saw public notebooks, after taking only the necessary points it's ~6 GB, which is very manageable."
        }
      ]
    },
    {
      "id": 2407288,
      "postDate": "2023-08-25T02:33:45.647Z",
      "content": "<p>Thanks for sharing your approach. It looks like big LB improvement in top performers coming from the ability to select the correct number of frames before padding. may i know the reason behind this </p>\n<blockquote>\n  <p>I trained on samples that had (No nan hands frames number)&gt;2*phrase_len with ~300 epochs. </p>\n</blockquote>",
      "rawMarkdown": "Thanks for sharing your approach. It looks like big LB improvement in top performers coming from the ability to select the correct number of frames before padding. may i know the reason behind this \n>I trained on samples that had (No nan hands frames number)>2*phrase_len with ~300 epochs. ",
      "replies": [
        {
          "id": 2407291,
          "postDate": "2023-08-25T02:40:33.680Z",
          "content": "<p>google for NAN for CTC loss.</p>\n<p>CTC throws NAN if the prediction length is too short. prediction length must be longer than target length. 2x is assuming that word is all repeated single character and you need to predict the same number of BLANK</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F687056ee7cda840c4fd160f29e43157c%2FSelection_999(2962).png?generation=1692931659483900&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "google for NAN for CTC loss.\n\nCTC throws NAN if the prediction length is too short. prediction length must be longer than target length. 2x is assuming that word is all repeated single character and you need to predict the same number of BLANK\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F687056ee7cda840c4fd160f29e43157c%2FSelection_999(2962).png?generation=1692931659483900&alt=media)",
          "replies": [
            {
              "id": 2407294,
              "postDate": "2023-08-25T02:43:08.507Z",
              "content": "<p>good implementation from scratch for CTC loss</p>\n<p><a href=\"https://github.com/vadimkantorov/ctc\" target=\"_blank\">https://github.com/vadimkantorov/ctc</a></p>\n<p>Slides: <a href=\"https://goo.gl/KwWR48\" target=\"_blank\">https://goo.gl/KwWR48</a><br>\n<a href=\"https://www.youtube.com/watch?v=eYIL4TMAeRI&amp;t=675s\" target=\"_blank\">https://www.youtube.com/watch?v=eYIL4TMAeRI&amp;t=675s</a></p>",
              "rawMarkdown": "good implementation from scratch for CTC loss\n\nhttps://github.com/vadimkantorov/ctc\n\nSlides: https://goo.gl/KwWR48\nhttps://www.youtube.com/watch?v=eYIL4TMAeRI&t=675s\n"
            }
          ]
        },
        {
          "id": 2407292,
          "postDate": "2023-08-25T02:41:20.030Z",
          "content": "<p>This is not selecting number of frames before padding. Rather, it is selecting good samples to train on- since there are bad samples that are better to throw away. Note that each sample is a sequence of a lot of frames.</p>",
          "rawMarkdown": "This is not selecting number of frames before padding. Rather, it is selecting good samples to train on- since there are bad samples that are better to throw away. Note that each sample is a sequence of a lot of frames.",
          "replies": [
            {
              "id": 2408271,
              "postDate": "2023-08-25T14:34:37.340Z",
              "content": "<p>Congratulations for your solo gold and thanks for sharing.</p>\n<p>I was surprised to see that picking a larger ratio of hand frames versus target length (2x) makes a difference. My first experiments in this competition were focused on the data, including how to handle missing hand frames. My expectation was that a high ratio would improve learning: how can the model learn to predict 20 characters if it only has 5 frames?</p>\n<p>My assumptions are often wrong, so I tested multiple ratios: 0, 0.1, 0.2, 0.5, 1, 2, 3, 4. To my surprise, it didn't seem to make much difference. Moreover, smaller ratios performed a bit better. I hypothesized that perhaps ctc could still learn by matching whatever characters it could. In the example above, if it can find say 5 characters using the 5 frames, it will miss the other 15, but it can still learn.</p>\n<p>A caveat of my experiments was that I used a relatively small model for just 30 of epochs. It was a risky strategy, but with a limited GPU quota something's gotta give. I wonder if the size of the model + number of epochs led me to the wrong conclusion. Did you measure performance with a smaller ratio?</p>",
              "rawMarkdown": "Congratulations for your solo gold and thanks for sharing.\n\nI was surprised to see that picking a larger ratio of hand frames versus target length (2x) makes a difference. My first experiments in this competition were focused on the data, including how to handle missing hand frames. My expectation was that a high ratio would improve learning: how can the model learn to predict 20 characters if it only has 5 frames?\n\nMy assumptions are often wrong, so I tested multiple ratios: 0, 0.1, 0.2, 0.5, 1, 2, 3, 4. To my surprise, it didn't seem to make much difference. Moreover, smaller ratios performed a bit better. I hypothesized that perhaps ctc could still learn by matching whatever characters it could. In the example above, if it can find say 5 characters using the 5 frames, it will miss the other 15, but it can still learn.\n\nA caveat of my experiments was that I used a relatively small model for just 30 of epochs. It was a risky strategy, but with a limited GPU quota something's gotta give. I wonder if the size of the model + number of epochs led me to the wrong conclusion. Did you measure performance with a smaller ratio?"
            },
            {
              "id": 2408538,
              "postDate": "2023-08-25T17:24:44.543Z",
              "content": "<p>No, I made only two experiments- one with training on all the data, and one with training on the cleaned set as I described. Which performed much better. Then there were also the public notebooks of <a href=\"https://www.kaggle.com/irohith\" target=\"_blank\">@irohith</a> that also found that a cleaned set perform better than whole set. Anyway, since the cleaned set was used in the latter phase to decide which samples from the small ratio are bad, it does not really matter what exact ratio I use in the first phase. I only need a model that is good enough to reasonably discriminate the extremely bad samples. (I used a threshold of 0.2 normalized levenshtein - so my first phase model only had to predict one or two letters correctly for a sample to be deemed good enough to be trained on, I.e. eventually I only threw away <em>really</em> bad samples that my model completely failed on)</p>",
              "rawMarkdown": "No, I made only two experiments- one with training on all the data, and one with training on the cleaned set as I described. Which performed much better. Then there were also the public notebooks of @irohith that also found that a cleaned set perform better than whole set. Anyway, since the cleaned set was used in the latter phase to decide which samples from the small ratio are bad, it does not really matter what exact ratio I use in the first phase. I only need a model that is good enough to reasonably discriminate the extremely bad samples. (I used a threshold of 0.2 normalized levenshtein - so my first phase model only had to predict one or two letters correctly for a sample to be deemed good enough to be trained on, I.e. eventually I only threw away *really* bad samples that my model completely failed on)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2407248,
      "postDate": "2023-08-25T01:36:34.280Z",
      "content": "<p>thanks for the writeup.</p>\n<p>did you use AWP in training? (automatic weight perturbation as in the previous competition) </p>",
      "rawMarkdown": "thanks for the writeup.\n\ndid you use AWP in training? (automatic weight perturbation as in the previous competition) ",
      "replies": [
        {
          "id": 2407257,
          "postDate": "2023-08-25T01:51:31.330Z",
          "content": "<p>Good point. My solution did not use AWP. I started to train an additional model yesterday with AWP, but it was already too close to the deadline, and I could not finish the training in time. Based on what I have seen, I think that the AWP model can get a better score if trained long enough.</p>",
          "rawMarkdown": "Good point. My solution did not use AWP. I started to train an additional model yesterday with AWP, but it was already too close to the deadline, and I could not finish the training in time. Based on what I have seen, I think that the AWP model can get a better score if trained long enough."
        },
        {
          "id": 2407260,
          "postDate": "2023-08-25T01:55:39.063Z",
          "content": "<p>I used AWP, but my model couldn't converge in AWP.</p>",
          "rawMarkdown": "I used AWP, but my model couldn't converge in AWP.",
          "replies": [
            {
              "id": 2407273,
              "postDate": "2023-08-25T02:11:42.913Z",
              "content": "<p>What was your lambda? With 0.2 (the original value Hoiso used) it did not work, but when I decreased it to 0.1, it converged.</p>",
              "rawMarkdown": "What was your lambda? With 0.2 (the original value Hoiso used) it did not work, but when I decreased it to 0.1, it converged."
            },
            {
              "id": 2407274,
              "postDate": "2023-08-25T02:14:18.337Z",
              "content": "<p>Yes, I used 0.2 ……😭</p>",
              "rawMarkdown": "Yes, I used 0.2 ......😭"
            },
            {
              "id": 2407281,
              "postDate": "2023-08-25T02:23:42.580Z",
              "content": "<p>I made the same mistake haha. I tried with 0.2 a week ago and saw it did not converge, so I moved on to other ideas. Yesterday, in the last desperate rush for the gold, I suddenly told myself…well, let's try to decrease it a bit. And it works! 🤣 but…too late for submission before the deadline. But my backup non-AWP model saved me 🤣</p>",
              "rawMarkdown": "I made the same mistake haha. I tried with 0.2 a week ago and saw it did not converge, so I moved on to other ideas. Yesterday, in the last desperate rush for the gold, I suddenly told myself...well, let's try to decrease it a bit. And it works! 🤣 but...too late for submission before the deadline. But my backup non-AWP model saved me 🤣"
            }
          ]
        }
      ]
    },
    {
      "id": 2407246,
      "postDate": "2023-08-25T01:30:01.267Z",
      "content": "<p>This game should be called more epochs competition…😂</p>",
      "rawMarkdown": "This game should be called more epochs competition...😂",
      "replies": [
        {
          "id": 2408436,
          "postDate": "2023-08-25T16:19:00.733Z",
          "content": "<p>There might be some participantant ids both in train and test, but still over 500 epochs training not improve LB, so not a game of overfit.</p>",
          "rawMarkdown": "There might be some participantant ids both in train and test, but still over 500 epochs training not improve LB, so not a game of overfit."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2407235,
      "author_name": "Yu Wu",
      "author_url": "",
      "post_date": "2023-08-25T01:19:51.220000",
      "content": "<p>Congrats on the solo gold! </p>\n<p>it seems I need more epochs (I only use 500 epochs) 😢</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2407256,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2023-08-25T01:47:13.857000",
          "content": "<p>Thank you! Sometimes more epochs is all you need 🤣 It was actually a bit strange; the validation CTC loss started to increase after several hundred epochs (i.e., overfitting), but the normalized Levenshtein distance (i.e., the metric) kept increasing, so I kept training. Maybe someone with more knowledge can explain this discrepancy.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2407311,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-08-25T02:57:35.133000",
              "content": "<p>\"t was actually a bit strange; the validation CTC loss started to increase after several hundred epochs \"</p>\n<p>i think it is class imbalance.</p>\n<ol>\n<li><p>first, any machine learning algorithm takes care of majority class first. in this case it is actually the BLANK token id.<br>\n(another example is segmentation with the majority class is background)</p></li>\n<li><p>after the  majority class is correct (and \"balanced\"), the learning starts to looks at the other character. actually both train loss and validation loss are always decreasing, just that the metric don't reflect that (for our case here).</p></li>\n</ol>\n<p>same is observed for image segementation: bce loss and mou metric. the jumps in long epoch are small and \"rare/difficult\" object in segemntation.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2407334,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2023-08-25T03:20:48.727000",
              "content": "<p>It's the opposite here. You suggest that 'train loss and validation loss are always decreasing just that the metric don't reflect that\", but in my case it is the validation loss that start to increase instead of decreasing, which is a clear indication of overfitting, and the metric is the one that keep increasing (which is good in our case- we need higher metric score).<br>\nActually it was a bit more general- I would say that, the correlation between ctc loss and the Levenshtein metric was far from good. I had models with 10 ctc loss and 0.7 Levenshtein distance, and models with 20 ctc loss and 0.78 Levenshtein distance (not exact numbers but approximately)- on the exact same validation set.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2408335,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-25T15:19:54.193000",
      "content": "<p>Congratulations. <br>\nWith respect to migrating from Kaggle notebook to colab notebook, can you share the details. I tried to move into colab but faced the issue of large data set.<br>\nWith the GPU limit on Kaggle notebooks, the number of experiments we could do was limited. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2408524,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2023-08-25T17:16:36.527000",
          "content": "<p>I created the tfrecords in kaggle. If you saw public notebooks, after taking only the necessary points it's ~6 GB, which is very manageable.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2407288,
      "author_name": "saidineshpola",
      "author_url": "",
      "post_date": "2023-08-25T02:33:45.647000",
      "content": "<p>Thanks for sharing your approach. It looks like big LB improvement in top performers coming from the ability to select the correct number of frames before padding. may i know the reason behind this </p>\n<blockquote>\n  <p>I trained on samples that had (No nan hands frames number)&gt;2*phrase_len with ~300 epochs. </p>\n</blockquote>",
      "votes": 0,
      "replies": [
        {
          "id": 2407291,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-08-25T02:40:33.680000",
          "content": "<p>google for NAN for CTC loss.</p>\n<p>CTC throws NAN if the prediction length is too short. prediction length must be longer than target length. 2x is assuming that word is all repeated single character and you need to predict the same number of BLANK</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F687056ee7cda840c4fd160f29e43157c%2FSelection_999(2962).png?generation=1692931659483900&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2407294,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-08-25T02:43:08.507000",
              "content": "<p>good implementation from scratch for CTC loss</p>\n<p><a href=\"https://github.com/vadimkantorov/ctc\" target=\"_blank\">https://github.com/vadimkantorov/ctc</a></p>\n<p>Slides: <a href=\"https://goo.gl/KwWR48\" target=\"_blank\">https://goo.gl/KwWR48</a><br>\n<a href=\"https://www.youtube.com/watch?v=eYIL4TMAeRI&amp;t=675s\" target=\"_blank\">https://www.youtube.com/watch?v=eYIL4TMAeRI&amp;t=675s</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2407292,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2023-08-25T02:41:20.030000",
          "content": "<p>This is not selecting number of frames before padding. Rather, it is selecting good samples to train on- since there are bad samples that are better to throw away. Note that each sample is a sequence of a lot of frames.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2408271,
              "author_name": "vialactea",
              "author_url": "",
              "post_date": "2023-08-25T14:34:37.340000",
              "content": "<p>Congratulations for your solo gold and thanks for sharing.</p>\n<p>I was surprised to see that picking a larger ratio of hand frames versus target length (2x) makes a difference. My first experiments in this competition were focused on the data, including how to handle missing hand frames. My expectation was that a high ratio would improve learning: how can the model learn to predict 20 characters if it only has 5 frames?</p>\n<p>My assumptions are often wrong, so I tested multiple ratios: 0, 0.1, 0.2, 0.5, 1, 2, 3, 4. To my surprise, it didn't seem to make much difference. Moreover, smaller ratios performed a bit better. I hypothesized that perhaps ctc could still learn by matching whatever characters it could. In the example above, if it can find say 5 characters using the 5 frames, it will miss the other 15, but it can still learn.</p>\n<p>A caveat of my experiments was that I used a relatively small model for just 30 of epochs. It was a risky strategy, but with a limited GPU quota something's gotta give. I wonder if the size of the model + number of epochs led me to the wrong conclusion. Did you measure performance with a smaller ratio?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2408538,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2023-08-25T17:24:44.543000",
              "content": "<p>No, I made only two experiments- one with training on all the data, and one with training on the cleaned set as I described. Which performed much better. Then there were also the public notebooks of <a href=\"https://www.kaggle.com/irohith\" target=\"_blank\">@irohith</a> that also found that a cleaned set perform better than whole set. Anyway, since the cleaned set was used in the latter phase to decide which samples from the small ratio are bad, it does not really matter what exact ratio I use in the first phase. I only need a model that is good enough to reasonably discriminate the extremely bad samples. (I used a threshold of 0.2 normalized levenshtein - so my first phase model only had to predict one or two letters correctly for a sample to be deemed good enough to be trained on, I.e. eventually I only threw away <em>really</em> bad samples that my model completely failed on)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2407248,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-25T01:36:34.280000",
      "content": "<p>thanks for the writeup.</p>\n<p>did you use AWP in training? (automatic weight perturbation as in the previous competition) </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2407257,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2023-08-25T01:51:31.330000",
          "content": "<p>Good point. My solution did not use AWP. I started to train an additional model yesterday with AWP, but it was already too close to the deadline, and I could not finish the training in time. Based on what I have seen, I think that the AWP model can get a better score if trained long enough.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2407260,
          "author_name": "Yu Wu",
          "author_url": "",
          "post_date": "2023-08-25T01:55:39.063000",
          "content": "<p>I used AWP, but my model couldn't converge in AWP.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2407273,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2023-08-25T02:11:42.913000",
              "content": "<p>What was your lambda? With 0.2 (the original value Hoiso used) it did not work, but when I decreased it to 0.1, it converged.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2407274,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-08-25T02:14:18.337000",
              "content": "<p>Yes, I used 0.2 ……😭</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2407281,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2023-08-25T02:23:42.580000",
              "content": "<p>I made the same mistake haha. I tried with 0.2 a week ago and saw it did not converge, so I moved on to other ideas. Yesterday, in the last desperate rush for the gold, I suddenly told myself…well, let's try to decrease it a bit. And it works! 🤣 but…too late for submission before the deadline. But my backup non-AWP model saved me 🤣</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2407246,
      "author_name": "Xiang Huang",
      "author_url": "",
      "post_date": "2023-08-25T01:30:01.267000",
      "content": "<p>This game should be called more epochs competition…😂</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2408436,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2023-08-25T16:19:00.733000",
          "content": "<p>There might be some participantant ids both in train and test, but still over 500 epochs training not improve LB, so not a game of overfit.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2407225": "I will clean and probably publish my colab notebook in a few days, but here is the gist of things:\nAt first, I took the 1st solution from the [previous competition](https://www.kaggle.com/competitions/asl-signs/discussion/406684), which includes a lot of augmentations. I made a few changes and some adjustments, partly based on ideas from [Rohith's notebook](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place) (mainly to use CTC loss and the size of the model). I trained on samples that had (No nan hands frames number)> 2*phrase_len with ~300 epochs. This model got a 0.779 public LB score.\nThen, I made my model much larger (~15M parameters) and trained for ~500 epochs. This change got me to 0.79 LB.\nFinally, I added to the training set all the remaining samples (besides my 3000 samples validation set) on one condition: that my 0.779 LB model prediction on them has normalized Levenshtein distance > 0.2. Then I trained for 1500 epochs, haha. Well, past ~1000 epochs, there was no improvement already. My submission was the ~1200 epoch or so with 0.794 LB.\nI want to thank @irohith and @hoyso48; their notebooks were a big part of my success. Thank you!",
    "2407235": "Congrats on the solo gold! \n\nit seems I need more epochs (I only use 500 epochs) 😢",
    "2408335": "Congratulations. \nWith respect to migrating from Kaggle notebook to colab notebook, can you share the details. I tried to move into colab but faced the issue of large data set.\nWith the GPU limit on Kaggle notebooks, the number of experiments we could do was limited. ",
    "2407288": "Thanks for sharing your approach. It looks like big LB improvement in top performers coming from the ability to select the correct number of frames before padding. may i know the reason behind this \n>I trained on samples that had (No nan hands frames number)>2*phrase_len with ~300 epochs. ",
    "2407248": "thanks for the writeup.\n\ndid you use AWP in training? (automatic weight perturbation as in the previous competition) ",
    "2407246": "This game should be called more epochs competition...😂"
  }
}