{
  "id": 433744,
  "title": "How low is your train loss/validation loss?",
  "url": "/competitions/asl-fingerspelling/discussion/433744",
  "author_name": "Yassine Alouini",
  "post_date": "2023-08-22T18:26:38.509000",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>For now I can get to <strong>17 train loss</strong> vs <strong>16 validation loss</strong> but I feel it can get much lower. <br>\nI am using <strong>CTC loss</strong> for now. By the way, here is a great explanation of the loss <a href=\"https://paperswithcode.com/method/ctc-loss\" target=\"_blank\">here</a>.</p>",
  "messages": [
    {
      "id": 2404166,
      "postDate": "2023-08-23T05:19:03.607Z",
      "content": "<p>Have around 9 train loss and 10 validation loss but hard to improve validation more.<br>\nMore epochs helps but perhaps better learning rate scheduler is needed.  There may be other tricks as hengck23 mentions. </p>",
      "rawMarkdown": "Have around 9 train loss and 10 validation loss but hard to improve validation more.\nMore epochs helps but perhaps better learning rate scheduler is needed.  There may be other tricks as hengck23 mentions. ",
      "votes": 1,
      "replies": [
        {
          "id": 2404191,
          "postDate": "2023-08-23T05:38:34.287Z",
          "content": "<p>Yes, I'm at 11/11 even with augmentation. Judging by the top lb scores, they would have managed to get the CTC loss below 2 or 1 (If they are using CTC that is) Perhaps they pretrained on another dataset (like the supp).</p>",
          "rawMarkdown": "Yes, I'm at 11/11 even with augmentation. Judging by the top lb scores, they would have managed to get the CTC loss below 2 or 1 (If they are using CTC that is) Perhaps they pretrained on another dataset (like the supp).",
          "replies": [
            {
              "id": 2404292,
              "postDate": "2023-08-23T06:56:41.790Z",
              "content": "<p>may not be true.<br>\nif you talking about long epoach training it works like this:</p>\n<pre><code> epoch training (i.e.  regularisation)\nloss decrease becuase  FEW corrected samples gets higher probabiility ... e.g. train pos samples  up  \n\n epoch training (i.e.  regularisation)\nloss decrease becuase  more samples  corrected,  necssary high probability  ... e.g. train pos samples  up  low , there is little fp\n</code></pre>\n<p>an example is is label smoothing. you restrict the target probability to 1-eps, rather than 1<br>\nlabel smooth has better accuray but higher loss</p>\n<p>the meaning of regularisation is ask the trainer \"not to focus on (only) reducing the loss\". and ask it to do somthing else. </p>\n<hr>\n<p>\"below 2 or 1 (If they are using CTC that is) \"</p>\n<p>it is not possible to go to very low error without overfitting. so if others are having better results, it could mean that they have better metric score than you at the same loss value (in additional to lower loss value than you) </p>",
              "rawMarkdown": "may not be true.\nif you talking about long epoach training it works like this:\n\n```\nshort epoch training (i.e. without regularisation)\nloss decrease becuase the FEW corrected samples gets higher probabiility ... e.g. train pos samples ends up with 0.999\n\nlong epoch training (i.e. with regularisation)\nloss decrease becuase the more samples get corrected, not necssary high probability  ... e.g. train pos samples ends up with low 0.6, there is little fp\n\n```\n\nan example is is label smoothing. you restrict the target probability to 1-eps, rather than 1\nlabel smooth has better accuray but higher loss\n\nthe meaning of regularisation is ask the trainer \"not to focus on (only) reducing the loss\". and ask it to do somthing else. \n\n---\n\"below 2 or 1 (If they are using CTC that is) \"\n\nit is not possible to go to very low error without overfitting. so if others are having better results, it could mean that they have better metric score than you at the same loss value (in additional to lower loss value than you) \n",
              "votes": 1
            },
            {
              "id": 2404398,
              "postDate": "2023-08-23T08:19:26.990Z",
              "content": "<p>Ah gotcha, makes sense.</p>",
              "rawMarkdown": "Ah gotcha, makes sense.\n\n",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2403940,
      "postDate": "2023-08-23T00:38:20.257Z",
      "content": "<p>in previous kaggle ASL hand sign recognition competition, the top 1st solution uses 400 epoches.</p>\n<p>I think one method (though not the only one) is to think how to train for long epoch without overfitting.<br>\nif you can achieve that (which is actually non-trivial #1) basically you have good generalisation as well.</p>\n<p>i say  \"non-trivial\" becuase if you did not use his tricks mentioned, you actually overfit at 60 epoch (in my experiments to repeat last competition results).</p>",
      "rawMarkdown": "in previous kaggle ASL hand sign recognition competition, the top 1st solution uses 400 epoches.\n\nI think one method (though not the only one) is to think how to train for long epoch without overfitting.\nif you can achieve that (which is actually non-trivial #1) basically you have good generalisation as well.\n\ni say  \"non-trivial\" becuase if you did not use his tricks mentioned, you actually overfit at 60 epoch (in my experiments to repeat last competition results).",
      "votes": 1,
      "replies": [
        {
          "id": 2404987,
          "postDate": "2023-08-23T16:03:57.267Z",
          "content": "<p>Thanks for the details. Indeed, I have tried a 200 epochs training but around 60~70 epochs, it was not improving anymore.<br>\nI was inspired by previous year's solution that trained for 400 epochs. We will find out soon if someone managed to do it without overfitting.</p>",
          "rawMarkdown": "Thanks for the details. Indeed, I have tried a 200 epochs training but around 60~70 epochs, it was not improving anymore.\nI was inspired by previous year's solution that trained for 400 epochs. We will find out soon if someone managed to do it without overfitting."
        }
      ]
    },
    {
      "id": 2403601,
      "postDate": "2023-08-22T18:26:38.510Z",
      "content": "<p>For now I can get to <strong>17 train loss</strong> vs <strong>16 validation loss</strong> but I feel it can get much lower. <br>\nI am using <strong>CTC loss</strong> for now. By the way, here is a great explanation of the loss <a href=\"https://paperswithcode.com/method/ctc-loss\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "For now I can get to **17 train loss** vs **16 validation loss** but I feel it can get much lower. \nI am using **CTC loss** for now. By the way, here is a great explanation of the loss [here](https://paperswithcode.com/method/ctc-loss).",
      "votes": 1
    },
    {
      "id": 2403613,
      "postDate": "2023-08-22T18:38:23.400Z",
      "content": "<p>Alright, it seems that the model can keep learning. Training for more epochs seems to be key. <br>\nGetting to 12 vs 12. </p>",
      "rawMarkdown": "Alright, it seems that the model can keep learning. Training for more epochs seems to be key. \nGetting to 12 vs 12. "
    }
  ],
  "comments": [
    {
      "id": 2404166,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2023-08-23T05:19:03.607000",
      "content": "<p>Have around 9 train loss and 10 validation loss but hard to improve validation more.<br>\nMore epochs helps but perhaps better learning rate scheduler is needed.  There may be other tricks as hengck23 mentions. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2404191,
          "author_name": "Euler",
          "author_url": "",
          "post_date": "2023-08-23T05:38:34.287000",
          "content": "<p>Yes, I'm at 11/11 even with augmentation. Judging by the top lb scores, they would have managed to get the CTC loss below 2 or 1 (If they are using CTC that is) Perhaps they pretrained on another dataset (like the supp).</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2404292,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-08-23T06:56:41.790000",
              "content": "<p>may not be true.<br>\nif you talking about long epoach training it works like this:</p>\n<pre><code> epoch training (i.e.  regularisation)\nloss decrease becuase  FEW corrected samples gets higher probabiility ... e.g. train pos samples  up  \n\n epoch training (i.e.  regularisation)\nloss decrease becuase  more samples  corrected,  necssary high probability  ... e.g. train pos samples  up  low , there is little fp\n</code></pre>\n<p>an example is is label smoothing. you restrict the target probability to 1-eps, rather than 1<br>\nlabel smooth has better accuray but higher loss</p>\n<p>the meaning of regularisation is ask the trainer \"not to focus on (only) reducing the loss\". and ask it to do somthing else. </p>\n<hr>\n<p>\"below 2 or 1 (If they are using CTC that is) \"</p>\n<p>it is not possible to go to very low error without overfitting. so if others are having better results, it could mean that they have better metric score than you at the same loss value (in additional to lower loss value than you) </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2404398,
              "author_name": "Euler",
              "author_url": "",
              "post_date": "2023-08-23T08:19:26.990000",
              "content": "<p>Ah gotcha, makes sense.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2403940,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-23T00:38:20.257000",
      "content": "<p>in previous kaggle ASL hand sign recognition competition, the top 1st solution uses 400 epoches.</p>\n<p>I think one method (though not the only one) is to think how to train for long epoch without overfitting.<br>\nif you can achieve that (which is actually non-trivial #1) basically you have good generalisation as well.</p>\n<p>i say  \"non-trivial\" becuase if you did not use his tricks mentioned, you actually overfit at 60 epoch (in my experiments to repeat last competition results).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2404987,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2023-08-23T16:03:57.267000",
          "content": "<p>Thanks for the details. Indeed, I have tried a 200 epochs training but around 60~70 epochs, it was not improving anymore.<br>\nI was inspired by previous year's solution that trained for 400 epochs. We will find out soon if someone managed to do it without overfitting.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2403613,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2023-08-22T18:38:23.400000",
      "content": "<p>Alright, it seems that the model can keep learning. Training for more epochs seems to be key. <br>\nGetting to 12 vs 12. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2404166": "Have around 9 train loss and 10 validation loss but hard to improve validation more.\nMore epochs helps but perhaps better learning rate scheduler is needed.  There may be other tricks as hengck23 mentions. ",
    "2403940": "in previous kaggle ASL hand sign recognition competition, the top 1st solution uses 400 epoches.\n\nI think one method (though not the only one) is to think how to train for long epoch without overfitting.\nif you can achieve that (which is actually non-trivial #1) basically you have good generalisation as well.\n\ni say  \"non-trivial\" becuase if you did not use his tricks mentioned, you actually overfit at 60 epoch (in my experiments to repeat last competition results).",
    "2403601": "For now I can get to **17 train loss** vs **16 validation loss** but I feel it can get much lower. \nI am using **CTC loss** for now. By the way, here is a great explanation of the loss [here](https://paperswithcode.com/method/ctc-loss).",
    "2403613": "Alright, it seems that the model can keep learning. Training for more epochs seems to be key. \nGetting to 12 vs 12. "
  }
}