{
  "id": 433906,
  "title": "Adding validation data (only 1000 samples) into training makes training process much slower",
  "url": "/competitions/asl-fingerspelling/discussion/433906",
  "author_name": "Yu Wu",
  "post_date": "2023-08-23T10:36:09.982000",
  "votes": 0,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi guys,</p>\n<p>Previously, I used 234418913.parquet (1000 samples) as the validation set, and used the remaining data as the training set.</p>\n<p>However, I found adding these 1000 samples into training significantly slows the convergence speed.</p>\n<p>With the validation set, the training score ends up with ~0.86.</p>\n<p>However, using full data ends up with ~0.84 training score under exactly the same setup (LB score also decreased).</p>\n<p>It seems to me I need to increase 1/5~1/6 epochs for training using all data. </p>\n<p>I'm so confused why 1000 samples (only 1/67 of all data) can make such a big difference. 🤯 </p>\n<p>Is this normal? </p>",
  "messages": [
    {
      "id": 2404651,
      "postDate": "2023-08-23T12:17:53.810Z",
      "content": "<p>use previous probability as regularization:</p>\n<pre><code>loss( sample) = ctc_loss_or other + \n</code></pre>",
      "rawMarkdown": "use previous probability as regularization:\n```\nloss(new sample) = ctc_loss_or other_loss(new sample|truth phrase) + KLDiv(old_probability, current_probability)\n\n```",
      "votes": 1,
      "replies": [
        {
          "id": 2404677,
          "postDate": "2023-08-23T12:42:08.437Z",
          "content": "<p>Thank you for your reply!</p>\n<p>what's the definition of probability here?</p>",
          "rawMarkdown": "Thank you for your reply!\n\nwhat's the definition of probability here?",
          "replies": [
            {
              "id": 2405547,
              "postDate": "2023-08-24T01:23:44.437Z",
              "content": "<p>I think he means something like the softmax of the logits there. I belive old_prob is from the past epoch and current_prob is for the current epoch</p>",
              "rawMarkdown": "I think he means something like the softmax of the logits there. I belive old_prob is from the past epoch and current_prob is for the current epoch",
              "votes": 1
            },
            {
              "id": 2405669,
              "postDate": "2023-08-24T03:53:39.393Z",
              "content": "<p>Thanks for your reply! It seems make sense</p>",
              "rawMarkdown": "Thanks for your reply! It seems make sense"
            },
            {
              "id": 2405683,
              "postDate": "2023-08-24T04:09:51.290Z",
              "content": "<p>the objective is to make the prediction close to the previous model (i.e. retain previous model accuracy) when possible.</p>\n<p>old_prob is from previous model.</p>",
              "rawMarkdown": "the objective is to make the prediction close to the previous model (i.e. retain previous model accuracy) when possible.\n\nold_prob is from previous model.\n",
              "votes": 1
            },
            {
              "id": 2405684,
              "postDate": "2023-08-24T04:10:52.400Z",
              "content": "<p>or you can use uniform probability for old_prob</p>\n<p><a href=\"https://discuss.pytorch.org/t/label-smoothing-with-ctcloss/103392\" target=\"_blank\">https://discuss.pytorch.org/t/label-smoothing-with-ctcloss/103392</a><br>\n<a href=\"https://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_label_smoothing_loss.py\" target=\"_blank\">https://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_label_smoothing_loss.py</a></p>",
              "rawMarkdown": "or you can use uniform probability for old_prob\n\nhttps://discuss.pytorch.org/t/label-smoothing-with-ctcloss/103392\nhttps://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_label_smoothing_loss.py",
              "votes": 1
            },
            {
              "id": 2405691,
              "postDate": "2023-08-24T04:17:14.423Z",
              "content": "<p>\"From Fig. 2, it becomes clear that KD acts as an informed label smoothing,\"<br>\nif prob_old is from a model, the model is acting as a teaching.<br>\nif prob_old is from a assumption(e.g. uniform prob), it is a prior </p>\n<p><a href=\"https://arxiv.org/pdf/2005.09310.pdf\" target=\"_blank\">https://arxiv.org/pdf/2005.09310.pdf</a></p>",
              "rawMarkdown": "\"From Fig. 2, it becomes clear that KD acts as an informed label smoothing,\"\nif prob_old is from a model, the model is acting as a teaching.\nif prob_old is from a assumption(e.g. uniform prob), it is a prior \n\nhttps://arxiv.org/pdf/2005.09310.pdf",
              "votes": 1
            },
            {
              "id": 2406019,
              "postDate": "2023-08-24T07:34:43.927Z",
              "content": "<p>looks so advanced. don't know if I have time to try it out at the last minute…</p>",
              "rawMarkdown": "looks so advanced. don't know if I have time to try it out at the last minute..."
            }
          ]
        }
      ]
    },
    {
      "id": 2404584,
      "postDate": "2023-08-23T10:36:09.983Z",
      "content": "<p>Hi guys,</p>\n<p>Previously, I used 234418913.parquet (1000 samples) as the validation set, and used the remaining data as the training set.</p>\n<p>However, I found adding these 1000 samples into training significantly slows the convergence speed.</p>\n<p>With the validation set, the training score ends up with ~0.86.</p>\n<p>However, using full data ends up with ~0.84 training score under exactly the same setup (LB score also decreased).</p>\n<p>It seems to me I need to increase 1/5~1/6 epochs for training using all data. </p>\n<p>I'm so confused why 1000 samples (only 1/67 of all data) can make such a big difference. 🤯 </p>\n<p>Is this normal? </p>",
      "rawMarkdown": "Hi guys,\n\nPreviously, I used 234418913.parquet (1000 samples) as the validation set, and used the remaining data as the training set.\n\nHowever, I found adding these 1000 samples into training significantly slows the convergence speed.\n\nWith the validation set, the training score ends up with ~0.86.\n\nHowever, using full data ends up with ~0.84 training score under exactly the same setup (LB score also decreased).\n\n It seems to me I need to increase 1/5~1/6 epochs for training using all data. \n\nI'm so confused why 1000 samples (only 1/67 of all data) can make such a big difference. 🤯 \n\nIs this normal? "
    }
  ],
  "comments": [
    {
      "id": 2404651,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-08-23T12:17:53.810000",
      "content": "<p>use previous probability as regularization:</p>\n<pre><code>loss( sample) = ctc_loss_or other + \n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 2404677,
          "author_name": "Yu Wu",
          "author_url": "",
          "post_date": "2023-08-23T12:42:08.437000",
          "content": "<p>Thank you for your reply!</p>\n<p>what's the definition of probability here?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2405547,
              "author_name": "Adriano Passos",
              "author_url": "",
              "post_date": "2023-08-24T01:23:44.437000",
              "content": "<p>I think he means something like the softmax of the logits there. I belive old_prob is from the past epoch and current_prob is for the current epoch</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2405669,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-08-24T03:53:39.393000",
              "content": "<p>Thanks for your reply! It seems make sense</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2405683,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-08-24T04:09:51.290000",
              "content": "<p>the objective is to make the prediction close to the previous model (i.e. retain previous model accuracy) when possible.</p>\n<p>old_prob is from previous model.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2405684,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-08-24T04:10:52.400000",
              "content": "<p>or you can use uniform probability for old_prob</p>\n<p><a href=\"https://discuss.pytorch.org/t/label-smoothing-with-ctcloss/103392\" target=\"_blank\">https://discuss.pytorch.org/t/label-smoothing-with-ctcloss/103392</a><br>\n<a href=\"https://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_label_smoothing_loss.py\" target=\"_blank\">https://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_label_smoothing_loss.py</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2405691,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-08-24T04:17:14.423000",
              "content": "<p>\"From Fig. 2, it becomes clear that KD acts as an informed label smoothing,\"<br>\nif prob_old is from a model, the model is acting as a teaching.<br>\nif prob_old is from a assumption(e.g. uniform prob), it is a prior </p>\n<p><a href=\"https://arxiv.org/pdf/2005.09310.pdf\" target=\"_blank\">https://arxiv.org/pdf/2005.09310.pdf</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2406019,
              "author_name": "Yu Wu",
              "author_url": "",
              "post_date": "2023-08-24T07:34:43.927000",
              "content": "<p>looks so advanced. don't know if I have time to try it out at the last minute…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2404651": "use previous probability as regularization:\n```\nloss(new sample) = ctc_loss_or other_loss(new sample|truth phrase) + KLDiv(old_probability, current_probability)\n\n```",
    "2404584": "Hi guys,\n\nPreviously, I used 234418913.parquet (1000 samples) as the validation set, and used the remaining data as the training set.\n\nHowever, I found adding these 1000 samples into training significantly slows the convergence speed.\n\nWith the validation set, the training score ends up with ~0.86.\n\nHowever, using full data ends up with ~0.84 training score under exactly the same setup (LB score also decreased).\n\n It seems to me I need to increase 1/5~1/6 epochs for training using all data. \n\nI'm so confused why 1000 samples (only 1/67 of all data) can make such a big difference. 🤯 \n\nIs this normal? "
  }
}