{
  "id": 223782,
  "title": "Overfitting & Cross-Validation questions",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/223782",
  "author_name": "Mostafa Ibrahim",
  "post_date": "2021-03-05T14:13:18.005000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>So for some reason my yolov5 model only gets 0.065 while most of the public yolov5 kernels are getting 0.11+. So while I was investigating why, I came across some of these questions:</p>\n<p>1) I am aware that yolo only saves your best weights (or checkpoints if you will), this is to avoid overfitting (I am assuming). So a question about this, if it only saves the weights when the validation loss decreases (right?) then the model can never overfit? Or can it? Cuz I trained mine for around 60 epochs while most notebooks are trained for about 30 (both with pre-trained weights) so am sensing it might have overfitted</p>\n<p>2) Some notebooks only train the models on 1 fold (1 variation of cross-validation). This doesn't make much sense to me… Why would you have cross validation in your notebook if you are only training the model 1 time? I thought the point was to rotate on different folds and then average..? Is there something I am missing here? (for instance, this one <a href=\"https://www.kaggle.com/awsaf49/vinbigdata-cxr-ad-yolov5-14-class-train\" target=\"_blank\">https://www.kaggle.com/awsaf49/vinbigdata-cxr-ad-yolov5-14-class-train</a> )</p>",
  "messages": [
    {
      "id": 1228782,
      "postDate": "2021-03-06T18:57:40.773Z",
      "content": "<p>There's the possibility that you're overfitting the validation data (I don't think that's necessarily the case, but it's possible). If you take a model and change things around somewhat randomly for a very long time and save the model whenever the validation performance improved, you're essentially fitting to the validation data (in a very inefficient way). </p>\n<p>The main reason notebooks would only fit one fold is usually that it would either not finish within the time a notebook can run on Kaggle, or the author was short on GPU / TPU time. Doing a multiple fold CV is, of course, usually the sensible way to go (unless the dataset is huge) - doing one out of, say, five folds is really the same as a 80:20 train-test-split.</p>\n<p>A proper CV has multiple benefits:</p>\n<ul>\n<li>It's a good evaluation of model performance and one of the more sensible ways to decide between models.</li>\n<li>It's a good basis for proper ensembling (whether it's finding a weighted average without overfitting the LB, or fitting another model on top of the predictions).</li>\n<li>Combining models fit to each fold is a way of doing bagging that should outperform a fit from a single train-test-split (how it does versus refitting on all data is less clear).</li>\n</ul>",
      "rawMarkdown": "There's the possibility that you're overfitting the validation data (I don't think that's necessarily the case, but it's possible). If you take a model and change things around somewhat randomly for a very long time and save the model whenever the validation performance improved, you're essentially fitting to the validation data (in a very inefficient way). \n\nThe main reason notebooks would only fit one fold is usually that it would either not finish within the time a notebook can run on Kaggle, or the author was short on GPU / TPU time. Doing a multiple fold CV is, of course, usually the sensible way to go (unless the dataset is huge) - doing one out of, say, five folds is really the same as a 80:20 train-test-split.\n\nA proper CV has multiple benefits:\n* It's a good evaluation of model performance and one of the more sensible ways to decide between models.\n* It's a good basis for proper ensembling (whether it's finding a weighted average without overfitting the LB, or fitting another model on top of the predictions).\n* Combining models fit to each fold is a way of doing bagging that should outperform a fit from a single train-test-split (how it does versus refitting on all data is less clear).",
      "votes": 1
    },
    {
      "id": 1227425,
      "postDate": "2021-03-05T14:13:18.007Z",
      "content": "<p>So for some reason my yolov5 model only gets 0.065 while most of the public yolov5 kernels are getting 0.11+. So while I was investigating why, I came across some of these questions:</p>\n<p>1) I am aware that yolo only saves your best weights (or checkpoints if you will), this is to avoid overfitting (I am assuming). So a question about this, if it only saves the weights when the validation loss decreases (right?) then the model can never overfit? Or can it? Cuz I trained mine for around 60 epochs while most notebooks are trained for about 30 (both with pre-trained weights) so am sensing it might have overfitted</p>\n<p>2) Some notebooks only train the models on 1 fold (1 variation of cross-validation). This doesn't make much sense to me… Why would you have cross validation in your notebook if you are only training the model 1 time? I thought the point was to rotate on different folds and then average..? Is there something I am missing here? (for instance, this one <a href=\"https://www.kaggle.com/awsaf49/vinbigdata-cxr-ad-yolov5-14-class-train\" target=\"_blank\">https://www.kaggle.com/awsaf49/vinbigdata-cxr-ad-yolov5-14-class-train</a> )</p>",
      "rawMarkdown": "So for some reason my yolov5 model only gets 0.065 while most of the public yolov5 kernels are getting 0.11+. So while I was investigating why, I came across some of these questions:\n\n1) I am aware that yolo only saves your best weights (or checkpoints if you will), this is to avoid overfitting (I am assuming). So a question about this, if it only saves the weights when the validation loss decreases (right?) then the model can never overfit? Or can it? Cuz I trained mine for around 60 epochs while most notebooks are trained for about 30 (both with pre-trained weights) so am sensing it might have overfitted\n\n2) Some notebooks only train the models on 1 fold (1 variation of cross-validation). This doesn't make much sense to me... Why would you have cross validation in your notebook if you are only training the model 1 time? I thought the point was to rotate on different folds and then average..? Is there something I am missing here? (for instance, this one https://www.kaggle.com/awsaf49/vinbigdata-cxr-ad-yolov5-14-class-train )",
      "votes": 1
    },
    {
      "id": 1227607,
      "postDate": "2021-03-05T17:13:16.540Z",
      "content": "<p>Probably you need to check again your data. I train yolov5x with default config can get ~ 0.12x</p>",
      "rawMarkdown": "Probably you need to check again your data. I train yolov5x with default config can get ~ 0.12x",
      "votes": 2
    },
    {
      "id": 1239362,
      "postDate": "2021-03-15T16:46:45.437Z",
      "content": "<p>Saving intermediate checkpoints gives you a few benefits:</p>\n<pre><code>Resilience: If you are training for a very long time, or doing distributed training on many machines, the likelihood of machine failure increases. If a machine fails, TensorFlow can resume from the last saved checkpoint instead of having to start from scratch. This behavior is automatic — TensorFlow looks for checkpoints and resumes from the last checkpoint.\nGeneralization: In general, the longer you train, the lower the loss on the training dataset. However, at some point, the error on the held-out, evaluation dataset might stop decreasing. If you have a very large model, and you are not doing sufficient regularization, the error on the evaluation dataset might even start to increase. If this happens, it can be helpful to go back and export the model that had the best validation error. This is also called early stopping because you could stop if you see the validation error start to increase. (A better idea is, of course, to decrease model complexity or increase the regularization so that this scenario doesn’t happen). The only way you can go back to the best validation error or do early stopping is if you have been periodically evaluating and checkpointing the model.\nTuneability: In a well-behaved training loop, gradient descent behaves such that you get to the neighborhood of the optimal error quickly on the basis of the majority of your data and then slowly converge towards the lowest error by optimizing on the corner cases. Now, imagine that you need to periodically retrain the model on fresh data \n</code></pre>",
      "rawMarkdown": "Saving intermediate checkpoints gives you a few benefits:\n\n    Resilience: If you are training for a very long time, or doing distributed training on many machines, the likelihood of machine failure increases. If a machine fails, TensorFlow can resume from the last saved checkpoint instead of having to start from scratch. This behavior is automatic — TensorFlow looks for checkpoints and resumes from the last checkpoint.\n    Generalization: In general, the longer you train, the lower the loss on the training dataset. However, at some point, the error on the held-out, evaluation dataset might stop decreasing. If you have a very large model, and you are not doing sufficient regularization, the error on the evaluation dataset might even start to increase. If this happens, it can be helpful to go back and export the model that had the best validation error. This is also called early stopping because you could stop if you see the validation error start to increase. (A better idea is, of course, to decrease model complexity or increase the regularization so that this scenario doesn’t happen). The only way you can go back to the best validation error or do early stopping is if you have been periodically evaluating and checkpointing the model.\n    Tuneability: In a well-behaved training loop, gradient descent behaves such that you get to the neighborhood of the optimal error quickly on the basis of the majority of your data and then slowly converge towards the lowest error by optimizing on the corner cases. Now, imagine that you need to periodically retrain the model on fresh data "
    }
  ],
  "comments": [
    {
      "id": 1228782,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-03-06T18:57:40.773000",
      "content": "<p>There's the possibility that you're overfitting the validation data (I don't think that's necessarily the case, but it's possible). If you take a model and change things around somewhat randomly for a very long time and save the model whenever the validation performance improved, you're essentially fitting to the validation data (in a very inefficient way). </p>\n<p>The main reason notebooks would only fit one fold is usually that it would either not finish within the time a notebook can run on Kaggle, or the author was short on GPU / TPU time. Doing a multiple fold CV is, of course, usually the sensible way to go (unless the dataset is huge) - doing one out of, say, five folds is really the same as a 80:20 train-test-split.</p>\n<p>A proper CV has multiple benefits:</p>\n<ul>\n<li>It's a good evaluation of model performance and one of the more sensible ways to decide between models.</li>\n<li>It's a good basis for proper ensembling (whether it's finding a weighted average without overfitting the LB, or fitting another model on top of the predictions).</li>\n<li>Combining models fit to each fold is a way of doing bagging that should outperform a fit from a single train-test-split (how it does versus refitting on all data is less clear).</li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1227607,
      "author_name": "Phat Tran",
      "author_url": "",
      "post_date": "2021-03-05T17:13:16.540000",
      "content": "<p>Probably you need to check again your data. I train yolov5x with default config can get ~ 0.12x</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1239362,
      "author_name": "Naim Mhedhbi",
      "author_url": "",
      "post_date": "2021-03-15T16:46:45.437000",
      "content": "<p>Saving intermediate checkpoints gives you a few benefits:</p>\n<pre><code>Resilience: If you are training for a very long time, or doing distributed training on many machines, the likelihood of machine failure increases. If a machine fails, TensorFlow can resume from the last saved checkpoint instead of having to start from scratch. This behavior is automatic — TensorFlow looks for checkpoints and resumes from the last checkpoint.\nGeneralization: In general, the longer you train, the lower the loss on the training dataset. However, at some point, the error on the held-out, evaluation dataset might stop decreasing. If you have a very large model, and you are not doing sufficient regularization, the error on the evaluation dataset might even start to increase. If this happens, it can be helpful to go back and export the model that had the best validation error. This is also called early stopping because you could stop if you see the validation error start to increase. (A better idea is, of course, to decrease model complexity or increase the regularization so that this scenario doesn’t happen). The only way you can go back to the best validation error or do early stopping is if you have been periodically evaluating and checkpointing the model.\nTuneability: In a well-behaved training loop, gradient descent behaves such that you get to the neighborhood of the optimal error quickly on the basis of the majority of your data and then slowly converge towards the lowest error by optimizing on the corner cases. Now, imagine that you need to periodically retrain the model on fresh data \n</code></pre>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1228782": "There's the possibility that you're overfitting the validation data (I don't think that's necessarily the case, but it's possible). If you take a model and change things around somewhat randomly for a very long time and save the model whenever the validation performance improved, you're essentially fitting to the validation data (in a very inefficient way). \n\nThe main reason notebooks would only fit one fold is usually that it would either not finish within the time a notebook can run on Kaggle, or the author was short on GPU / TPU time. Doing a multiple fold CV is, of course, usually the sensible way to go (unless the dataset is huge) - doing one out of, say, five folds is really the same as a 80:20 train-test-split.\n\nA proper CV has multiple benefits:\n* It's a good evaluation of model performance and one of the more sensible ways to decide between models.\n* It's a good basis for proper ensembling (whether it's finding a weighted average without overfitting the LB, or fitting another model on top of the predictions).\n* Combining models fit to each fold is a way of doing bagging that should outperform a fit from a single train-test-split (how it does versus refitting on all data is less clear).",
    "1227425": "So for some reason my yolov5 model only gets 0.065 while most of the public yolov5 kernels are getting 0.11+. So while I was investigating why, I came across some of these questions:\n\n1) I am aware that yolo only saves your best weights (or checkpoints if you will), this is to avoid overfitting (I am assuming). So a question about this, if it only saves the weights when the validation loss decreases (right?) then the model can never overfit? Or can it? Cuz I trained mine for around 60 epochs while most notebooks are trained for about 30 (both with pre-trained weights) so am sensing it might have overfitted\n\n2) Some notebooks only train the models on 1 fold (1 variation of cross-validation). This doesn't make much sense to me... Why would you have cross validation in your notebook if you are only training the model 1 time? I thought the point was to rotate on different folds and then average..? Is there something I am missing here? (for instance, this one https://www.kaggle.com/awsaf49/vinbigdata-cxr-ad-yolov5-14-class-train )",
    "1227607": "Probably you need to check again your data. I train yolov5x with default config can get ~ 0.12x",
    "1239362": "Saving intermediate checkpoints gives you a few benefits:\n\n    Resilience: If you are training for a very long time, or doing distributed training on many machines, the likelihood of machine failure increases. If a machine fails, TensorFlow can resume from the last saved checkpoint instead of having to start from scratch. This behavior is automatic — TensorFlow looks for checkpoints and resumes from the last checkpoint.\n    Generalization: In general, the longer you train, the lower the loss on the training dataset. However, at some point, the error on the held-out, evaluation dataset might stop decreasing. If you have a very large model, and you are not doing sufficient regularization, the error on the evaluation dataset might even start to increase. If this happens, it can be helpful to go back and export the model that had the best validation error. This is also called early stopping because you could stop if you see the validation error start to increase. (A better idea is, of course, to decrease model complexity or increase the regularization so that this scenario doesn’t happen). The only way you can go back to the best validation error or do early stopping is if you have been periodically evaluating and checkpointing the model.\n    Tuneability: In a well-behaved training loop, gradient descent behaves such that you get to the neighborhood of the optimal error quickly on the basis of the majority of your data and then slowly converge towards the lowest error by optimizing on the corner cases. Now, imagine that you need to periodically retrain the model on fresh data "
  }
}