{
  "id": 217958,
  "title": "Cross validation and ensembling for object detection",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/217958",
  "author_name": "Mostafa Ibrahim",
  "post_date": "2021-02-08T22:56:18.312000",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am a bit confused about cross-validation in this competition. From my basic understanding of cross-validation, u typically have 5 folds, for example, 4 training folds and 1 validation fold, and then u would run the model on all possible training/validation folds and average the results (so that the model would run 5 times in this example). (Please correct me if I am wrong, and I did read a lot about cross-validation).</p>\n<p>My question is, generally, if you have x folds then you must train the model x times right? Because when I was looking at this yolo notebook (<a href=\"https://www.kaggle.com/awsaf49/vinbigdata-cxr-ad-yolov5-14-class-train\" target=\"_blank\">https://www.kaggle.com/awsaf49/vinbigdata-cxr-ad-yolov5-14-class-train</a>) I didn't understand how it was implementing cross-validation because although the author was correctly placing the training images from the 3 training folds and the validation images from the 1 validation fold, he only ran the model once.</p>\n<p>Another question is about ensembling. There seems to be a widely used ensembling method for objection detection called Non maximum suppression and a bunch of others.Are standard averaging methods for predictions (ensembling) valid here? Or are they invalid and that's why there are a bunch  of other methods (<a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">https://github.com/ZFTurbo/Weighted-Boxes-Fusion</a>)</p>",
  "messages": [
    {
      "id": 1192093,
      "postDate": "2021-02-08T23:24:56.950Z",
      "content": "<p>Cross-validation does indeed involve fitting your model for each fold that you use and creating the out of fold predictions on the records the model was not trained on. For the notebook you linked, I believe the author trained one fold per notebook version (so, yes, the code you see in one version does indeed only train a single fold).</p>\n<p>For how to combine predictions of different models, I have a lot less experience with that than for other prediction tasks, but I found <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/70421\" target=\"_blank\">this winning solution</a> to a past competition interesting that used one approach that sounds pretty sensible. They used <a href=\"https://github.com/ahrnbom/ensemble-objdet\" target=\"_blank\">a relatively simple approach that takes averages of the positions and sums the confidences of bounding boxes for the same class</a> (with a threshold in terms of IOU - which you can tune - where two boxes are considered the same) and weighted by the fraction of test-time-augmentations and/or models that contain a box. With cross-validation, you would see what choice of averaging weights and IOU threshold gives you the best fit to the out-of-fold predictions you created.</p>",
      "rawMarkdown": "Cross-validation does indeed involve fitting your model for each fold that you use and creating the out of fold predictions on the records the model was not trained on. For the notebook you linked, I believe the author trained one fold per notebook version (so, yes, the code you see in one version does indeed only train a single fold).\n\nFor how to combine predictions of different models, I have a lot less experience with that than for other prediction tasks, but I found [this winning solution](https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/70421) to a past competition interesting that used one approach that sounds pretty sensible. They used [a relatively simple approach that takes averages of the positions and sums the confidences of bounding boxes for the same class](https://github.com/ahrnbom/ensemble-objdet) (with a threshold in terms of IOU - which you can tune - where two boxes are considered the same) and weighted by the fraction of test-time-augmentations and/or models that contain a box. With cross-validation, you would see what choice of averaging weights and IOU threshold gives you the best fit to the out-of-fold predictions you created.",
      "votes": 3,
      "replies": [
        {
          "id": 1205063,
          "postDate": "2021-02-16T14:13:25.157Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1192084,
      "postDate": "2021-02-08T22:56:18.313Z",
      "content": "<p>I am a bit confused about cross-validation in this competition. From my basic understanding of cross-validation, u typically have 5 folds, for example, 4 training folds and 1 validation fold, and then u would run the model on all possible training/validation folds and average the results (so that the model would run 5 times in this example). (Please correct me if I am wrong, and I did read a lot about cross-validation).</p>\n<p>My question is, generally, if you have x folds then you must train the model x times right? Because when I was looking at this yolo notebook (<a href=\"https://www.kaggle.com/awsaf49/vinbigdata-cxr-ad-yolov5-14-class-train\" target=\"_blank\">https://www.kaggle.com/awsaf49/vinbigdata-cxr-ad-yolov5-14-class-train</a>) I didn't understand how it was implementing cross-validation because although the author was correctly placing the training images from the 3 training folds and the validation images from the 1 validation fold, he only ran the model once.</p>\n<p>Another question is about ensembling. There seems to be a widely used ensembling method for objection detection called Non maximum suppression and a bunch of others.Are standard averaging methods for predictions (ensembling) valid here? Or are they invalid and that's why there are a bunch  of other methods (<a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">https://github.com/ZFTurbo/Weighted-Boxes-Fusion</a>)</p>",
      "rawMarkdown": "I am a bit confused about cross-validation in this competition. From my basic understanding of cross-validation, u typically have 5 folds, for example, 4 training folds and 1 validation fold, and then u would run the model on all possible training/validation folds and average the results (so that the model would run 5 times in this example). (Please correct me if I am wrong, and I did read a lot about cross-validation).\n\nMy question is, generally, if you have x folds then you must train the model x times right? Because when I was looking at this yolo notebook (https://www.kaggle.com/awsaf49/vinbigdata-cxr-ad-yolov5-14-class-train) I didn't understand how it was implementing cross-validation because although the author was correctly placing the training images from the 3 training folds and the validation images from the 1 validation fold, he only ran the model once.\n\nAnother question is about ensembling. There seems to be a widely used ensembling method for objection detection called Non maximum suppression and a bunch of others.Are standard averaging methods for predictions (ensembling) valid here? Or are they invalid and that's why there are a bunch  of other methods (https://github.com/ZFTurbo/Weighted-Boxes-Fusion)",
      "votes": 4
    }
  ],
  "comments": [
    {
      "id": 1192093,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-02-08T23:24:56.950000",
      "content": "<p>Cross-validation does indeed involve fitting your model for each fold that you use and creating the out of fold predictions on the records the model was not trained on. For the notebook you linked, I believe the author trained one fold per notebook version (so, yes, the code you see in one version does indeed only train a single fold).</p>\n<p>For how to combine predictions of different models, I have a lot less experience with that than for other prediction tasks, but I found <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/70421\" target=\"_blank\">this winning solution</a> to a past competition interesting that used one approach that sounds pretty sensible. They used <a href=\"https://github.com/ahrnbom/ensemble-objdet\" target=\"_blank\">a relatively simple approach that takes averages of the positions and sums the confidences of bounding boxes for the same class</a> (with a threshold in terms of IOU - which you can tune - where two boxes are considered the same) and weighted by the fraction of test-time-augmentations and/or models that contain a box. With cross-validation, you would see what choice of averaging weights and IOU threshold gives you the best fit to the out-of-fold predictions you created.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1205063,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-16T14:13:25.157000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1192093": "Cross-validation does indeed involve fitting your model for each fold that you use and creating the out of fold predictions on the records the model was not trained on. For the notebook you linked, I believe the author trained one fold per notebook version (so, yes, the code you see in one version does indeed only train a single fold).\n\nFor how to combine predictions of different models, I have a lot less experience with that than for other prediction tasks, but I found [this winning solution](https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/70421) to a past competition interesting that used one approach that sounds pretty sensible. They used [a relatively simple approach that takes averages of the positions and sums the confidences of bounding boxes for the same class](https://github.com/ahrnbom/ensemble-objdet) (with a threshold in terms of IOU - which you can tune - where two boxes are considered the same) and weighted by the fraction of test-time-augmentations and/or models that contain a box. With cross-validation, you would see what choice of averaging weights and IOU threshold gives you the best fit to the out-of-fold predictions you created.",
    "1192084": "I am a bit confused about cross-validation in this competition. From my basic understanding of cross-validation, u typically have 5 folds, for example, 4 training folds and 1 validation fold, and then u would run the model on all possible training/validation folds and average the results (so that the model would run 5 times in this example). (Please correct me if I am wrong, and I did read a lot about cross-validation).\n\nMy question is, generally, if you have x folds then you must train the model x times right? Because when I was looking at this yolo notebook (https://www.kaggle.com/awsaf49/vinbigdata-cxr-ad-yolov5-14-class-train) I didn't understand how it was implementing cross-validation because although the author was correctly placing the training images from the 3 training folds and the validation images from the 1 validation fold, he only ran the model once.\n\nAnother question is about ensembling. There seems to be a widely used ensembling method for objection detection called Non maximum suppression and a bunch of others.Are standard averaging methods for predictions (ensembling) valid here? Or are they invalid and that's why there are a bunch  of other methods (https://github.com/ZFTurbo/Weighted-Boxes-Fusion)"
  }
}