{
  "id": 610708,
  "title": "Ensemble Works On Val But Fails On Test",
  "url": "/competitions/grand-xray-slam-division-a/discussion/610708",
  "author_name": "Eli Haciyev",
  "post_date": "2025-10-05T17:23:06.013000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I have two models, both of which produce good results individually: AUC ~0.912/0.921 in validation and 0.923/0.927 in testing. I assembled an ensemble (averaging/optimizing weights) and saw an improvement in validation (0.924), but the AUC dropped significantly in testing (0.89). I'd like to understand the reasons and get advice on proper ensemble design.</p>\n<p>I want to understand why individual models predict better on a test sample than on a validation sample, but on an ensemble of models it turns out to be the opposite.</p>",
  "messages": [
    {
      "id": 3298502,
      "postDate": "2025-10-05T17:23:06.013Z",
      "content": "<p>I have two models, both of which produce good results individually: AUC ~0.912/0.921 in validation and 0.923/0.927 in testing. I assembled an ensemble (averaging/optimizing weights) and saw an improvement in validation (0.924), but the AUC dropped significantly in testing (0.89). I'd like to understand the reasons and get advice on proper ensemble design.</p>\n<p>I want to understand why individual models predict better on a test sample than on a validation sample, but on an ensemble of models it turns out to be the opposite.</p>",
      "rawMarkdown": "I have two models, both of which produce good results individually: AUC ~0.912/0.921 in validation and 0.923/0.927 in testing. I assembled an ensemble (averaging/optimizing weights) and saw an improvement in validation (0.924), but the AUC dropped significantly in testing (0.89). I'd like to understand the reasons and get advice on proper ensemble design.\n\nI want to understand why individual models predict better on a test sample than on a validation sample, but on an ensemble of models it turns out to be the opposite.",
      "votes": 1
    },
    {
      "id": 3298746,
      "postDate": "2025-10-06T08:47:38.703Z",
      "content": "<p><a href=\"https://www.kaggle.com/elihaciyev\" target=\"_blank\">@elihaciyev</a> ,It can happen when models in the ensemble are highly correlated or when the ensemble weights are tuned too closely to the validation data.<br>\nTry checking the correlation between model outputs. if they’re too similar, the ensemble may amplify small errors.<br>\nAlso, ensure the validation split is representative of the test set, and test simpler averaging or cross-validated weights to reduce overfitting.</p>",
      "rawMarkdown": "@elihaciyev ,It can happen when models in the ensemble are highly correlated or when the ensemble weights are tuned too closely to the validation data.\nTry checking the correlation between model outputs. if they’re too similar, the ensemble may amplify small errors.\nAlso, ensure the validation split is representative of the test set, and test simpler averaging or cross-validated weights to reduce overfitting.",
      "replies": [
        {
          "id": 3299235,
          "postDate": "2025-10-07T14:47:32.697Z",
          "content": "<p><a href=\"https://www.kaggle.com/fathibenamor\" target=\"_blank\">@fathibenamor</a>  Thanks for the answer.<br>\n I think the main reason is to make the validation data as similar to the test data as possible, but it's difficult because the task is multi-label classification and you don't know exactly how the data is broken down. </p>",
          "rawMarkdown": "@fathibenamor  Thanks for the answer.\n I think the main reason is to make the validation data as similar to the test data as possible, but it's difficult because the task is multi-label classification and you don't know exactly how the data is broken down. ",
          "replies": [
            {
              "id": 3299257,
              "postDate": "2025-10-07T15:32:22.670Z",
              "content": "<p><a href=\"https://www.kaggle.com/elihaciyev\" target=\"_blank\">@elihaciyev</a> You’re right — in multi-label setups, validation and test distributions can differ quite a bit, especially if the split isn’t perfectly stratified across all label combinations. That can make ensembles overfit to validation patterns that aren’t fully representative.</p>\n<p>You could try building folds using multi-label stratification or verify label co-occurrence balance between train/val. Also, simple averaging often performs more consistently than heavily optimized weights in this case.</p>",
              "rawMarkdown": "@elihaciyev You’re right — in multi-label setups, validation and test distributions can differ quite a bit, especially if the split isn’t perfectly stratified across all label combinations. That can make ensembles overfit to validation patterns that aren’t fully representative.\n\nYou could try building folds using multi-label stratification or verify label co-occurrence balance between train/val. Also, simple averaging often performs more consistently than heavily optimized weights in this case."
            },
            {
              "id": 3299290,
              "postDate": "2025-10-07T17:51:53.417Z",
              "content": "<p><a href=\"https://www.kaggle.com/fathibenamor\" target=\"_blank\">@fathibenamor</a> Thank you for your answer, yes, that's exactly how I broke it with help <strong>MultilabelStratifiedShuffleSplit</strong>, but this is the case I was talking about.<br>\nYes, when I just average the validation results, the estimate is better than when I select weights for each model.</p>",
              "rawMarkdown": "@fathibenamor Thank you for your answer, yes, that's exactly how I broke it with help **MultilabelStratifiedShuffleSplit**, but this is the case I was talking about.\nYes, when I just average the validation results, the estimate is better than when I select weights for each model."
            },
            {
              "id": 3299496,
              "postDate": "2025-10-08T08:52:52.123Z",
              "content": "<p><a href=\"https://www.kaggle.com/elihaciyev\" target=\"_blank\">@elihaciyev</a> Exactly, that makes sense. Weighted ensembles can easily overfit to validation noise, especially in multi-label tasks. Simple averaging often generalizes better unless you have a very large validation set or use cross-validation to tune weights more reliably.</p>",
              "rawMarkdown": "@elihaciyev Exactly, that makes sense. Weighted ensembles can easily overfit to validation noise, especially in multi-label tasks. Simple averaging often generalizes better unless you have a very large validation set or use cross-validation to tune weights more reliably."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3298746,
      "author_name": "Alpha",
      "author_url": "",
      "post_date": "2025-10-06T08:47:38.703000",
      "content": "<p><a href=\"https://www.kaggle.com/elihaciyev\" target=\"_blank\">@elihaciyev</a> ,It can happen when models in the ensemble are highly correlated or when the ensemble weights are tuned too closely to the validation data.<br>\nTry checking the correlation between model outputs. if they’re too similar, the ensemble may amplify small errors.<br>\nAlso, ensure the validation split is representative of the test set, and test simpler averaging or cross-validated weights to reduce overfitting.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3299235,
          "author_name": "Eli Haciyev",
          "author_url": "",
          "post_date": "2025-10-07T14:47:32.697000",
          "content": "<p><a href=\"https://www.kaggle.com/fathibenamor\" target=\"_blank\">@fathibenamor</a>  Thanks for the answer.<br>\n I think the main reason is to make the validation data as similar to the test data as possible, but it's difficult because the task is multi-label classification and you don't know exactly how the data is broken down. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3299257,
              "author_name": "Alpha",
              "author_url": "",
              "post_date": "2025-10-07T15:32:22.670000",
              "content": "<p><a href=\"https://www.kaggle.com/elihaciyev\" target=\"_blank\">@elihaciyev</a> You’re right — in multi-label setups, validation and test distributions can differ quite a bit, especially if the split isn’t perfectly stratified across all label combinations. That can make ensembles overfit to validation patterns that aren’t fully representative.</p>\n<p>You could try building folds using multi-label stratification or verify label co-occurrence balance between train/val. Also, simple averaging often performs more consistently than heavily optimized weights in this case.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3299290,
              "author_name": "Eli Haciyev",
              "author_url": "",
              "post_date": "2025-10-07T17:51:53.417000",
              "content": "<p><a href=\"https://www.kaggle.com/fathibenamor\" target=\"_blank\">@fathibenamor</a> Thank you for your answer, yes, that's exactly how I broke it with help <strong>MultilabelStratifiedShuffleSplit</strong>, but this is the case I was talking about.<br>\nYes, when I just average the validation results, the estimate is better than when I select weights for each model.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3299496,
              "author_name": "Alpha",
              "author_url": "",
              "post_date": "2025-10-08T08:52:52.123000",
              "content": "<p><a href=\"https://www.kaggle.com/elihaciyev\" target=\"_blank\">@elihaciyev</a> Exactly, that makes sense. Weighted ensembles can easily overfit to validation noise, especially in multi-label tasks. Simple averaging often generalizes better unless you have a very large validation set or use cross-validation to tune weights more reliably.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3298502": "I have two models, both of which produce good results individually: AUC ~0.912/0.921 in validation and 0.923/0.927 in testing. I assembled an ensemble (averaging/optimizing weights) and saw an improvement in validation (0.924), but the AUC dropped significantly in testing (0.89). I'd like to understand the reasons and get advice on proper ensemble design.\n\nI want to understand why individual models predict better on a test sample than on a validation sample, but on an ensemble of models it turns out to be the opposite.",
    "3298746": "@elihaciyev ,It can happen when models in the ensemble are highly correlated or when the ensemble weights are tuned too closely to the validation data.\nTry checking the correlation between model outputs. if they’re too similar, the ensemble may amplify small errors.\nAlso, ensure the validation split is representative of the test set, and test simpler averaging or cross-validated weights to reduce overfitting."
  }
}