{
  "id": 225391,
  "title": "Best strategy to ensemble detection models",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/225391",
  "author_name": "NakedKoala",
  "post_date": "2021-03-12T03:25:42.849000",
  "votes": 16,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Just wondering what are everyone's ensembling strategies.  Any insight beyond just concat different model's bboxes and then perform NMS ? <br>\n<strong>My observation</strong><br>\nMy best single fold model with TTA achieves 0.26+ on LB.   With NMS-thres=0.5, this model predicts, on average, 49 bboxes per test image.  In single fold setting, TTA leads to incremental gain.<br>\nEnsembling multiple folds the naive way  - concat all bboxes predictions and performing NMS - with the same inference setting as the single fold model lead to a score of 0.257.    The ensemble prediction produces an average of 77 bboxes per test image.<br>\nMy theory for the worse LB score is that there are probably too many bboxes. As a result, the ensemble prediction boosts recall slightly but decreases precision a lot.<br>\nThen, for the same ensemble weights, I use NMS-thres=0.4 and turn off TTA. This change significantly decreases the number of bboxes per test image to 46.  On submission, my LB score improves by 0.01. <br>\n<strong>Lesson Learnt</strong> ( Not sure about its correctness ):<br>\nIf Ensembling model the naive way ( concat bboxes + NMS),  # of bboxes per test image increases along with number of sub-models. <br>\n**So we probably have to do some post-processing to ensure the # of boxes per test image doesn't grow with the number of sub-model and get out of control ? **<br>\nRight now I am just tweaking inference config to bring down the bboxes per test image.  Are there better ways ? <br>\n<strong>Question</strong><br>\nDo you guys observe stable incremental gains when going from single fold -&gt; multiple folds ?<br>\nDo you guys observe stable incremental gains when going from single arch  -&gt; multiple arch?<br>\n**People belonging to large team ( &gt;= 3 persons), how do you guys combine different base models ? **</p>",
  "messages": [
    {
      "id": 1235297,
      "postDate": "2021-03-12T03:25:42.850Z",
      "content": "<p>Just wondering what are everyone's ensembling strategies.  Any insight beyond just concat different model's bboxes and then perform NMS ? <br>\n<strong>My observation</strong><br>\nMy best single fold model with TTA achieves 0.26+ on LB.   With NMS-thres=0.5, this model predicts, on average, 49 bboxes per test image.  In single fold setting, TTA leads to incremental gain.<br>\nEnsembling multiple folds the naive way  - concat all bboxes predictions and performing NMS - with the same inference setting as the single fold model lead to a score of 0.257.    The ensemble prediction produces an average of 77 bboxes per test image.<br>\nMy theory for the worse LB score is that there are probably too many bboxes. As a result, the ensemble prediction boosts recall slightly but decreases precision a lot.<br>\nThen, for the same ensemble weights, I use NMS-thres=0.4 and turn off TTA. This change significantly decreases the number of bboxes per test image to 46.  On submission, my LB score improves by 0.01. <br>\n<strong>Lesson Learnt</strong> ( Not sure about its correctness ):<br>\nIf Ensembling model the naive way ( concat bboxes + NMS),  # of bboxes per test image increases along with number of sub-models. <br>\n**So we probably have to do some post-processing to ensure the # of boxes per test image doesn't grow with the number of sub-model and get out of control ? **<br>\nRight now I am just tweaking inference config to bring down the bboxes per test image.  Are there better ways ? <br>\n<strong>Question</strong><br>\nDo you guys observe stable incremental gains when going from single fold -&gt; multiple folds ?<br>\nDo you guys observe stable incremental gains when going from single arch  -&gt; multiple arch?<br>\n**People belonging to large team ( &gt;= 3 persons), how do you guys combine different base models ? **</p>",
      "rawMarkdown": "\nJust wondering what are everyone's ensembling strategies.  Any insight beyond just concat different model's bboxes and then perform NMS ? \n\n**My observation**\n\nMy best single fold model with TTA achieves 0.26+ on LB.   With NMS-thres=0.5, this model predicts, on average, 49 bboxes per test image.  In single fold setting, TTA leads to incremental gain.\n\n\nEnsembling multiple folds the naive way  - concat all bboxes predictions and performing NMS - with the same inference setting as the single fold model lead to a score of 0.257.    The ensemble prediction produces an average of 77 bboxes per test image.\n\n\nMy theory for the worse LB score is that there are probably too many bboxes. As a result, the ensemble prediction boosts recall slightly but decreases precision a lot.\n\n\nThen, for the same ensemble weights, I use NMS-thres=0.4 and turn off TTA. This change significantly decreases the number of bboxes per test image to 46.  On submission, my LB score improves by 0.01. \n\n\n**Lesson Learnt** ( Not sure about its correctness ):\n\nIf Ensembling model the naive way ( concat bboxes + NMS),  # of bboxes per test image increases along with number of sub-models. \n\n\n**So we probably have to do some post-processing to ensure the # of boxes per test image doesn't grow with the number of sub-model and get out of control ? **\n\n\nRight now I am just tweaking inference config to bring down the bboxes per test image.  Are there better ways ? \n\n\n\n**Question**\n\n\nDo you guys observe stable incremental gains when going from single fold -> multiple folds ?\n\n\nDo you guys observe stable incremental gains when going from single arch  -> multiple arch?\n\n**People belonging to large team ( >= 3 persons), how do you guys combine different base models ? **",
      "votes": 15
    },
    {
      "id": 1249917,
      "postDate": "2021-03-23T16:13:24.203Z",
      "content": "<p>Ensembling with different folds and models definitely helps. For us, the boost is around 0.03LB. TTA seems to hurt because some diseases are concentrated on a single side? Does someone have any idea?</p>",
      "rawMarkdown": "Ensembling with different folds and models definitely helps. For us, the boost is around 0.03LB. TTA seems to hurt because some diseases are concentrated on a single side? Does someone have any idea?",
      "votes": 3,
      "replies": [
        {
          "id": 1249947,
          "postDate": "2021-03-23T16:40:29.937Z",
          "content": "<p>If you are saying TTA with Hflip may hurt, how about Hflip augmentation during training</p>",
          "rawMarkdown": "If you are saying TTA with Hflip may hurt, how about Hflip augmentation during training"
        },
        {
          "id": 1253210,
          "postDate": "2021-03-26T13:22:31.663Z",
          "content": "<p>Hi,how to ensemble different models?</p>",
          "rawMarkdown": "Hi,how to ensemble different models?"
        }
      ]
    },
    {
      "id": 1235704,
      "postDate": "2021-03-12T11:51:32.823Z",
      "content": "<p>i tryed and lb increase 0.002 , thanks for sharing !!!</p>\n<p>first question: yes, lb with single fold 0.248 ensemble 0.263<br>\nsecond question: I did not understand 😂</p>",
      "rawMarkdown": "i tryed and lb increase 0.002 , thanks for sharing !!!\n\nfirst question: yes, lb with single fold 0.248 ensemble 0.263\nsecond question: I did not understand 😂",
      "votes": 1,
      "replies": [
        {
          "id": 1236211,
          "postDate": "2021-03-12T22:11:50.343Z",
          "content": "<p><a href=\"https://www.kaggle.com/adrielcabral\" target=\"_blank\">@adrielcabral</a> is this 5 fold or less? </p>",
          "rawMarkdown": "@adrielcabral is this 5 fold or less? "
        },
        {
          "id": 1236216,
          "postDate": "2021-03-12T22:22:26.237Z",
          "content": "<p>only 5 folds</p>",
          "rawMarkdown": "only 5 folds"
        },
        {
          "id": 1236385,
          "postDate": "2021-03-13T05:00:15.093Z",
          "content": "<p>Better single model could be obtained with 5 folds or with less folds (for example 4)?</p>",
          "rawMarkdown": "Better single model could be obtained with 5 folds or with less folds (for example 4)?"
        },
        {
          "id": 1236834,
          "postDate": "2021-03-13T13:50:31.520Z",
          "content": "<p>with 5  folds … never tried less</p>",
          "rawMarkdown": "with 5  folds ... never tried less"
        },
        {
          "id": 1241168,
          "postDate": "2021-03-16T23:49:00.983Z",
          "content": "<p><a href=\"https://www.kaggle.com/adrielcabral\" target=\"_blank\">@adrielcabral</a> </p>\n<p>how did you ensemble your 5 folds ?  Did you just use the yolov5 default behavior ( concat predictions from all folds and then NMS) ? Any further post processing ? </p>\n<p>Did you observe that the average number of predicted bounding bboxes grow with the number of models ?  ( IE: 5 folds ensemble produces a lot more bboxes than 1 fold only)</p>",
          "rawMarkdown": "@adrielcabral \n\nhow did you ensemble your 5 folds ?  Did you just use the yolov5 default behavior ( concat predictions from all folds and then NMS) ? Any further post processing ? \n\nDid you observe that the average number of predicted bounding bboxes grow with the number of models ?  ( IE: 5 folds ensemble produces a lot more bboxes than 1 fold only)"
        },
        {
          "id": 1242071,
          "postDate": "2021-03-17T11:37:32.610Z",
          "content": "<p>i just do ensemble of weights !python detect.py --weights fold0.pt fold1.pt fold2.pt fold3.pt ..</p>\n<p>with ensemble, sometimes the model predict 100 boxes<br>\nedit: sorry, i said wrong, my model ensemble predict around of 60-70 boxes per image, </p>",
          "rawMarkdown": "i just do ensemble of weights !python detect.py --weights fold0.pt fold1.pt fold2.pt fold3.pt ..\n\nwith ensemble, sometimes the model predict 100 boxes\nedit: sorry, i said wrong, my model ensemble predict around of 60-70 boxes per image, ",
          "votes": 1
        },
        {
          "id": 1242280,
          "postDate": "2021-03-17T14:10:25.807Z",
          "content": "<p><a href=\"https://www.kaggle.com/adrielcabral\" target=\"_blank\">@adrielcabral</a> My ensemble is up to 300 boxes,  did you turn on multi-label, augment? and is your score threshold 0.001?</p>",
          "rawMarkdown": "@adrielcabral My ensemble is up to 300 boxes,  did you turn on multi-label, augment? and is your score threshold 0.001?"
        },
        {
          "id": 1242708,
          "postDate": "2021-03-17T18:48:17.747Z",
          "content": "<p>i turn on multi-label and conf_thresh: 0.001</p>",
          "rawMarkdown": "i turn on multi-label and conf_thresh: 0.001",
          "votes": 1
        },
        {
          "id": 1243186,
          "postDate": "2021-03-18T04:35:13.373Z",
          "content": "<p><a href=\"https://www.kaggle.com/adrielcabral\" target=\"_blank\">@adrielcabral</a> did you do NMS for overlapping training boxes before yolo training?</p>",
          "rawMarkdown": "@adrielcabral did you do NMS for overlapping training boxes before yolo training?"
        },
        {
          "id": 1243698,
          "postDate": "2021-03-18T12:30:02.553Z",
          "content": "<p>yes, fusion bboxes with IOU &gt; 0.4</p>",
          "rawMarkdown": "yes, fusion bboxes with IOU > 0.4",
          "votes": 1
        },
        {
          "id": 1243858,
          "postDate": "2021-03-18T14:33:21.643Z",
          "content": "<p>THanks <a href=\"https://www.kaggle.com/adrielcabral\" target=\"_blank\">@adrielcabral</a> I am following your way</p>",
          "rawMarkdown": "THanks @adrielcabral I am following your way"
        },
        {
          "id": 1243872,
          "postDate": "2021-03-18T14:47:47.720Z",
          "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> Well, not is the best way (42º place), but is good and I can still fall more in the end 😔😂</p>",
          "rawMarkdown": "@deepkim Well, not is the best way (42º place), but is good and I can still fall more in the end 😔😂",
          "votes": 1
        },
        {
          "id": 1244769,
          "postDate": "2021-03-19T08:05:14.163Z",
          "content": "<p><a href=\"https://www.kaggle.com/adrielcabral\" target=\"_blank\">@adrielcabral</a> hope you improve more</p>",
          "rawMarkdown": "@adrielcabral hope you improve more",
          "votes": 1
        },
        {
          "id": 1253222,
          "postDate": "2021-03-26T13:32:09.080Z",
          "content": "<p>HI,where is multi-label?</p>",
          "rawMarkdown": "HI,where is multi-label?",
          "votes": 1
        },
        {
          "id": 1253301,
          "postDate": "2021-03-26T15:19:16.383Z",
          "content": "<p>general.py in yolov5</p>",
          "rawMarkdown": "general.py in yolov5"
        }
      ]
    },
    {
      "id": 1235312,
      "postDate": "2021-03-12T03:37:50.250Z",
      "content": "<p>It is very difficult, looking for the advice of others. Once, I ensembled 6 models, average of 148 bboxes per image. Best model was 0.234, ensemble scored 0.252. Now, my best single models are 0.24x and 0.25x but ensemble doesn't increase score, or decreases…. I'm starting to think I got lucky. I'm so confused how we can ensemble for a better score.</p>",
      "rawMarkdown": "It is very difficult, looking for the advice of others. Once, I ensembled 6 models, average of 148 bboxes per image. Best model was 0.234, ensemble scored 0.252. Now, my best single models are 0.24x and 0.25x but ensemble doesn't increase score, or decreases.... I'm starting to think I got lucky. I'm so confused how we can ensemble for a better score.",
      "replies": [
        {
          "id": 1235350,
          "postDate": "2021-03-12T04:36:00.730Z",
          "content": "<p>How many bboxes do you get for best single model ? Is it a lot less ? </p>",
          "rawMarkdown": "How many bboxes do you get for best single model ? Is it a lot less ? "
        },
        {
          "id": 1235358,
          "postDate": "2021-03-12T04:44:11.327Z",
          "content": "<p>A lot less, ~50. I am almost 100% sure I just got crazy lucky with my ensemble, haven't gotten anywhere near it.</p>",
          "rawMarkdown": "A lot less, ~50. I am almost 100% sure I just got crazy lucky with my ensemble, haven't gotten anywhere near it.",
          "votes": 1
        },
        {
          "id": 1236386,
          "postDate": "2021-03-13T05:02:31.080Z",
          "content": "<p>Your best single model is obtained by yolov5, or other model?</p>",
          "rawMarkdown": "Your best single model is obtained by yolov5, or other model?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1249917,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "2021-03-23T16:13:24.203000",
      "content": "<p>Ensembling with different folds and models definitely helps. For us, the boost is around 0.03LB. TTA seems to hurt because some diseases are concentrated on a single side? Does someone have any idea?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1249947,
          "author_name": "Phat Tran",
          "author_url": "",
          "post_date": "2021-03-23T16:40:29.937000",
          "content": "<p>If you are saying TTA with Hflip may hurt, how about Hflip augmentation during training</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1253210,
          "author_name": "HAaHAa",
          "author_url": "",
          "post_date": "2021-03-26T13:22:31.663000",
          "content": "<p>Hi,how to ensemble different models?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1235704,
      "author_name": "adriel cabral",
      "author_url": "",
      "post_date": "2021-03-12T11:51:32.823000",
      "content": "<p>i tryed and lb increase 0.002 , thanks for sharing !!!</p>\n<p>first question: yes, lb with single fold 0.248 ensemble 0.263<br>\nsecond question: I did not understand 😂</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1236211,
          "author_name": "Trushant Kalyanpur",
          "author_url": "",
          "post_date": "2021-03-12T22:11:50.343000",
          "content": "<p><a href=\"https://www.kaggle.com/adrielcabral\" target=\"_blank\">@adrielcabral</a> is this 5 fold or less? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1236216,
          "author_name": "adriel cabral",
          "author_url": "",
          "post_date": "2021-03-12T22:22:26.237000",
          "content": "<p>only 5 folds</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1236385,
          "author_name": "Kuan Zhang",
          "author_url": "",
          "post_date": "2021-03-13T05:00:15.093000",
          "content": "<p>Better single model could be obtained with 5 folds or with less folds (for example 4)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1236834,
          "author_name": "adriel cabral",
          "author_url": "",
          "post_date": "2021-03-13T13:50:31.520000",
          "content": "<p>with 5  folds … never tried less</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1241168,
          "author_name": "NakedKoala",
          "author_url": "",
          "post_date": "2021-03-16T23:49:00.983000",
          "content": "<p><a href=\"https://www.kaggle.com/adrielcabral\" target=\"_blank\">@adrielcabral</a> </p>\n<p>how did you ensemble your 5 folds ?  Did you just use the yolov5 default behavior ( concat predictions from all folds and then NMS) ? Any further post processing ? </p>\n<p>Did you observe that the average number of predicted bounding bboxes grow with the number of models ?  ( IE: 5 folds ensemble produces a lot more bboxes than 1 fold only)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1242071,
          "author_name": "adriel cabral",
          "author_url": "",
          "post_date": "2021-03-17T11:37:32.610000",
          "content": "<p>i just do ensemble of weights !python detect.py --weights fold0.pt fold1.pt fold2.pt fold3.pt ..</p>\n<p>with ensemble, sometimes the model predict 100 boxes<br>\nedit: sorry, i said wrong, my model ensemble predict around of 60-70 boxes per image, </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1242280,
          "author_name": "Phat Tran",
          "author_url": "",
          "post_date": "2021-03-17T14:10:25.807000",
          "content": "<p><a href=\"https://www.kaggle.com/adrielcabral\" target=\"_blank\">@adrielcabral</a> My ensemble is up to 300 boxes,  did you turn on multi-label, augment? and is your score threshold 0.001?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1242708,
          "author_name": "adriel cabral",
          "author_url": "",
          "post_date": "2021-03-17T18:48:17.747000",
          "content": "<p>i turn on multi-label and conf_thresh: 0.001</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1243186,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2021-03-18T04:35:13.373000",
          "content": "<p><a href=\"https://www.kaggle.com/adrielcabral\" target=\"_blank\">@adrielcabral</a> did you do NMS for overlapping training boxes before yolo training?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1243698,
          "author_name": "adriel cabral",
          "author_url": "",
          "post_date": "2021-03-18T12:30:02.553000",
          "content": "<p>yes, fusion bboxes with IOU &gt; 0.4</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1243858,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2021-03-18T14:33:21.643000",
          "content": "<p>THanks <a href=\"https://www.kaggle.com/adrielcabral\" target=\"_blank\">@adrielcabral</a> I am following your way</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1243872,
          "author_name": "adriel cabral",
          "author_url": "",
          "post_date": "2021-03-18T14:47:47.720000",
          "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> Well, not is the best way (42º place), but is good and I can still fall more in the end 😔😂</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1244769,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2021-03-19T08:05:14.163000",
          "content": "<p><a href=\"https://www.kaggle.com/adrielcabral\" target=\"_blank\">@adrielcabral</a> hope you improve more</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1253222,
          "author_name": "HAaHAa",
          "author_url": "",
          "post_date": "2021-03-26T13:32:09.080000",
          "content": "<p>HI,where is multi-label?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1253301,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2021-03-26T15:19:16.383000",
          "content": "<p>general.py in yolov5</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1235312,
      "author_name": "Stanley Zheng",
      "author_url": "",
      "post_date": "2021-03-12T03:37:50.250000",
      "content": "<p>It is very difficult, looking for the advice of others. Once, I ensembled 6 models, average of 148 bboxes per image. Best model was 0.234, ensemble scored 0.252. Now, my best single models are 0.24x and 0.25x but ensemble doesn't increase score, or decreases…. I'm starting to think I got lucky. I'm so confused how we can ensemble for a better score.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1235350,
          "author_name": "NakedKoala",
          "author_url": "",
          "post_date": "2021-03-12T04:36:00.730000",
          "content": "<p>How many bboxes do you get for best single model ? Is it a lot less ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1235358,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2021-03-12T04:44:11.327000",
          "content": "<p>A lot less, ~50. I am almost 100% sure I just got crazy lucky with my ensemble, haven't gotten anywhere near it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1236386,
          "author_name": "Kuan Zhang",
          "author_url": "",
          "post_date": "2021-03-13T05:02:31.080000",
          "content": "<p>Your best single model is obtained by yolov5, or other model?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1235297": "\nJust wondering what are everyone's ensembling strategies.  Any insight beyond just concat different model's bboxes and then perform NMS ? \n\n**My observation**\n\nMy best single fold model with TTA achieves 0.26+ on LB.   With NMS-thres=0.5, this model predicts, on average, 49 bboxes per test image.  In single fold setting, TTA leads to incremental gain.\n\n\nEnsembling multiple folds the naive way  - concat all bboxes predictions and performing NMS - with the same inference setting as the single fold model lead to a score of 0.257.    The ensemble prediction produces an average of 77 bboxes per test image.\n\n\nMy theory for the worse LB score is that there are probably too many bboxes. As a result, the ensemble prediction boosts recall slightly but decreases precision a lot.\n\n\nThen, for the same ensemble weights, I use NMS-thres=0.4 and turn off TTA. This change significantly decreases the number of bboxes per test image to 46.  On submission, my LB score improves by 0.01. \n\n\n**Lesson Learnt** ( Not sure about its correctness ):\n\nIf Ensembling model the naive way ( concat bboxes + NMS),  # of bboxes per test image increases along with number of sub-models. \n\n\n**So we probably have to do some post-processing to ensure the # of boxes per test image doesn't grow with the number of sub-model and get out of control ? **\n\n\nRight now I am just tweaking inference config to bring down the bboxes per test image.  Are there better ways ? \n\n\n\n**Question**\n\n\nDo you guys observe stable incremental gains when going from single fold -> multiple folds ?\n\n\nDo you guys observe stable incremental gains when going from single arch  -> multiple arch?\n\n**People belonging to large team ( >= 3 persons), how do you guys combine different base models ? **",
    "1249917": "Ensembling with different folds and models definitely helps. For us, the boost is around 0.03LB. TTA seems to hurt because some diseases are concentrated on a single side? Does someone have any idea?",
    "1235704": "i tryed and lb increase 0.002 , thanks for sharing !!!\n\nfirst question: yes, lb with single fold 0.248 ensemble 0.263\nsecond question: I did not understand 😂",
    "1235312": "It is very difficult, looking for the advice of others. Once, I ensembled 6 models, average of 148 bboxes per image. Best model was 0.234, ensemble scored 0.252. Now, my best single models are 0.24x and 0.25x but ensemble doesn't increase score, or decreases.... I'm starting to think I got lucky. I'm so confused how we can ensemble for a better score."
  }
}