{
  "id": 229786,
  "title": "[4th place solution] YOLOv5x",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/229786",
  "author_name": "fantastic_hirarin",
  "post_date": "2021-03-31T18:25:07.697000",
  "votes": 34,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Thank you to all participants and organizers!<br>\nAnd congratulations to all the winners!<br>\nHere's the review of my solution.</p>\n<h1>Summary</h1>\n<ul>\n<li>model: yolov5x</li>\n<li>image size: 640</li>\n<li>TTA: 3 scale patterns and horizontal flip</li>\n<li>ensemble: (4fold cv * 3 different preprocessed labels ) = 12 models</li>\n</ul>\n<h1>Preprocessing</h1>\n<p>In the preprocessing, I used wbf and merged the boxes with iou=0.4.<br>\nThis makes a single box created by multiple annotators.<br>\nOne thing to keep in mind is that how many of the three annotators judged the disease. <br>\nThere is a difference in how much confidence one can have in a disease box if only one person or two people or three people judges it to be a disease.<br>\nI thought it is necessary to create a model that could account for the difference.</p>\n<p>First, I created three different training labels using the following methods.</p>\n<ul>\n<li>labels-1: Leave boxes that are considered a disease by one or more annotators.</li>\n<li>labels-2: Leave boxes that are considered a disease by two or more annotators.</li>\n<li>labels-3: Leave boxes that are considered a disease by all three annotators.</li>\n</ul>\n<h1>Training</h1>\n<p>The above three patterns of labels were trained with yolov5x.<br>\nI trained 50 epochs with image size 640, and the only hyperparameter change was to change the mixup from 0 to 0.5.<br>\nHere are the four models I created.</p>\n<ul>\n<li>model-1: trained labels-1 on all images</li>\n<li>model-2: trained labels-2 on all images</li>\n<li>model-3: trained labels-3 on all images</li>\n<li>model-4: trained labels-1 only on images with disease</li>\n</ul>\n<h1>Inference</h1>\n<p>I modified the TTA implemented in yolov5 a bit.<br>\nTTA was performed on six image patterns using a combination of three scale patterns [1, 0.83, 0.67] and horizontal flip.</p>\n<p>The final output is an ensemble of the model's outputs using wbf with iou=0.5.<br>\nHere are the scores for each model in the ensemble.</p>\n<ul>\n<li>model-1: public 0.264, private 0.278</li>\n<li>model-1 + model-2: public 0.271, private 0.295</li>\n<li>model-1 + model-2 + model-3: public 0.267, private 0.303</li>\n<li>model-1 + model-2 + model-4: public 0.278, private 0.297</li>\n<li>model-1 + model-2 + model-3 + model-4: public 0.276, private 0.306</li>\n</ul>\n<p><strong>Why the ensemble improved the score:</strong><br>\nFor the ensemble using wbf, the box's confidence score detected by more models will be higher. The box's confidence score detected by fewer models will be lower.<br>\nThe boxes that are judged to be a disease by more annotators are trained to be detected by more models (model-1, model-2, model-3). Therefore, the more likely a box is to be diseased, the higher its confidence score will be. As a result, the mAP is improved.</p>\n<h1>Class 14, no finding</h1>\n<p>I considered the highest confidence score among the boxes detected by model-1 and model-2 as the <em>finding_score</em> of the image.<br>\nThen, I appended f'14 {1-<em>finding_score</em>} 0 0 1 1' to the results of all images.<br>\nAs a result, the class 14 score will be higher for images where the box is not easily detected by the detection models.<br>\nSince mAP only considers the order based on the confidence score, it is safe to add f'14 {1-finding_score} 0 0 1 1' to all images.</p>\n<h1>Final selection</h1>\n<p>I chose the following two outputs for my final selection</p>\n<ul>\n<li>model-1 + model-2 + model-4: public 0.278, private 0.297</li>\n<li>model-1 + model-2 + model-4 + some public notebook's output: public 0.285, private 0.303</li>\n</ul>\n<p>I couldn't choose the output with the model-3 ensemble.<br>\nThis is because model-3 had a low mAP due to the extremely small sample size of minor diseases, and the public LB dropped when model-3 was ensembled.<br>\nSo sad :(</p>",
  "messages": [
    {
      "id": 1258585,
      "postDate": "2021-03-31T18:25:07.697Z",
      "content": "<p>Thank you to all participants and organizers!<br>\nAnd congratulations to all the winners!<br>\nHere's the review of my solution.</p>\n<h1>Summary</h1>\n<ul>\n<li>model: yolov5x</li>\n<li>image size: 640</li>\n<li>TTA: 3 scale patterns and horizontal flip</li>\n<li>ensemble: (4fold cv * 3 different preprocessed labels ) = 12 models</li>\n</ul>\n<h1>Preprocessing</h1>\n<p>In the preprocessing, I used wbf and merged the boxes with iou=0.4.<br>\nThis makes a single box created by multiple annotators.<br>\nOne thing to keep in mind is that how many of the three annotators judged the disease. <br>\nThere is a difference in how much confidence one can have in a disease box if only one person or two people or three people judges it to be a disease.<br>\nI thought it is necessary to create a model that could account for the difference.</p>\n<p>First, I created three different training labels using the following methods.</p>\n<ul>\n<li>labels-1: Leave boxes that are considered a disease by one or more annotators.</li>\n<li>labels-2: Leave boxes that are considered a disease by two or more annotators.</li>\n<li>labels-3: Leave boxes that are considered a disease by all three annotators.</li>\n</ul>\n<h1>Training</h1>\n<p>The above three patterns of labels were trained with yolov5x.<br>\nI trained 50 epochs with image size 640, and the only hyperparameter change was to change the mixup from 0 to 0.5.<br>\nHere are the four models I created.</p>\n<ul>\n<li>model-1: trained labels-1 on all images</li>\n<li>model-2: trained labels-2 on all images</li>\n<li>model-3: trained labels-3 on all images</li>\n<li>model-4: trained labels-1 only on images with disease</li>\n</ul>\n<h1>Inference</h1>\n<p>I modified the TTA implemented in yolov5 a bit.<br>\nTTA was performed on six image patterns using a combination of three scale patterns [1, 0.83, 0.67] and horizontal flip.</p>\n<p>The final output is an ensemble of the model's outputs using wbf with iou=0.5.<br>\nHere are the scores for each model in the ensemble.</p>\n<ul>\n<li>model-1: public 0.264, private 0.278</li>\n<li>model-1 + model-2: public 0.271, private 0.295</li>\n<li>model-1 + model-2 + model-3: public 0.267, private 0.303</li>\n<li>model-1 + model-2 + model-4: public 0.278, private 0.297</li>\n<li>model-1 + model-2 + model-3 + model-4: public 0.276, private 0.306</li>\n</ul>\n<p><strong>Why the ensemble improved the score:</strong><br>\nFor the ensemble using wbf, the box's confidence score detected by more models will be higher. The box's confidence score detected by fewer models will be lower.<br>\nThe boxes that are judged to be a disease by more annotators are trained to be detected by more models (model-1, model-2, model-3). Therefore, the more likely a box is to be diseased, the higher its confidence score will be. As a result, the mAP is improved.</p>\n<h1>Class 14, no finding</h1>\n<p>I considered the highest confidence score among the boxes detected by model-1 and model-2 as the <em>finding_score</em> of the image.<br>\nThen, I appended f'14 {1-<em>finding_score</em>} 0 0 1 1' to the results of all images.<br>\nAs a result, the class 14 score will be higher for images where the box is not easily detected by the detection models.<br>\nSince mAP only considers the order based on the confidence score, it is safe to add f'14 {1-finding_score} 0 0 1 1' to all images.</p>\n<h1>Final selection</h1>\n<p>I chose the following two outputs for my final selection</p>\n<ul>\n<li>model-1 + model-2 + model-4: public 0.278, private 0.297</li>\n<li>model-1 + model-2 + model-4 + some public notebook's output: public 0.285, private 0.303</li>\n</ul>\n<p>I couldn't choose the output with the model-3 ensemble.<br>\nThis is because model-3 had a low mAP due to the extremely small sample size of minor diseases, and the public LB dropped when model-3 was ensembled.<br>\nSo sad :(</p>",
      "rawMarkdown": "Thank you to all participants and organizers!\nAnd congratulations to all the winners!\nHere's the review of my solution.\n\n# Summary\n- model: yolov5x\n- image size: 640\n- TTA: 3 scale patterns and horizontal flip\n- ensemble: (4fold cv * 3 different preprocessed labels ) = 12 models\n\n\n# Preprocessing\nIn the preprocessing, I used wbf and merged the boxes with iou=0.4.\nThis makes a single box created by multiple annotators.\nOne thing to keep in mind is that how many of the three annotators judged the disease. \nThere is a difference in how much confidence one can have in a disease box if only one person or two people or three people judges it to be a disease.\nI thought it is necessary to create a model that could account for the difference.\n\nFirst, I created three different training labels using the following methods.\n- labels-1: Leave boxes that are considered a disease by one or more annotators.\n- labels-2: Leave boxes that are considered a disease by two or more annotators.\n- labels-3: Leave boxes that are considered a disease by all three annotators.\n\n\n# Training\nThe above three patterns of labels were trained with yolov5x.\nI trained 50 epochs with image size 640, and the only hyperparameter change was to change the mixup from 0 to 0.5.\nHere are the four models I created.\n- model-1: trained labels-1 on all images\n- model-2: trained labels-2 on all images\n- model-3: trained labels-3 on all images\n- model-4: trained labels-1 only on images with disease\n\n\n# Inference\nI modified the TTA implemented in yolov5 a bit.\nTTA was performed on six image patterns using a combination of three scale patterns [1, 0.83, 0.67] and horizontal flip.\n\nThe final output is an ensemble of the model's outputs using wbf with iou=0.5.\nHere are the scores for each model in the ensemble.\n- model-1: public 0.264, private 0.278\n- model-1 + model-2: public 0.271, private 0.295\n- model-1 + model-2 + model-3: public 0.267, private 0.303\n- model-1 + model-2 + model-4: public 0.278, private 0.297\n- model-1 + model-2 + model-3 + model-4: public 0.276, private 0.306\n\n**Why the ensemble improved the score:**\nFor the ensemble using wbf, the box's confidence score detected by more models will be higher. The box's confidence score detected by fewer models will be lower.\nThe boxes that are judged to be a disease by more annotators are trained to be detected by more models (model-1, model-2, model-3). Therefore, the more likely a box is to be diseased, the higher its confidence score will be. As a result, the mAP is improved.\n\n# Class 14, no finding\nI considered the highest confidence score among the boxes detected by model-1 and model-2 as the *finding_score* of the image.\nThen, I appended f'14 {1-*finding_score*} 0 0 1 1' to the results of all images.\nAs a result, the class 14 score will be higher for images where the box is not easily detected by the detection models.\nSince mAP only considers the order based on the confidence score, it is safe to add f'14 {1-finding_score} 0 0 1 1' to all images.\n\n# Final selection\nI chose the following two outputs for my final selection\n- model-1 + model-2 + model-4: public 0.278, private 0.297\n- model-1 + model-2 + model-4 + some public notebook's output: public 0.285, private 0.303\n\nI couldn't choose the output with the model-3 ensemble.\nThis is because model-3 had a low mAP due to the extremely small sample size of minor diseases, and the public LB dropped when model-3 was ensembled.\nSo sad :(",
      "votes": 34
    },
    {
      "id": 1259461,
      "postDate": "2021-04-01T12:31:03.997Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/hiraiitsuki\" target=\"_blank\">@hiraiitsuki</a>  ! Yet another Yolo win ;-)</p>",
      "rawMarkdown": "Congrats @hiraiitsuki  ! Yet another Yolo win ;-)",
      "votes": 1,
      "replies": [
        {
          "id": 1259652,
          "postDate": "2021-04-01T15:04:52.100Z",
          "content": "<p>Thank you!!</p>",
          "rawMarkdown": "Thank you!!"
        }
      ]
    },
    {
      "id": 1258891,
      "postDate": "2021-04-01T01:37:30.690Z",
      "content": "<p>Congrats on solo 4th place and gold medal <a href=\"https://www.kaggle.com/hiraiitsuki\" target=\"_blank\">@hiraiitsuki</a> </p>",
      "rawMarkdown": "Congrats on solo 4th place and gold medal @hiraiitsuki ",
      "votes": 1,
      "replies": [
        {
          "id": 1259181,
          "postDate": "2021-04-01T07:49:34.833Z",
          "content": "<p>Thank you!<br>\nCongrats on your silver medal too!</p>",
          "rawMarkdown": "Thank you!\nCongrats on your silver medal too!"
        }
      ]
    },
    {
      "id": 1258714,
      "postDate": "2021-03-31T20:28:19.940Z",
      "content": "<p>It is a nice and simple solution! And Yolov5 is cool!!! 💪</p>",
      "rawMarkdown": "It is a nice and simple solution! And Yolov5 is cool!!! 💪",
      "votes": 1,
      "replies": [
        {
          "id": 1259177,
          "postDate": "2021-04-01T07:43:56.213Z",
          "content": "<p>Thank you!!<br>\nYes! Yolov5 is fantastic :)</p>",
          "rawMarkdown": "Thank you!!\nYes! Yolov5 is fantastic :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1258631,
      "postDate": "2021-03-31T19:02:30.947Z",
      "content": "<p>Congrats for the great solution using different confidence levels of labels. I also used a similar solution with two levels (weak=&gt;1 annotator vs strong=&gt;2 or 3 annotators) of labels but I just trained 2 models from only one fold division which wasn`t sufficient to perform. For your final ensemble, did you modify wbf ? Because I thought wbf was simply averaging the confidence bbox scores when fused. </p>",
      "rawMarkdown": "Congrats for the great solution using different confidence levels of labels. I also used a similar solution with two levels (weak=>1 annotator vs strong=>2 or 3 annotators) of labels but I just trained 2 models from only one fold division which wasn`t sufficient to perform. For your final ensemble, did you modify wbf ? Because I thought wbf was simply averaging the confidence bbox scores when fused. ",
      "votes": 1,
      "replies": [
        {
          "id": 1259155,
          "postDate": "2021-04-01T07:21:54.567Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1259159,
          "postDate": "2021-04-01T07:24:38.390Z",
          "content": "<p>Thank you!<br>\nI didn't modified the wbf source code.</p>\n<blockquote>\n  <p>Because I thought wbf was simply averaging the confidence bbox scores when fused.</p>\n</blockquote>\n<p>That's right.<br>\nHowever, when the detected and undetected bboxes are merged, the confidence score will be lower.<br>\nFor example, a bbox score detected by only one of the three models will be reduced to a third when the outputs of the three models are merged.<br>\nAs a result, the score of bboxes detected by fewer models will be lower, and the score of bboxes detected by many models will be relatively higher.</p>",
          "rawMarkdown": "Thank you!\nI didn't modified the wbf source code.\n\n> Because I thought wbf was simply averaging the confidence bbox scores when fused.\n\nThat's right.\nHowever, when the detected and undetected bboxes are merged, the confidence score will be lower.\nFor example, a bbox score detected by only one of the three models will be reduced to a third when the outputs of the three models are merged.\nAs a result, the score of bboxes detected by fewer models will be lower, and the score of bboxes detected by many models will be relatively higher."
        }
      ]
    },
    {
      "id": 1258806,
      "postDate": "2021-03-31T22:44:36.843Z",
      "content": "<p>Congratulations. Great job achieving such a great score with only YoloV5. I like how you built diversion versions.</p>",
      "rawMarkdown": "Congratulations. Great job achieving such a great score with only YoloV5. I like how you built diversion versions.",
      "votes": 2,
      "replies": [
        {
          "id": 1259180,
          "postDate": "2021-04-01T07:46:37.723Z",
          "content": "<p>Thank you!<br>\nI think the fun part of Kaggle is that simple solution can get a high score depending on ideas!</p>",
          "rawMarkdown": "Thank you!\nI think the fun part of Kaggle is that simple solution can get a high score depending on ideas!",
          "votes": 2
        }
      ]
    },
    {
      "id": 2697513,
      "postDate": "2024-03-14T23:58:48.910Z",
      "content": "<p>How did you keep track of which radiologists annoted which bbox after using wbf in the preprocessing?</p>",
      "rawMarkdown": "How did you keep track of which radiologists annoted which bbox after using wbf in the preprocessing?"
    }
  ],
  "comments": [
    {
      "id": 1259461,
      "author_name": "Sidney Ng",
      "author_url": "",
      "post_date": "2021-04-01T12:31:03.997000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/hiraiitsuki\" target=\"_blank\">@hiraiitsuki</a>  ! Yet another Yolo win ;-)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1259652,
          "author_name": "fantastic_hirarin",
          "author_url": "",
          "post_date": "2021-04-01T15:04:52.100000",
          "content": "<p>Thank you!!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1258891,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-04-01T01:37:30.690000",
      "content": "<p>Congrats on solo 4th place and gold medal <a href=\"https://www.kaggle.com/hiraiitsuki\" target=\"_blank\">@hiraiitsuki</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1259181,
          "author_name": "fantastic_hirarin",
          "author_url": "",
          "post_date": "2021-04-01T07:49:34.833000",
          "content": "<p>Thank you!<br>\nCongrats on your silver medal too!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1258714,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2021-03-31T20:28:19.940000",
      "content": "<p>It is a nice and simple solution! And Yolov5 is cool!!! 💪</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1259177,
          "author_name": "fantastic_hirarin",
          "author_url": "",
          "post_date": "2021-04-01T07:43:56.213000",
          "content": "<p>Thank you!!<br>\nYes! Yolov5 is fantastic :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1258631,
      "author_name": "Alexandre Cadrin-Chênevert",
      "author_url": "",
      "post_date": "2021-03-31T19:02:30.947000",
      "content": "<p>Congrats for the great solution using different confidence levels of labels. I also used a similar solution with two levels (weak=&gt;1 annotator vs strong=&gt;2 or 3 annotators) of labels but I just trained 2 models from only one fold division which wasn`t sufficient to perform. For your final ensemble, did you modify wbf ? Because I thought wbf was simply averaging the confidence bbox scores when fused. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1259155,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-01T07:21:54.567000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1259159,
          "author_name": "fantastic_hirarin",
          "author_url": "",
          "post_date": "2021-04-01T07:24:38.390000",
          "content": "<p>Thank you!<br>\nI didn't modified the wbf source code.</p>\n<blockquote>\n  <p>Because I thought wbf was simply averaging the confidence bbox scores when fused.</p>\n</blockquote>\n<p>That's right.<br>\nHowever, when the detected and undetected bboxes are merged, the confidence score will be lower.<br>\nFor example, a bbox score detected by only one of the three models will be reduced to a third when the outputs of the three models are merged.<br>\nAs a result, the score of bboxes detected by fewer models will be lower, and the score of bboxes detected by many models will be relatively higher.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1258806,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-03-31T22:44:36.843000",
      "content": "<p>Congratulations. Great job achieving such a great score with only YoloV5. I like how you built diversion versions.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1259180,
          "author_name": "fantastic_hirarin",
          "author_url": "",
          "post_date": "2021-04-01T07:46:37.723000",
          "content": "<p>Thank you!<br>\nI think the fun part of Kaggle is that simple solution can get a high score depending on ideas!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2697513,
      "author_name": "Gabriel Maire",
      "author_url": "",
      "post_date": "2024-03-14T23:58:48.910000",
      "content": "<p>How did you keep track of which radiologists annoted which bbox after using wbf in the preprocessing?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1258585": "Thank you to all participants and organizers!\nAnd congratulations to all the winners!\nHere's the review of my solution.\n\n# Summary\n- model: yolov5x\n- image size: 640\n- TTA: 3 scale patterns and horizontal flip\n- ensemble: (4fold cv * 3 different preprocessed labels ) = 12 models\n\n\n# Preprocessing\nIn the preprocessing, I used wbf and merged the boxes with iou=0.4.\nThis makes a single box created by multiple annotators.\nOne thing to keep in mind is that how many of the three annotators judged the disease. \nThere is a difference in how much confidence one can have in a disease box if only one person or two people or three people judges it to be a disease.\nI thought it is necessary to create a model that could account for the difference.\n\nFirst, I created three different training labels using the following methods.\n- labels-1: Leave boxes that are considered a disease by one or more annotators.\n- labels-2: Leave boxes that are considered a disease by two or more annotators.\n- labels-3: Leave boxes that are considered a disease by all three annotators.\n\n\n# Training\nThe above three patterns of labels were trained with yolov5x.\nI trained 50 epochs with image size 640, and the only hyperparameter change was to change the mixup from 0 to 0.5.\nHere are the four models I created.\n- model-1: trained labels-1 on all images\n- model-2: trained labels-2 on all images\n- model-3: trained labels-3 on all images\n- model-4: trained labels-1 only on images with disease\n\n\n# Inference\nI modified the TTA implemented in yolov5 a bit.\nTTA was performed on six image patterns using a combination of three scale patterns [1, 0.83, 0.67] and horizontal flip.\n\nThe final output is an ensemble of the model's outputs using wbf with iou=0.5.\nHere are the scores for each model in the ensemble.\n- model-1: public 0.264, private 0.278\n- model-1 + model-2: public 0.271, private 0.295\n- model-1 + model-2 + model-3: public 0.267, private 0.303\n- model-1 + model-2 + model-4: public 0.278, private 0.297\n- model-1 + model-2 + model-3 + model-4: public 0.276, private 0.306\n\n**Why the ensemble improved the score:**\nFor the ensemble using wbf, the box's confidence score detected by more models will be higher. The box's confidence score detected by fewer models will be lower.\nThe boxes that are judged to be a disease by more annotators are trained to be detected by more models (model-1, model-2, model-3). Therefore, the more likely a box is to be diseased, the higher its confidence score will be. As a result, the mAP is improved.\n\n# Class 14, no finding\nI considered the highest confidence score among the boxes detected by model-1 and model-2 as the *finding_score* of the image.\nThen, I appended f'14 {1-*finding_score*} 0 0 1 1' to the results of all images.\nAs a result, the class 14 score will be higher for images where the box is not easily detected by the detection models.\nSince mAP only considers the order based on the confidence score, it is safe to add f'14 {1-finding_score} 0 0 1 1' to all images.\n\n# Final selection\nI chose the following two outputs for my final selection\n- model-1 + model-2 + model-4: public 0.278, private 0.297\n- model-1 + model-2 + model-4 + some public notebook's output: public 0.285, private 0.303\n\nI couldn't choose the output with the model-3 ensemble.\nThis is because model-3 had a low mAP due to the extremely small sample size of minor diseases, and the public LB dropped when model-3 was ensembled.\nSo sad :(",
    "1259461": "Congrats @hiraiitsuki  ! Yet another Yolo win ;-)",
    "1258891": "Congrats on solo 4th place and gold medal @hiraiitsuki ",
    "1258714": "It is a nice and simple solution! And Yolov5 is cool!!! 💪",
    "1258631": "Congrats for the great solution using different confidence levels of labels. I also used a similar solution with two levels (weak=>1 annotator vs strong=>2 or 3 annotators) of labels but I just trained 2 models from only one fold division which wasn`t sufficient to perform. For your final ensemble, did you modify wbf ? Because I thought wbf was simply averaging the confidence bbox scores when fused. ",
    "1258806": "Congratulations. Great job achieving such a great score with only YoloV5. I like how you built diversion versions.",
    "2697513": "How did you keep track of which radiologists annoted which bbox after using wbf in the preprocessing?"
  }
}