{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":24800,"databundleVersionId":1831594,"sourceType":"competition"},{"sourceId":1950595,"sourceType":"datasetVersion","datasetId":1164135}],"dockerImageVersionId":30822,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<span style=\"color: #ae1400; Segoe UI; font-size: 2.0em; font-weight: 300;\">Overview</span>\n\n<p style='text-align: justify;'><span style=\"font-family: Trebuchet MS; font-size: 1.2em;\"> This notebook explores the different ways to select or fuse multiple bboxes of the same chest abnormality, annotated by several radiologists</span></p>\n\n<p style='text-align: justify;'><span style=\"font-family: Trebuchet MS; font-size: 1.2em;\"> It also covers the conversion of custom bbox annotations to the YoLo format for frameworks such as YOLO etc</span></p>\n\n\n**A few key intricacies in this competetion includes the way the training data is provided. To state a few:** \n\n- **The abnormalities are labelled by multiple radiologists and there seems to be multiple bounding boxes for some abnormalities.**\n\n- **Another issue being that some dense abnormality/lesion area may contain multiple labels. The radiologists creates boxes and then they may assign many labels to a single bounding box. This was stated by one of the competetion hosts.**\n\nSo the challenge of this competetion includes handling these issues before or after model training.Some ways to handle this would be to use suppression, selection or fusion techniques below which are covered in this notebook:\n\n- **Non-maximum Suppression (NMS)**\n- **Soft-NMS**\n- **Non-maximum Weighted (NMW)**\n- **Weighted Bboxes Fusion (WBF)**\n\n\n\nThese are generally used after model scoring to get to a consensus of a bounding box where the abnormality is based on the confidence scores, weightage of different models (if ensembling is used) etc.\n\nHere, the challenge lies in the fact that we don't have confidence scores or metrics to assign weightage to the multiple annotations by radiologists. So all radiologists are treated equally and the suppression or fusion of bounding boxes has to be done with these factors out of the picture. This seems to show a behaviour by which the bounding boxes which appear alone, seems to get suppressed. This issue is handled in this notebook by separating out single bounding boxes before any technique is applied.\n\n**Finally, please feel free to suggest any other novel methods to address these issues. Hope everyone finds this useful!**\n\n\n\n<br />\n<br />\n\n\n<p style='text-align: justify;'><span style=\"color: #ae1400; Segoe UI; font-size: 1.2em; font-weight: 300;\">Check out the training notebook which uses the Yolo dataset generated here. It explains the Installation, Data preparation, Training and Inference using the Yolov8 model l. It can be adapted to various other models with just 1-2 lines of code change.</span></p>\n\n\n<p style='text-align: justify;'><span style=\"font-family: Trebuchet MS; font-size: 1.1em;\">vincxr-yolov8l⚡📈</span></p>\n\n\nTRAINING NOTEBOOK - https://www.kaggle.com/code/buithanhxuan/vincxr-yolov8l\n\n<br />\n<br />\n\n<p style='text-align: justify;'><span style=\"color: #ae1400; Segoe UI; font-size: 1.2em; font-weight: 300;\">The Yolo dataset with Fused Boxes generated as a part of this notebook is public. Please do check it out.</span></p>\n\n\n<p style='text-align: justify;'><span style=\"font-family: Trebuchet MS; font-size: 1.1em;\">VinBigData - Yolo Dataset with WBF 3x Downscaled</span></p>\n\n\nDATASET LINK - https://www.kaggle.com/datasets/buithanhxuan/vinbigdata-yolo-dataset-with-wbf-3x-downscaled\n\n**Please note that All images with Chest Abnormalities are present in this dataset.**\n\n<br />\n<br />\n\n**Class Mapping in the Annotations File:**\n\n    0 - Aortic enlargement\n    1 - Atelectasis\n    2 - Calcification\n    3 - Cardiomegaly\n    4 - Consolidation\n    5 - ILD\n    6 - Infiltration\n    7 - Lung Opacity\n    8 - Nodule/Mass\n    9 - Other lesion\n    10 - Pleural effusion\n    11 - Pleural thickening\n    12 - Pneumothorax\n    13 - Pulmonary fibrosis\n    14 - No finding\n \n \n### Citations:\n\n**Thanks to [corochann](http://https://www.kaggle.com/corochann) for creating the vinbigdata-chest-xray-original-png which is used here:\nhttps://www.kaggle.com/datasets/corochann/vinbigdata-chest-xray-original-png**\n\n**[Weighted Boxes Fusion: ensembling boxes for object detection models paper](https://arxiv.org/abs/1910.13302)**\n\n\n<br />\n<br />\n\n[![Ask Me Anything !](https://img.shields.io/badge/Ask%20me-something-1abc9c.svg?style=flat-square&logo=kaggle)](https://www.kaggle.com/buithanhxuan)\n<br />\n\n![Upvote!](https://img.shields.io/badge/Upvote-If%20you%20like%20my%20work-07b3c8?style=for-the-badge&logo=kaggle)","metadata":{}},{"cell_type":"markdown","source":"![](https://www.futuretimeline.net/blog/images/1466-chest-xray-ai-technology.jpg)\n\n<p style='text-align: center;'><span style=\"color: #0D0D0D; font-family: Segoe UI; font-size: 2.6em; font-weight: 300;\">VINBIGDATA - FUSING BBOXES + BUILDING YOLO DATASET 640px</span></p>\n\n\n","metadata":{}},{"cell_type":"code","source":"!pip install -q ensemble-boxes","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-12-30T02:01:25.219917Z","iopub.execute_input":"2024-12-30T02:01:25.220277Z","iopub.status.idle":"2024-12-30T02:01:31.689257Z","shell.execute_reply.started":"2024-12-30T02:01:25.220244Z","shell.execute_reply":"2024-12-30T02:01:31.687931Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%matplotlib inline\n\nimport os\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom matplotlib import rcParams\nsns.set(rc={\"font.size\":9,\"axes.titlesize\":15,\"axes.labelsize\":9,\n            \"axes.titlepad\":11, \"axes.labelpad\":9, \"legend.fontsize\":7,\n            \"legend.title_fontsize\":7, 'axes.grid' : False})\nimport cv2\nimport json\nimport pandas as pd\nimport glob\nimport os.path as osp\nfrom path import Path\nimport datetime\nimport numpy as np\nfrom tqdm.auto import tqdm\nimport random\nimport shutil\nfrom sklearn.model_selection import train_test_split\n\nfrom ensemble_boxes import *\nimport warnings\nfrom collections import Counter","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:01:31.6905Z","iopub.execute_input":"2024-12-30T02:01:31.690913Z","iopub.status.idle":"2024-12-30T02:01:34.56052Z","shell.execute_reply.started":"2024-12-30T02:01:31.690878Z","shell.execute_reply":"2024-12-30T02:01:34.55934Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Loading the Annotation CSV","metadata":{}},{"cell_type":"code","source":"train_annotations = pd.read_csv(\"/kaggle/input/vinbigdata-chest-xray-abnormalities-detection/train.csv\")\ntrain_annotations.head(5)","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:01:34.56177Z","iopub.execute_input":"2024-12-30T02:01:34.562316Z","iopub.status.idle":"2024-12-30T02:01:34.76025Z","shell.execute_reply.started":"2024-12-30T02:01:34.562283Z","shell.execute_reply":"2024-12-30T02:01:34.759075Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Selecting Images with Abnormalities","metadata":{}},{"cell_type":"code","source":"train_annotations = train_annotations[train_annotations.class_id!=14]\ntrain_annotations['image_path'] = train_annotations['image_id'].map(lambda x:os.path.join('/kaggle/input/vinbigdata-chest-xray-original-png/train', str(x)+'.png'))\ntrain_annotations.head(5)","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:01:34.761372Z","iopub.execute_input":"2024-12-30T02:01:34.761638Z","iopub.status.idle":"2024-12-30T02:01:34.838914Z","shell.execute_reply.started":"2024-12-30T02:01:34.761617Z","shell.execute_reply":"2024-12-30T02:01:34.837867Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"img_size_df = pd.read_csv('/kaggle/input/vinbigdata-chest-xray-original-png/train_meta.csv', names=['image_id', 'orig_height', 'orig_width'])\n\n# Merge the original image size information with the training data\ntrain_annotations = train_annotations.merge(img_size_df, on='image_id', how='left')\ntrain_annotations.head(5)","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:01:34.839973Z","iopub.execute_input":"2024-12-30T02:01:34.840353Z","iopub.status.idle":"2024-12-30T02:01:34.930157Z","shell.execute_reply.started":"2024-12-30T02:01:34.840326Z","shell.execute_reply":"2024-12-30T02:01:34.929081Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"imagepaths = train_annotations['image_path'].unique()\nprint(\"Number of Images with abnormalities:\",len(imagepaths))\nanno_count = train_annotations.shape[0]\nprint(\"Number of Annotations with abnormalities:\", anno_count)","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:01:34.933086Z","iopub.execute_input":"2024-12-30T02:01:34.933424Z","iopub.status.idle":"2024-12-30T02:01:34.952862Z","shell.execute_reply.started":"2024-12-30T02:01:34.933396Z","shell.execute_reply":"2024-12-30T02:01:34.95172Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Helper Functions","metadata":{}},{"cell_type":"code","source":"def plot_img(img, size=(18, 18), is_rgb=True, title=\"\", cmap='gray'):\n    plt.figure(figsize=size)\n    plt.imshow(img, cmap=cmap)\n    plt.suptitle(title)\n    plt.show()\n\ndef plot_imgs(imgs, cols=2, size=10, is_rgb=True, title=\"\", cmap='gray', img_size=None):\n    rows = len(imgs)//cols + 1\n    fig = plt.figure(figsize=(cols*size, rows*size))\n    for i, img in enumerate(imgs):\n        if img_size is not None:\n            img = cv2.resize(img, img_size)\n        fig.add_subplot(rows, cols, i+1)\n        plt.imshow(img, cmap=cmap)\n    plt.suptitle(title)\n    \ndef draw_bbox(image, box, label, color):   \n    alpha = 0.1\n    alpha_box = 0.4\n    overlay_bbox = image.copy()\n    overlay_text = image.copy()\n    output = image.copy()\n\n    text_width, text_height = cv2.getTextSize(label.upper(), cv2.FONT_HERSHEY_SIMPLEX, 0.6, 1)[0]\n    cv2.rectangle(overlay_bbox, (box[0], box[1]), (box[2], box[3]),\n                color, -1)\n    cv2.addWeighted(overlay_bbox, alpha, output, 1 - alpha, 0, output)\n    cv2.rectangle(overlay_text, (box[0], box[1]-7-text_height), (box[0]+text_width+2, box[1]),\n                (0, 0, 0), -1)\n    cv2.addWeighted(overlay_text, alpha_box, output, 1 - alpha_box, 0, output)\n    cv2.rectangle(output, (box[0], box[1]), (box[2], box[3]),\n                    color, thickness)\n    cv2.putText(output, label.upper(), (box[0], box[1]-5),\n            cv2.FONT_HERSHEY_SIMPLEX, 0.6, (255, 255, 255), 1, cv2.LINE_AA)\n    return output","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:01:34.954518Z","iopub.execute_input":"2024-12-30T02:01:34.954885Z","iopub.status.idle":"2024-12-30T02:01:34.967291Z","shell.execute_reply.started":"2024-12-30T02:01:34.954857Z","shell.execute_reply":"2024-12-30T02:01:34.966105Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Define Classes","metadata":{}},{"cell_type":"code","source":"labels =  [\n            \"__ignore__\",\n            \"Aortic_enlargement\",\n            \"Atelectasis\",\n            \"Calcification\",\n            \"Cardiomegaly\",\n            \"Consolidation\",\n            \"ILD\",\n            \"Infiltration\",\n            \"Lung_Opacity\",\n            \"Nodule/Mass\",\n            \"Other_lesion\",\n            \"Pleural_effusion\",\n            \"Pleural_thickening\",\n            \"Pneumothorax\",\n            \"Pulmonary_fibrosis\"\n            ]\nviz_labels = labels[1:]","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:01:34.968639Z","iopub.execute_input":"2024-12-30T02:01:34.969083Z","iopub.status.idle":"2024-12-30T02:01:34.991534Z","shell.execute_reply.started":"2024-12-30T02:01:34.969043Z","shell.execute_reply":"2024-12-30T02:01:34.990534Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Visualize Original Bboxes","metadata":{}},{"cell_type":"code","source":"# map label_id to specify color\n#label2color = [[random.randint(0,255) for i in range(3)] for class_id in viz_labels]\nlabel2color = [[59, 238, 119], [222, 21, 229], [94, 49, 164], [206, 221, 133], [117, 75, 3],\n                 [210, 224, 119], [211, 176, 166], [63, 7, 197], [102, 65, 77], [194, 134, 175],\n                 [209, 219, 50], [255, 44, 47], [89, 125, 149], [110, 27, 100]]\n\nthickness = 3\nimgs = []\n\nfor img_id, path in zip(train_annotations['image_id'][:6], train_annotations['image_path'][:6]):\n\n    boxes = train_annotations.loc[train_annotations['image_id'] == img_id,\n                                  ['x_min', 'y_min', 'x_max', 'y_max']].values\n    img_labels = train_annotations.loc[train_annotations['image_id'] == img_id, ['class_id']].values.squeeze()\n    \n    img = cv2.imread(path)\n    \n    for label_id, box in zip(img_labels, boxes):\n        color = label2color[label_id]\n        img = draw_bbox(img, list(np.int_(box)), viz_labels[label_id], color)\n    imgs.append(img)\n\nplot_imgs(imgs, size=9, cmap=None)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:01:34.992364Z","iopub.execute_input":"2024-12-30T02:01:34.992738Z","iopub.status.idle":"2024-12-30T02:01:43.857538Z","shell.execute_reply.started":"2024-12-30T02:01:34.992699Z","shell.execute_reply":"2024-12-30T02:01:43.85571Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Exploring Techniques to Combine Bboxes\n## Non-maximum Suppression (NMS)\n### Non-maximum Suppression (NMS): Loại bỏ các bbox chồng chéo dựa trên ngưỡng Intersection over Union (IoU) và điểm tin cậy.","metadata":{}},{"cell_type":"code","source":"iou_thr = 0.5\nskip_box_thr = 0.0001\nviz_images = []\n\nfor i, path in tqdm(enumerate(imagepaths[5:8])):\n    img_array  = cv2.imread(path)\n    image_basename = Path(path).stem\n    print(f\"(\\'{image_basename}\\', \\'{path}\\')\")\n    img_annotations = train_annotations[train_annotations.image_id==image_basename]\n\n    boxes_viz = img_annotations[['x_min', 'y_min', 'x_max', 'y_max']].to_numpy().tolist()\n    labels_viz = img_annotations['class_id'].to_numpy().tolist()\n    \n    print(\"Bboxes before nms:\\n\", boxes_viz)\n    print(\"Labels before nms:\\n\", labels_viz)\n    \n    ## Visualize Original Bboxes\n    img_before = img_array.copy()\n    for box, label in zip(boxes_viz, labels_viz):\n        x_min, y_min, x_max, y_max = (box[0], box[1], box[2], box[3])\n        color = label2color[int(label)]\n        img_before = draw_bbox(img_before, list(np.int_(box)), viz_labels[label], color)\n    viz_images.append(img_before)\n    \n    boxes_list = []\n    scores_list = []\n    labels_list = []\n    weights = []\n    \n    boxes_single = []\n    labels_single = []\n    \n    cls_ids = img_annotations['class_id'].unique().tolist()\n    count_dict = Counter(img_annotations['class_id'].tolist())\n    print(count_dict)\n\n    for cid in cls_ids:       \n        ## Performing Fusing operation only for multiple bboxes with the same label\n        if count_dict[cid]==1:\n            labels_single.append(cid)\n            boxes_single.append(img_annotations[img_annotations.class_id==cid][['x_min', 'y_min', 'x_max', 'y_max']].to_numpy().squeeze().tolist())\n\n        else:\n            cls_list =img_annotations[img_annotations.class_id==cid]['class_id'].tolist()\n            labels_list.append(cls_list)\n            bbox = img_annotations[img_annotations.class_id==cid][['x_min', 'y_min', 'x_max', 'y_max']].to_numpy()\n            ## Normalizing Bbox by Image Width and Height\n            bbox = bbox/(img_array.shape[1], img_array.shape[0], img_array.shape[1], img_array.shape[0])\n            bbox = np.clip(bbox, 0, 1)\n            boxes_list.append(bbox.tolist())\n            scores_list.append(np.ones(len(cls_list)).tolist())\n\n            weights.append(1)\n            \n    # Perform NMS\n    boxes, scores, box_labels = nms(boxes_list, scores_list, labels_list, weights=weights,\n                                    iou_thr=iou_thr)\n    \n    boxes = boxes*(img_array.shape[1], img_array.shape[0], img_array.shape[1], img_array.shape[0])\n    boxes = boxes.round(1).tolist()\n    box_labels = box_labels.astype(int).tolist()\n\n    boxes.extend(boxes_single)\n    box_labels.extend(labels_single)\n    \n    print(\"Bboxes after nms:\\n\", boxes)\n    print(\"Labels after nms:\\n\", box_labels)\n    \n    ## Visualize Bboxes after operation\n    img_after = img_array.copy()\n    for box, label in zip(boxes, box_labels):\n        color = label2color[int(label)]\n        img_after = draw_bbox(img_after, list(np.int_(box)), viz_labels[label], color)\n    viz_images.append(img_after)\n    print()\n        \nplot_imgs(viz_images, cmap=None)\nplt.figtext(0.3, 0.9,\"Original Bboxes\", va=\"top\", ha=\"center\", size=25)\nplt.figtext(0.73, 0.9,\"Non-max Suppression\", va=\"top\", ha=\"center\", size=25)\nplt.savefig('nms.png', bbox_inches='tight')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:01:43.858778Z","iopub.execute_input":"2024-12-30T02:01:43.859334Z","iopub.status.idle":"2024-12-30T02:02:07.37946Z","shell.execute_reply.started":"2024-12-30T02:01:43.859289Z","shell.execute_reply":"2024-12-30T02:02:07.377832Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Soft-NMS\n### Soft-NMS: Tương tự như NMS nhưng giảm dần điểm tin cậy của các bbox chồng chéo thay vì loại bỏ chúng hoàn toàn.","metadata":{}},{"cell_type":"code","source":"iou_thr = 0.5\nskip_box_thr = 0.0001\nviz_images = []\nsigma = 0.1\n\nfor i, path in tqdm(enumerate(imagepaths[5:8])):\n    img_array  = cv2.imread(path)\n    image_basename = Path(path).stem\n    print(f\"(\\'{image_basename}\\', \\'{path}\\')\")\n    img_annotations = train_annotations[train_annotations.image_id==image_basename]\n    \n    boxes_viz = img_annotations[['x_min', 'y_min', 'x_max', 'y_max']].to_numpy().tolist()\n    labels_viz = img_annotations['class_id'].to_numpy().tolist()\n    \n    print(\"Bboxes before soft_nms:\\n\", boxes_viz)\n    print(\"Labels before soft_nms:\\n\", labels_viz)\n    \n    ## Visualize Original Bboxes\n    img_before = img_array.copy()\n    for box, label in zip(boxes_viz, labels_viz):\n        x_min, y_min, x_max, y_max = (box[0], box[1], box[2], box[3])\n        color = label2color[int(label)]\n        img_before = draw_bbox(img_before, list(np.int_(box)), viz_labels[label], color)\n    viz_images.append(img_before)\n    \n    boxes_list = []\n    scores_list = []\n    labels_list = []\n    weights = []\n    \n    boxes_single = []\n    labels_single = []\n    \n    cls_ids = img_annotations['class_id'].unique().tolist()\n    count_dict = Counter(img_annotations['class_id'].tolist())\n    print(count_dict)\n\n    for cid in cls_ids:       \n        ## Performing Fusing operation only for multiple bboxes with the same label\n        if count_dict[cid]==1:\n            labels_single.append(cid)\n            boxes_single.append(img_annotations[img_annotations.class_id==cid][['x_min', 'y_min', 'x_max', 'y_max']].to_numpy().squeeze().tolist())\n\n        else:\n            cls_list =img_annotations[img_annotations.class_id==cid]['class_id'].tolist()\n            labels_list.append(cls_list)\n            bbox = img_annotations[img_annotations.class_id==cid][['x_min', 'y_min', 'x_max', 'y_max']].to_numpy()\n            ## Normalizing Bbox by Image Width and Height\n            bbox = bbox/(img_array.shape[1], img_array.shape[0], img_array.shape[1], img_array.shape[0])\n            bbox = np.clip(bbox, 0, 1)\n            boxes_list.append(bbox.tolist())\n            scores_list.append(np.ones(len(cls_list)).tolist())\n\n            weights.append(1)\n            \n        \n    # Perform Soft-NMS\n    boxes, scores, box_labels = soft_nms(boxes_list, scores_list, labels_list, weights=weights,\n                                         iou_thr=iou_thr, sigma=sigma, thresh=skip_box_thr)\n    \n    \n    boxes = boxes*(img_array.shape[1], img_array.shape[0], img_array.shape[1], img_array.shape[0])\n    boxes = boxes.round(1).tolist()\n    box_labels = box_labels.astype(int).tolist()\n    \n    boxes.extend(boxes_single)\n    box_labels.extend(labels_single)\n    \n    print(\"Bboxes after soft_nms:\\n\", boxes)\n    print(\"Labels after soft_nms:\\n\", box_labels)\n    \n    ## Visualize Bboxes after operation\n    img_after = img_array.copy()\n    for box, label in zip(boxes, box_labels):\n        color = label2color[int(label)]\n        img_after = draw_bbox(img_after, list(np.int_(box)), viz_labels[label], color)\n    viz_images.append(img_after)\n    print()\n        \nplot_imgs(viz_images, cmap=None)\nplt.figtext(0.3, 0.9,\"Original Bboxes\", va=\"top\", ha=\"center\", size=25)\nplt.figtext(0.73, 0.9,\"Soft NMS\", va=\"top\", ha=\"center\", size=25)\nplt.savefig('snms.png', bbox_inches='tight')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:02:07.380964Z","iopub.execute_input":"2024-12-30T02:02:07.3815Z","iopub.status.idle":"2024-12-30T02:02:26.11736Z","shell.execute_reply.started":"2024-12-30T02:02:07.381448Z","shell.execute_reply":"2024-12-30T02:02:26.115951Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Non-maximum Weighted\n### Non-maximum Weighted (NMW): Tương tự như NMS nhưng sử dụng trọng số để ưu tiên các bbox từ các mô hình hoặc nguồn khác nhau.","metadata":{}},{"cell_type":"code","source":"iou_thr = 0.5\nskip_box_thr = 0.0001\nviz_images = []\n\nfor i, path in tqdm(enumerate(imagepaths[5:8])):\n    img_array  = cv2.imread(path)\n    image_basename = Path(path).stem\n    print(f\"(\\'{image_basename}\\', \\'{path}\\')\")\n    img_annotations = train_annotations[train_annotations.image_id==image_basename]\n\n    boxes_viz = img_annotations[['x_min', 'y_min', 'x_max', 'y_max']].to_numpy().tolist()\n    labels_viz = img_annotations['class_id'].to_numpy().tolist()\n    \n    print(\"Bboxes before non_maximum_weighted:\\n\", boxes_viz)\n    print(\"Labels before non_maximum_weighted:\\n\", labels_viz)\n    \n    ## Visualize Original Bboxes\n    img_before = img_array.copy()\n    for box, label in zip(boxes_viz, labels_viz):\n        x_min, y_min, x_max, y_max = (box[0], box[1], box[2], box[3])\n        color = label2color[int(label)]\n        img_before = draw_bbox(img_before, list(np.int_(box)), viz_labels[label], color)\n    viz_images.append(img_before)\n    \n    boxes_list = []\n    scores_list = []\n    labels_list = []\n    weights = []\n    \n    boxes_single = []\n    labels_single = []\n    \n    cls_ids = img_annotations['class_id'].unique().tolist()\n    count_dict = Counter(img_annotations['class_id'].tolist())\n    print(count_dict)\n\n    for cid in cls_ids:       \n        ## Performing Fusing operation only for multiple bboxes with the same label\n        if count_dict[cid]==1:\n            labels_single.append(cid)\n            boxes_single.append(img_annotations[img_annotations.class_id==cid][['x_min', 'y_min', 'x_max', 'y_max']].to_numpy().squeeze().tolist())\n\n        else:\n            cls_list =img_annotations[img_annotations.class_id==cid]['class_id'].tolist()\n            labels_list.append(cls_list)\n            bbox = img_annotations[img_annotations.class_id==cid][['x_min', 'y_min', 'x_max', 'y_max']].to_numpy()\n            ## Normalizing Bbox by Image Width and Height\n            bbox = bbox/(img_array.shape[1], img_array.shape[0], img_array.shape[1], img_array.shape[0])\n            bbox = np.clip(bbox, 0, 1)\n            boxes_list.append(bbox.tolist())\n            scores_list.append(np.ones(len(cls_list)).tolist())\n\n            weights.append(1)\n            \n\n    # Perform Non-maximum Weighted\n    boxes, scores, box_labels = non_maximum_weighted(boxes_list, scores_list, labels_list,\n                                                     weights=weights, iou_thr=iou_thr,skip_box_thr=skip_box_thr)\n    \n    boxes = boxes*(img_array.shape[1], img_array.shape[0], img_array.shape[1], img_array.shape[0])\n    boxes = boxes.round(1).tolist()\n    box_labels = box_labels.astype(int).tolist()\n\n    boxes.extend(boxes_single)\n    box_labels.extend(labels_single)\n    \n    print(\"Bboxes after non_maximum_weighted:\\n\", boxes)\n    print(\"Labels after non_maximum_weighted:\\n\", box_labels)\n    \n    ## Visualize Bboxes after operation\n    img_after = img_array.copy()\n    for box, label in zip(boxes, box_labels):\n        color = label2color[int(label)]\n        img_after = draw_bbox(img_after, list(np.int_(box)), viz_labels[label], color)\n    viz_images.append(img_after)\n    print()\n        \nplot_imgs(viz_images, cmap=None)\nplt.figtext(0.3, 0.9,\"Original Bboxes\", va=\"top\", ha=\"center\", size=25)\nplt.figtext(0.73, 0.9,\"Non-maximum Weighted\", va=\"top\", ha=\"center\", size=25)\nplt.savefig('nmw.png', bbox_inches='tight')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:02:26.118827Z","iopub.execute_input":"2024-12-30T02:02:26.119267Z","iopub.status.idle":"2024-12-30T02:02:45.117862Z","shell.execute_reply.started":"2024-12-30T02:02:26.11923Z","shell.execute_reply":"2024-12-30T02:02:45.116464Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Weighted boxes fusion (WBF)\n### Weighted Boxes Fusion (WBF) là một kỹ thuật hiệu quả để kết hợp các dự đoán từ nhiều mô hình phát hiện đối tượng. Không giống như Non-Maximum Suppression (NMS) loại bỏ các hộp chồng chéo dựa trên điểm tin cậy, WBF sử dụng tất cả các hộp dự đoán để tính toán các hộp trung bình có trọng số.","metadata":{}},{"cell_type":"code","source":"iou_thr = 0.5\nskip_box_thr = 0.0001\nviz_images = []\nsigma = 0.1\n\nfor i, path in tqdm(enumerate(imagepaths[5:8])):\n    img_array  = cv2.imread(path)\n    image_basename = Path(path).stem\n    print(f\"(\\'{image_basename}\\', \\'{path}\\')\")\n    img_annotations = train_annotations[train_annotations.image_id==image_basename]\n\n    boxes_viz = img_annotations[['x_min', 'y_min', 'x_max', 'y_max']].to_numpy().tolist()\n    labels_viz = img_annotations['class_id'].to_numpy().tolist()\n    \n    print(\"Bboxes before WBF:\\n\", boxes_viz)\n    print(\"Labels before WBF:\\n\", labels_viz)\n    \n    ## Visualize Original Bboxes\n    img_before = img_array.copy()\n    for box, label in zip(boxes_viz, labels_viz):\n        x_min, y_min, x_max, y_max = (box[0], box[1], box[2], box[3])\n        color = label2color[int(label)]\n        img_before = draw_bbox(img_before, list(np.int_(box)), viz_labels[label], color)\n    viz_images.append(img_before)\n    \n    boxes_list = []\n    scores_list = []\n    labels_list = []\n    weights = []\n    \n    boxes_single = []\n    labels_single = []\n    \n    cls_ids = img_annotations['class_id'].unique().tolist()\n    count_dict = Counter(img_annotations['class_id'].tolist())\n    print(count_dict)\n\n    for cid in cls_ids:       \n        ## Performing Fusing operation only for multiple bboxes with the same label\n        if count_dict[cid]==1:\n            labels_single.append(cid)\n            boxes_single.append(img_annotations[img_annotations.class_id==cid][['x_min', 'y_min', 'x_max', 'y_max']].to_numpy().squeeze().tolist())\n\n        else:\n            cls_list =img_annotations[img_annotations.class_id==cid]['class_id'].tolist()\n            labels_list.append(cls_list)\n            bbox = img_annotations[img_annotations.class_id==cid][['x_min', 'y_min', 'x_max', 'y_max']].to_numpy()\n            ## Normalizing Bbox by Image Width and Height\n            bbox = bbox/(img_array.shape[1], img_array.shape[0], img_array.shape[1], img_array.shape[0])\n            bbox = np.clip(bbox, 0, 1)\n            boxes_list.append(bbox.tolist())\n            scores_list.append(np.ones(len(cls_list)).tolist())\n\n            weights.append(1)\n            \n\n    # Perform WBF\n    boxes, scores, box_labels= weighted_boxes_fusion(boxes_list, scores_list, labels_list, weights=weights,\n                                                     iou_thr=iou_thr, skip_box_thr=skip_box_thr)\n    \n    \n    boxes = boxes*(img_array.shape[1], img_array.shape[0], img_array.shape[1], img_array.shape[0])\n    boxes = boxes.round(1).tolist()\n    box_labels = box_labels.astype(int).tolist()\n\n    boxes.extend(boxes_single)\n    box_labels.extend(labels_single)\n    \n    print(\"Bboxes after WBF:\\n\", boxes)\n    print(\"Labels after WBF:\\n\", box_labels)\n    \n    ## Visualize Bboxes after operation\n    img_after = img_array.copy()\n    for box, label in zip(boxes, box_labels):\n        color = label2color[int(label)]\n        img_after = draw_bbox(img_after, list(np.int_(box)), viz_labels[label], color)\n    viz_images.append(img_after)\n    print()\n        \nplot_imgs(viz_images, cmap=None)\nplt.figtext(0.3, 0.9,\"Original Bboxes\", va=\"top\", ha=\"center\", size=25)\nplt.figtext(0.73, 0.9,\"WBF\", va=\"top\", ha=\"center\", size=25)\nplt.savefig('wbf.png', bbox_inches='tight')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:02:45.119462Z","iopub.execute_input":"2024-12-30T02:02:45.119883Z","iopub.status.idle":"2024-12-30T02:03:03.651552Z","shell.execute_reply.started":"2024-12-30T02:02:45.119847Z","shell.execute_reply":"2024-12-30T02:03:03.650026Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Weighted Boxes Fusion seems to give  gives better results comparing to others in this situation considering that we don't have the confidence/weights of the annotations done by different radiologists","metadata":{}},{"cell_type":"markdown","source":"# Building YOLO DATASET\n## Train, Validation & Test Split","metadata":{}},{"cell_type":"code","source":"# First, let's add back the images with class_id 14 (no findings)\nall_annotations = pd.read_csv(\"/kaggle/input/vinbigdata-chest-xray-abnormalities-detection/train.csv\")\nall_annotations['image_path'] = all_annotations['image_id'].map(lambda x:os.path.join('/kaggle/input/vinbigdata-chest-xray-original-png/train', str(x)+'.png'))\nall_annotations.head(5)","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:03:03.653022Z","iopub.execute_input":"2024-12-30T02:03:03.65352Z","iopub.status.idle":"2024-12-30T02:03:03.84648Z","shell.execute_reply.started":"2024-12-30T02:03:03.653473Z","shell.execute_reply":"2024-12-30T02:03:03.845377Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Calculate Class Weights\n\nimport numpy as np\n\n# Count the number of annotations per class\nclass_counts = all_annotations['class_id'].value_counts().to_dict()\n\n# Total number of samples\ntotal_samples = sum(class_counts.values())\n\n# Calculate class weights based on inverse frequency\nclass_weights = {cls: total_samples / (len(class_counts) * count) for cls, count in class_counts.items()}\n\nprint(f\"Class Weights: {class_weights}\")","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:03:03.847608Z","iopub.execute_input":"2024-12-30T02:03:03.847938Z","iopub.status.idle":"2024-12-30T02:03:03.858939Z","shell.execute_reply.started":"2024-12-30T02:03:03.847914Z","shell.execute_reply":"2024-12-30T02:03:03.857361Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Set random seed for reproducibility\nnp.random.seed(42)\n\n# Get unique image IDs\nimage_ids = all_annotations['image_id'].unique()\n\n# Split the data\ntrain_ids, temp_ids = train_test_split(image_ids, test_size=0.25, random_state=42)\nval_ids, test_ids = train_test_split(temp_ids, test_size=0.2, random_state=42)\n\nprint(f\"Train set: {len(train_ids)} images\")\nprint(f\"Validation set: {len(val_ids)} images\")\nprint(f\"Test set: {len(test_ids)} images\")","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:03:03.860206Z","iopub.execute_input":"2024-12-30T02:03:03.860566Z","iopub.status.idle":"2024-12-30T02:03:03.896486Z","shell.execute_reply.started":"2024-12-30T02:03:03.860538Z","shell.execute_reply":"2024-12-30T02:03:03.895168Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from pathlib import Path\n\n# Create directories for YOLO dataset\nyolo_dir = Path('./vinbigdata-yolo-dataset-with-wbf-640px-1class')\nyolo_dir.mkdir(exist_ok=True)\n\nfor split in ['train', 'val', 'test']:\n    (yolo_dir / split / 'images').mkdir(parents=True, exist_ok=True)\n    (yolo_dir / split / 'labels').mkdir(parents=True, exist_ok=True)","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:03:03.897664Z","iopub.execute_input":"2024-12-30T02:03:03.897969Z","iopub.status.idle":"2024-12-30T02:03:03.907713Z","shell.execute_reply.started":"2024-12-30T02:03:03.89794Z","shell.execute_reply":"2024-12-30T02:03:03.906507Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from typing import List\ndef ls(path: Path) -> List[Path]:\n    return list(path.iterdir())","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:03:03.908886Z","iopub.execute_input":"2024-12-30T02:03:03.909343Z","iopub.status.idle":"2024-12-30T02:03:03.933448Z","shell.execute_reply.started":"2024-12-30T02:03:03.909298Z","shell.execute_reply":"2024-12-30T02:03:03.93231Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"ls(yolo_dir)","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:03:03.934511Z","iopub.execute_input":"2024-12-30T02:03:03.934866Z","iopub.status.idle":"2024-12-30T02:03:03.957781Z","shell.execute_reply.started":"2024-12-30T02:03:03.934827Z","shell.execute_reply":"2024-12-30T02:03:03.956537Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## YoLo Format Overview\n### Basic Syntax\n","metadata":{}},{"cell_type":"markdown","source":"Cú pháp cơ bản của định dạng YOLO\nMỗi file nhãn (label file) của YOLO sẽ có cùng tên với file ảnh tương ứng nhưng có phần mở rộng .txt. File này chứa một hoặc nhiều dòng, mỗi dòng mô tả một bounding box cho một đối tượng trong ảnh. Dưới đây là cú pháp cơ bản cho mỗi dòng trong file nhãn:\n\ncsharp\nCopy code\n<object-class> <x_center> <y_center> <width> <height>\nTrong đó:\n\n<object-class>: Chỉ số của lớp đối tượng (bắt đầu từ 0). Ví dụ: nếu bạn có 3 lớp: \"cat\", \"dog\", và \"person\", thì \"cat\" có thể có giá trị 0, \"dog\" có giá trị 1, và \"person\" có giá trị 2.\n<x_center>: Tọa độ x của tâm bounding box, được chuẩn hóa bằng chiều rộng của ảnh (nằm trong khoảng từ 0 đến 1).\n<y_center>: Tọa độ y của tâm bounding box, được chuẩn hóa bằng chiều cao của ảnh (nằm trong khoảng từ 0 đến 1).\n<width>: Chiều rộng của bounding box, được chuẩn hóa bằng chiều rộng của ảnh (nằm trong khoảng từ 0 đến 1).\n<height>: Chiều cao của bounding box, được chuẩn hóa bằng chiều cao của ảnh (nằm trong khoảng từ 0 đến 1).\nVí dụ về một file YOLO label\nGiả sử bạn có một ảnh image1.jpg với kích thước 1024x1024 pixels và file nhãn tương ứng image1.txt có nội dung sau:\n\nCopy code\n0 0.5 0.5 0.2 0.3\n1 0.7 0.8 0.1 0.1\nGiải thích:\n\nDòng 1: 0 0.5 0.5 0.2 0.3\n\n0: Đây là lớp đối tượng đầu tiên (ví dụ: \"cat\").\n0.5: Tọa độ x của tâm bounding box (50% từ cạnh trái của ảnh).\n0.5: Tọa độ y của tâm bounding box (50% từ cạnh trên của ảnh).\n0.2: Chiều rộng của bounding box (20% của chiều rộng ảnh).\n0.3: Chiều cao của bounding box (30% của chiều cao ảnh).\nDòng 2: 1 0.7 0.8 0.1 0.1\n\n1: Đây là lớp đối tượng thứ hai (ví dụ: \"dog\").\n0.7: Tọa độ x của tâm bounding box (70% từ cạnh trái của ảnh).\n0.8: Tọa độ y của tâm bounding box (80% từ cạnh trên của ảnh).\n0.1: Chiều rộng của bounding box (10% của chiều rộng ảnh).\n0.1: Chiều cao của bounding box (10% của chiều cao ảnh).","metadata":{}},{"cell_type":"code","source":"# Function to convert bbox to YOLO format\ndef convert_to_yolo_format(box, img_width, img_height):\n    x_center = (box[0] + box[2]) / 2 / img_width\n    y_center = (box[1] + box[3]) / 2 / img_height\n    width = (box[2] - box[0]) / img_width\n    height = (box[3] - box[1]) / img_height\n    return [x_center, y_center, width, height]","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:03:03.961968Z","iopub.execute_input":"2024-12-30T02:03:03.962406Z","iopub.status.idle":"2024-12-30T02:03:03.978044Z","shell.execute_reply.started":"2024-12-30T02:03:03.962347Z","shell.execute_reply":"2024-12-30T02:03:03.976617Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Prepare the datasets\nviz_images = {\n    'train': [],\n    'val': [],\n    'test': []\n}","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:03:03.979577Z","iopub.execute_input":"2024-12-30T02:03:03.980012Z","iopub.status.idle":"2024-12-30T02:03:04.000647Z","shell.execute_reply.started":"2024-12-30T02:03:03.979955Z","shell.execute_reply":"2024-12-30T02:03:03.999219Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# # Prepare the datasets\n# for split, ids in [('train', train_ids), ('val', val_ids), ('test', test_ids)]:\n#     saved_images_count = 0  # Counter to track saved images per split\n#     for img_id in tqdm(ids, desc=f\"Processing {split} set\"):\n#         img_annotations = all_annotations[all_annotations.image_id == img_id]\n#         img_path = img_annotations.iloc[0]['image_path']\n#         img = cv2.imread(img_path)\n        \n#         # Get original image dimensions\n#         orig_height, orig_width = img.shape[:2]\n        \n#         # Resize the image to 640x640 pixels\n#         img = cv2.resize(img, (640, 640))\n#         img_height, img_width = 640, 640  # Set the new width and height after resizing\n        \n#         # Calculate scale factors for width and height\n#         width_scale = img_width / orig_width\n#         height_scale = img_height / orig_height\n        \n#         # Save the resized image to the appropriate directory\n#         output_img_path = yolo_dir / split / 'images' / f\"{img_id}.jpg\"\n#         cv2.imwrite(str(output_img_path), img)  # Save the resized image\n        \n#         # Prepare annotations\n#         with open(yolo_dir / split / 'labels' / f\"{img_id}.txt\", 'w') as f:\n#             if 14 in img_annotations['class_id'].values:\n#                 # No findings case\n#                 f.write(\"14 0.5 0.5 1 1\\n\")\n#             else:\n#                 boxes_list = []\n#                 scores_list = []\n#                 labels_list = []\n                \n#                 cls_ids = img_annotations['class_id'].unique().tolist()\n#                 for cid in cls_ids:\n#                     cls_annotations = img_annotations[img_annotations.class_id == cid]\n#                     cls_boxes = cls_annotations[['x_min', 'y_min', 'x_max', 'y_max']].values\n                    \n#                     # Scale bounding box coordinates based on original image size before resize\n#                     cls_boxes[:, [0, 2]] = cls_boxes[:, [0, 2]] * width_scale  # Adjust x_min and x_max\n#                     cls_boxes[:, [1, 3]] = cls_boxes[:, [1, 3]] * height_scale  # Adjust y_min and y_max\n                    \n#                     # Normalize box coordinates based on resized image dimensions (1280x1280)\n#                     cls_boxes = cls_boxes / [img_width, img_height, img_width, img_height]\n#                     cls_boxes = np.clip(cls_boxes, 0, 1)  # Ensure boxes are within bounds\n                    \n#                     boxes_list.append(cls_boxes.tolist())\n#                     scores_list.append(np.ones(len(cls_boxes)).tolist())\n#                     labels_list.append([cid] * len(cls_boxes))\n                \n#                 # Visualize original bounding boxes before WBF\n#                 if saved_images_count < 5:  # Only save up to 5 images per split\n#                     img_before = img.copy()\n                    \n#                     # Flatten and ensure all boxes have 4 elements\n#                     flat_boxes = [box for sublist in boxes_list for box in sublist if len(box) == 4]\n#                     flat_labels = [label for sublist in labels_list for label in sublist]\n\n#                     for box, label in zip(flat_boxes, flat_labels):\n#                         # Calculate bounding box coordinates based on resized image dimensions\n#                         x_min, y_min, x_max, y_max = np.array(box) * [img_width, img_height, img_width, img_height]\n#                         color = label2color[int(label)]\n#                         img_before = draw_bbox(img_before, [int(x_min), int(y_min), int(x_max), int(y_max)], viz_labels[int(label)], color)\n                    \n#                 # Apply WBF (Weighted Boxes Fusion)\n#                 boxes, scores, labels = weighted_boxes_fusion(\n#                     boxes_list, scores_list, labels_list, \n#                     weights=None, iou_thr=0.5, skip_box_thr=0.0001\n#                 )\n                \n#                 # Write YOLO format annotations\n#                 for box, label in zip(boxes, labels):\n#                     # YOLO expects normalized coordinates, so we don't need to scale them back\n#                     yolo_box = convert_to_yolo_format(box, 1, 1)  # already normalized\n#                     f.write(f\"{int(label)} {' '.join(map(str, yolo_box))}\\n\")\n#                 # Findings case\n#                 f.write(\"15 0.5 0.5 1 1\\n\")\n                \n#                 # Visualize bounding boxes after WBF\n#                 if saved_images_count < 5:  # Only save up to 5 images per split\n#                     img_after = img.copy()\n#                     for box, label in zip(boxes, labels):\n#                         # Calculate bounding box coordinates based on resized image dimensions\n#                         x_min, y_min, x_max, y_max = np.array(box) * [img_width, img_height, img_width, img_height]\n#                         color = label2color[int(label)]\n#                         img_after = draw_bbox(img_after, [int(x_min), int(y_min), int(x_max), int(y_max)], viz_labels[int(label)], color)\n                    \n#                     # Append both images (before and after WBF) to visualization list\n#                     viz_images[split].append((img_before, img_after))\n#                     saved_images_count += 1\n\n# print(\"YOLO dataset creation and visualization completed.\")","metadata":{"execution":{"iopub.status.busy":"2024-12-30T02:03:04.0019Z","iopub.execute_input":"2024-12-30T02:03:04.002275Z","iopub.status.idle":"2024-12-30T02:03:04.021119Z","shell.execute_reply.started":"2024-12-30T02:03:04.002246Z","shell.execute_reply":"2024-12-30T02:03:04.019916Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Prepare the datasets\nfor split, ids in [('train', train_ids), ('val', val_ids), ('test', test_ids)]:\n    saved_images_count = 0  # Counter to track saved images per split\n    for img_id in tqdm(ids, desc=f\"Processing {split} set\"):\n        img_annotations = all_annotations[all_annotations.image_id == img_id]\n        img_path = img_annotations.iloc[0]['image_path']\n        img = cv2.imread(img_path)\n\n        if img is None:\n            print(f\"Could not read image: {img_path}\")\n            continue\n\n        # Get original image dimensions\n        orig_height, orig_width = img.shape[:2]\n\n        # Resize the image to 640x640 pixels\n        img = cv2.resize(img, (640, 640))\n        img_height, img_width = 640, 640  # Set the new width and height after resizing\n\n        # Calculate scale factors for width and height\n        width_scale = img_width / orig_width\n        height_scale = img_height / orig_height\n\n        # Save the resized image to the appropriate directory\n        output_img_path = yolo_dir / split / 'images' / f\"{img_id}.jpg\"\n        cv2.imwrite(str(output_img_path), img)  # Save the resized image\n\n        # Prepare annotations\n        label_path = yolo_dir / split / 'labels' / f\"{img_id}.txt\"\n        annotations_exist = False  # Flag to check if there are valid annotations\n\n        with open(label_path, 'w') as f:\n            # Case 1: No findings (class_id = 14), vì xuất ra 2class nên chuyển thành id 1\n            if 14 in img_annotations['class_id'].values:\n                # f.write(\"14 0.5 0.5 1 1\\n\")\n                f.write(\"1 0.5 0.5 1 1\\n\")\n                annotations_exist = True\n\n            # Case 2: Aortic enlargement (class_id = 0)\n            aortic_annotations = img_annotations[img_annotations['class_id'] == 0]\n            boxes_list = []  # Store bounding boxes for visualization\n\n            for _, ann in aortic_annotations.iterrows():\n                x_min, y_min, x_max, y_max = ann[['x_min', 'y_min', 'x_max', 'y_max']].values\n\n                # Scale bounding box coordinates based on original image size before resize\n                x_min *= width_scale\n                x_max *= width_scale\n                y_min *= height_scale\n                y_max *= height_scale\n\n                # Normalize box coordinates based on resized image dimensions\n                x_center = (x_min + x_max) / 2 / img_width\n                y_center = (y_min + y_max) / 2 / img_height\n                box_width = (x_max - x_min) / img_width\n                box_height = (y_max - y_min) / img_height\n\n                # Append box to list for visualization\n                boxes_list.append([x_min, y_min, x_max, y_max])\n\n                # Write normalized YOLO format annotation\n                f.write(f\"0 {x_center} {y_center} {box_width} {box_height}\\n\")\n                annotations_exist = True\n\n            # Visualization for each image in Case 2: Aortic enlargement\n            if saved_images_count < 5 and annotations_exist:\n                img_before = img.copy()\n                img_after = img.copy()  # Placeholder for after WBF visualization\n\n                for box in boxes_list:\n                    # Draw bounding boxes on img_before\n                    x_min, y_min, x_max, y_max = map(int, box)\n                    color = (0, 255, 0)  # Green for visualization\n                    img_before = cv2.rectangle(img_before, (x_min, y_min), (x_max, y_max), color, 2)\n                    img_before = cv2.putText(img_before, \"Aortic_enlargement\", (x_min, y_min - 10),\n                                             cv2.FONT_HERSHEY_SIMPLEX, 0.5, color, 1)\n\n                # Example for img_after (you can apply WBF here if needed)\n                for box in boxes_list:\n                    # Draw bounding boxes on img_after\n                    x_min, y_min, x_max, y_max = map(int, box)\n                    color = (255, 0, 0)  # Red for visualization after WBF\n                    img_after = cv2.rectangle(img_after, (x_min, y_min), (x_max, y_max), color, 2)\n                    img_after = cv2.putText(img_after, \"Aortic_enlargement (After)\", (x_min, y_min - 10),\n                                            cv2.FONT_HERSHEY_SIMPLEX, 0.5, color, 1)\n\n                # Append both images (before and after) to visualization list\n                viz_images[split].append((img_before, img_after))\n                saved_images_count += 1\n\n        # Remove label file if no annotations exist\n        if not annotations_exist:\n            os.remove(label_path)\n\nprint(\"YOLO dataset creation with visualization for Aortic enlargement completed.\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# # Prepare the datasets\n# for split, ids in [('train', train_ids), ('val', val_ids), ('test', test_ids)]:\n#     saved_images_count = 0  # Counter to track saved images per split\n#     for img_id in tqdm(ids, desc=f\"Processing {split} set\"):\n#         img_annotations = all_annotations[all_annotations.image_id == img_id]\n#         img_path = img_annotations.iloc[0]['image_path']\n#         img = cv2.imread(img_path)\n        \n#         # Get original image dimensions\n#         orig_height, orig_width = img.shape[:2]\n        \n#         # Resize the image to 640x640 pixels\n#         img = cv2.resize(img, (640, 640))\n#         img_height, img_width = 640, 640  # Set the new width and height after resizing\n        \n#         # Calculate scale factors for width and height\n#         width_scale = img_width / orig_width\n#         height_scale = img_height / orig_height\n        \n#         # Save the resized image to the appropriate directory\n#         output_img_path = yolo_dir / split / 'images' / f\"{img_id}.jpg\"\n#         cv2.imwrite(str(output_img_path), img)  # Save the resized image\n        \n#         # Prepare annotations\n#         with open(yolo_dir / split / 'labels' / f\"{img_id}.txt\", 'w') as f:\n#             if 14 in img_annotations['class_id'].values:\n#                 # No findings case\n#                 f.write(\"14 0.5 0.5 1 1\\n\")\n#             else:\n#                 boxes_list = []\n#                 scores_list = []\n#                 labels_list = []\n                \n#                 aortic_annotations = img_annotations[img_annotations['class_id'] == 0]\n#                 for cid in aortic_annotations:\n#                     cls_annotations = img_annotations[img_annotations.class_id == cid]\n#                     cls_boxes = cls_annotations[['x_min', 'y_min', 'x_max', 'y_max']].values\n                    \n#                     # Scale bounding box coordinates based on original image size before resize\n#                     cls_boxes[:, [0, 2]] = cls_boxes[:, [0, 2]] * width_scale  # Adjust x_min and x_max\n#                     cls_boxes[:, [1, 3]] = cls_boxes[:, [1, 3]] * height_scale  # Adjust y_min and y_max\n                    \n#                     # Normalize box coordinates based on resized image dimensions (1280x1280)\n#                     cls_boxes = cls_boxes / [img_width, img_height, img_width, img_height]\n#                     cls_boxes = np.clip(cls_boxes, 0, 1)  # Ensure boxes are within bounds\n                    \n#                     boxes_list.append(cls_boxes.tolist())\n#                     scores_list.append(np.ones(len(cls_boxes)).tolist())\n#                     labels_list.append([cid] * len(cls_boxes))\n                \n#                 # Visualize original bounding boxes before WBF\n#                 if saved_images_count < 5:  # Only save up to 5 images per split\n#                     img_before = img.copy()\n                    \n#                     # Flatten and ensure all boxes have 4 elements\n#                     flat_boxes = [box for sublist in boxes_list for box in sublist if len(box) == 4]\n#                     flat_labels = [label for sublist in labels_list for label in sublist]\n\n#                     for box, label in zip(flat_boxes, flat_labels):\n#                         # Calculate bounding box coordinates based on resized image dimensions\n#                         x_min, y_min, x_max, y_max = np.array(box) * [img_width, img_height, img_width, img_height]\n#                         color = label2color[int(label)]\n#                         img_before = draw_bbox(img_before, [int(x_min), int(y_min), int(x_max), int(y_max)], viz_labels[int(label)], color)\n                    \n#                 # Apply WBF (Weighted Boxes Fusion)\n#                 boxes, scores, labels = weighted_boxes_fusion(\n#                     boxes_list, scores_list, labels_list, \n#                     weights=None, iou_thr=0.5, skip_box_thr=0.0001\n#                 )\n                \n#                 # Write YOLO format annotations\n#                 for box, label in zip(boxes, labels):\n#                     # YOLO expects normalized coordinates, so we don't need to scale them back\n#                     yolo_box = convert_to_yolo_format(box, 1, 1)  # already normalized\n#                     f.write(f\"{int(label)} {' '.join(map(str, yolo_box))}\\n\")\n#                 # Findings case\n#                 f.write(\"15 0.5 0.5 1 1\\n\")\n                \n#                 # Visualize bounding boxes after WBF\n#                 if saved_images_count < 5:  # Only save up to 5 images per split\n#                     img_after = img.copy()\n#                     for box, label in zip(boxes, labels):\n#                         # Calculate bounding box coordinates based on resized image dimensions\n#                         x_min, y_min, x_max, y_max = np.array(box) * [img_width, img_height, img_width, img_height]\n#                         color = label2color[int(label)]\n#                         img_after = draw_bbox(img_after, [int(x_min), int(y_min), int(x_max), int(y_max)], viz_labels[int(label)], color)\n                    \n#                     # Append both images (before and after WBF) to visualization list\n#                     viz_images[split].append((img_before, img_after))\n#                     saved_images_count += 1\n\n# print(\"YOLO dataset creation and visualization completed.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T02:03:04.022365Z","iopub.execute_input":"2024-12-30T02:03:04.02272Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Visualize and save images before and after WBF\nfor split, images in viz_images.items():\n    for idx, (img_before, img_after) in enumerate(images):\n        plt.figure(figsize=(20, 10))\n\n        # Plot image before WBF\n        plt.subplot(1, 2, 1)\n        plt.imshow(cv2.cvtColor(img_before, cv2.COLOR_BGR2RGB))\n        plt.title(f\"{split.capitalize()} Image {idx+1} Before WBF\")\n        plt.axis('off')\n\n        # Plot image after WBF\n        plt.subplot(1, 2, 2)\n        plt.imshow(cv2.cvtColor(img_after, cv2.COLOR_BGR2RGB))\n        plt.title(f\"{split.capitalize()} Image {idx+1} After WBF\")\n        plt.axis('off')\n\n        plt.savefig(f'{split}_image_{idx+1}_comparison.png', bbox_inches='tight')\n        plt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# # Visualize and save images before and after WBF\n# for split, images in viz_images.items():\n#     for idx, (img_before, img_after) in enumerate(images):\n#         plt.figure(figsize=(20, 10))\n\n#         # Plot image before WBF\n#         plt.subplot(1, 2, 1)\n#         plt.imshow(cv2.cvtColor(img_before, cv2.COLOR_BGR2RGB))\n#         plt.title(f\"{split.capitalize()} Image {idx+1} Before WBF\")\n#         plt.axis('off')\n\n#         # Plot image after WBF\n#         plt.subplot(1, 2, 2)\n#         plt.imshow(cv2.cvtColor(img_after, cv2.COLOR_BGR2RGB))\n#         plt.title(f\"{split.capitalize()} Image {idx+1} After WBF\")\n#         plt.axis('off')\n\n#         plt.savefig(f'{split}_image_{idx+1}_comparison.png', bbox_inches='tight')\n#         plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-12T01:53:53.52418Z","iopub.execute_input":"2024-09-12T01:53:53.524818Z","iopub.status.idle":"2024-09-12T01:54:33.060948Z","shell.execute_reply.started":"2024-09-12T01:53:53.524745Z","shell.execute_reply":"2024-09-12T01:54:33.059781Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create train.txt, val.txt, and test.txt\nfor split in ['train', 'val', 'test']:\n    with open(yolo_dir / f\"{split}.txt\", 'w') as f:\n        for img_path in (yolo_dir / split / 'images').glob('*.jpg'):\n            # Replace \"working\" in img_path with the specified value\n            modified_img_path = str(img_path.absolute()).replace(\"working\", \"input/vinbigdata-yolo-dataset-with-wbf-640px-1class\")\n            f.write(f\"{modified_img_path}\\n\")\n\nprint(\"Dataset split files created.\")","metadata":{"execution":{"iopub.status.busy":"2024-09-12T01:54:33.062509Z","iopub.execute_input":"2024-09-12T01:54:33.062921Z","iopub.status.idle":"2024-09-12T01:54:33.46601Z","shell.execute_reply.started":"2024-09-12T01:54:33.062876Z","shell.execute_reply":"2024-09-12T01:54:33.464793Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_input_dir = Path(\"/kaggle/input/vinbigdata-yolo-dataset-with-wbf-640px-1class/vinbigdata-yolo-dataset-with-wbf-640px-1class\")\n\n# # Create data.yaml file\n# data_yaml = f\"\"\"\n# train: {(data_input_dir / 'train.txt')}\n# val: {(data_input_dir / 'val.txt')}\n# test: {(data_input_dir / 'test.txt')}\n\n# nc: {len(viz_labels) + 1}  # +1 for the \"no finding\" class\n# names: {viz_labels + ['No finding']}\n# \"\"\"\n\n# with open(yolo_dir / 'data.yaml', 'w') as f:\n#     f.write(data_yaml)\n\n# print(\"data.yaml file created.\")\n\n# Add class weights to the YAML configuration\nclass_weight_list = [class_weights[cls] for cls in sorted(class_weights.keys())]\n\ndata_yaml = f\"\"\"\ntrain: {(data_input_dir / 'train.txt')}\nval: {(data_input_dir / 'val.txt')}\ntest: {(data_input_dir / 'test.txt')}\n\nnc: {2}  # 2 for bệnh động mạch vành + ['No finding']\nnames: {['Aortic_enlargement'] + ['No finding']}\n# class_weights: {class_weight_list}\n\"\"\"\n\nwith open(yolo_dir / 'data.yaml', 'w') as f:\n    f.write(data_yaml)\n\nprint(\"data.yaml file with class weights created.\")","metadata":{"execution":{"iopub.status.busy":"2024-09-12T01:54:33.467637Z","iopub.execute_input":"2024-09-12T01:54:33.468023Z","iopub.status.idle":"2024-09-12T01:54:33.478828Z","shell.execute_reply.started":"2024-09-12T01:54:33.467983Z","shell.execute_reply":"2024-09-12T01:54:33.477213Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<object-class> <x_center> <y_center> <width> <height>\n","metadata":{}},{"cell_type":"code","source":"warnings.filterwarnings(\"ignore\", category=UserWarning)","metadata":{"execution":{"iopub.status.busy":"2024-09-12T01:54:33.480665Z","iopub.execute_input":"2024-09-12T01:54:33.481334Z","iopub.status.idle":"2024-09-12T01:54:33.491809Z","shell.execute_reply.started":"2024-09-12T01:54:33.481274Z","shell.execute_reply":"2024-09-12T01:54:33.490482Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!find ./vinbigdata-yolo-dataset-with-wbf-640px-1class/train -type f | wc -l","metadata":{"execution":{"iopub.status.busy":"2024-09-12T01:54:33.493594Z","iopub.execute_input":"2024-09-12T01:54:33.493996Z","iopub.status.idle":"2024-09-12T01:54:34.788235Z","shell.execute_reply.started":"2024-09-12T01:54:33.493956Z","shell.execute_reply":"2024-09-12T01:54:34.786822Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!find ./vinbigdata-yolo-dataset-with-wbf-640px-1class/val -type f | wc -l","metadata":{"execution":{"iopub.status.busy":"2024-09-12T01:54:34.790128Z","iopub.execute_input":"2024-09-12T01:54:34.790561Z","iopub.status.idle":"2024-09-12T01:54:35.960916Z","shell.execute_reply.started":"2024-09-12T01:54:34.790512Z","shell.execute_reply":"2024-09-12T01:54:35.959269Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!find ./vinbigdata-yolo-dataset-with-wbf-640px-1class/test -type f | wc -l","metadata":{"execution":{"iopub.status.busy":"2024-09-12T01:54:35.963305Z","iopub.execute_input":"2024-09-12T01:54:35.963852Z","iopub.status.idle":"2024-09-12T01:54:37.131759Z","shell.execute_reply.started":"2024-09-12T01:54:35.96379Z","shell.execute_reply":"2024-09-12T01:54:37.130146Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Packaging Dataset into Zip for Upload","metadata":{}},{"cell_type":"code","source":"%%bash\ncd ./vinbigdata-yolo-dataset-with-wbf-640px-1class\nzip -rq ../vinbigdata-yolo-dataset-with-wbf-640px-1class.zip ./*\necho \"Zipping completed successfully.\"","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-09-12T01:54:37.133981Z","iopub.execute_input":"2024-09-12T01:54:37.134473Z","iopub.status.idle":"2024-09-12T01:59:41.978326Z","shell.execute_reply.started":"2024-09-12T01:54:37.134418Z","shell.execute_reply":"2024-09-12T01:59:41.976883Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%bash\nrm -r ./vinbigdata-yolo-dataset-with-wbf-640px-1class\nls -ahl","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-09-12T01:59:41.980582Z","iopub.execute_input":"2024-09-12T01:59:41.980976Z","iopub.status.idle":"2024-09-12T01:59:43.639893Z","shell.execute_reply.started":"2024-09-12T01:59:41.980938Z","shell.execute_reply":"2024-09-12T01:59:43.638529Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style='text-align: center;'><span style=\"color: #0D0D0D; font-family: Segoe UI; font-size: 2.0em; font-weight: 300;\">THANK YOU! PLEASE UPVOTE</span></p>\n\n<p style='text-align: center;'><span style=\"color: #0D0D0D; font-family: Segoe UI; font-size: 2.5em; font-weight: 300;\">HOPE IT WAS USEFUL</span></p>\n\n<p style='text-align: center;'><span style=\"color: #009BAE; font-family: Segoe UI; font-size: 1.2em; font-weight: 300;\">Check out the Train Notebook below</span></p>\n\n\n\n<p style='text-align: center;'><span style=\"font-family: Trebuchet MS; font-size: 1.3em;\"><a href=\"https://www.kaggle.com/code/buithanhxuan/vincxr-yolov8l\" target=\"_top\">vincxr-yolov8l⚡📈</a></span></p>\n\n<p style='text-align: center;'></p>\n","metadata":{}}]}