{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Table of Contents\n1. Introduction\n2. Objectives\n3. Data\n4. Methods(Implementation)\n5. Exploratory Data Analysis\n6. Results\n7. Discussion\n8. Furture Improvement\n9. References","metadata":{}},{"cell_type":"markdown","source":"**1. Introduction**\n\nWhen you have a broken arm, radiologists help save the day—and the bone. These doctors diagnose and treat medical conditions using imaging techniques like CT and PET scans, MRIs, and, of course, X-rays. Yet, as it happens when working with such a wide variety of medical tools, radiologists face many daily challenges, perhaps the most difficult being the chest radiograph. The interpretation of chest X-rays can lead to medical misdiagnosis, even for the best practicing doctor. Computer-aided detection and diagnosis systems (CADe/CADx) would help reduce the pressure on doctors at metropolitan hospitals and improve diagnostic quality in rural areas.\n\nThe annotations were collected via VinBigData's web-based platform, VinLab. Details on building the dataset can be found in the organizer's recent paper “VinDr-CXR: An open dataset of chest X-rays with radiologist's annotations”.und in the organizer's recent paper “VinDr-CXR: An open dataset of chest X-rays with radiologist's annotations”.","metadata":{}},{"cell_type":"markdown","source":"**2. Objective**\n\nExisting methods of interpreting chest X-ray images classify them into a list of findings. There is currently no specification of their locations on the image which sometimes leads to inexplicable results. A solution for localizing findings on chest X-ray images is needed for providing doctors with more meaningful diagnostic assistance\n\nIn this competition, we are classifying common thoracic lung diseases and localizing critical findings. This is an object detection and classification problem from chest x-ray image (radiographs)","metadata":{}},{"cell_type":"markdown","source":"**3. Data**\n\n**3.1 Intro**\n\nIn this competition:\n\nTask: Automatically localize and classify 14 types of thoracic abnormalities from chest radiographs.\nDataset: Consisting of 18,000 scans: 15,000 train images and will be evaluated on a test set of 3,000 images.\nFor each test image, you will be predicting a bounding box and class for all findings. If you predict that there are no findings, you should create a prediction of \"14 1 0 0 1 1\" (14 is the class ID for no finding, and this provides a one-pixel bounding box with a confidence of 1.0).\n\nThe images are in DICOM format, which means they contain additional data that might be useful for visualizing and classifying","metadata":{}},{"cell_type":"markdown","source":"**3.2 Dataset information**\n\nThe dataset comprises 18,000 postero-anterior (PA) CXR scans in DICOM format, which were de-identified to protect patient privacy. All images were labeled by a panel of experienced radiologists for the presence of 14 critical radiographic findings as listed below:\n\nWe consider 14 critical radiographic findings as listed below (click for further informations):\n\n0 - [Aortic enlargement](https://en.wikipedia.org/wiki/Aortic_aneurysm) <br>\n1 - [Atelectasis](https://en.wikipedia.org/wiki/Atelectasis) <br>\n2 - [Calcification](https://en.wikipedia.org/wiki/Calcification) <br>\n3 - [Cardiomegaly](https://en.wikipedia.org/wiki/Cardiomegaly) <br>\n4 - [Consolidation](https://en.wikipedia.org/wiki/Pulmonary_consolidation) <br>\n5 - [ILD](https://en.wikipedia.org/wiki/Interstitial_lung_disease) <br>\n6 - [Infiltration](https://en.wikipedia.org/wiki/Infiltration_(medical)) <br>\n7 - [Lung Opacity](https://en.wikipedia.org/wiki/Ground-glass_opacity) <br>\n8 - [Nodule/Mass](https://en.wikipedia.org/wiki/Lung_nodule) <br>\n9 - Other lesion <br>\n10 - [Pleural effusion](https://en.wikipedia.org/wiki/Pleural_effusion) <br>\n11 - [Pleural thickening](https://en.wikipedia.org/wiki/Pleural_thickening) <br>\n12 - [Pneumothorax](https://en.wikipedia.org/wiki/Pneumothorax) <br>\n13 - [Pulmonary fibrosis](https://en.wikipedia.org/wiki/Pulmonary_fibrosis#:~:text=Pulmonary%20fibrosis%20is%20a%20condition,%2C%20pneumothorax%2C%20and%20lung%20cancer.)\n14 - No Finding\n\nThe \"No finding\" observation (14) was intended to capture the absence of all findings above.","metadata":{}},{"cell_type":"markdown","source":"**4. Methods (Implementation)**","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport matplotlib\nimport pydicom as dicom\nimport cv2\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path = '/kaggle/input/vinbigdata-chest-xray-abnormalities-detection/'\nos.listdir(path)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = pd.read_csv(path+'train.csv')\nsamp_subm = pd.read_csv(path+'sample_submission.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**5. Exploratory Data Analysis**","metadata":{}},{"cell_type":"code","source":"print('Number train samples:', len(train_data.index))\nprint('Number test samples:', len(samp_subm.index))","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 1, figsize=(12, 4))\nx = train_data['class_name'].value_counts().keys()\ny = train_data['class_name'].value_counts().values\nax.bar(x, y)\nax.set_xticklabels(x, rotation=90)\nax.set_title('Distribution of the labels')\nplt.grid()\nplt.show()","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see the dataset is inbalanced.","metadata":{}},{"cell_type":"code","source":"# Read DICOM Files\nidnum = 2\nimage_id = train_data.loc[idnum, 'image_id']\ndata_file = dicom.dcmread(path+'train/'+image_id+'.dicom')\nimg = data_file.pixel_array","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(data_file)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Image shape:', img.shape)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bbox = [train_data.loc[idnum, 'x_min'],\n        train_data.loc[idnum, 'y_min'],\n        train_data.loc[idnum, 'x_max'],\n        train_data.loc[idnum, 'y_max']]\nfig, ax = plt.subplots(1, 1, figsize=(20, 4))\nax.imshow(img, cmap='gray')\np = matplotlib.patches.Rectangle((bbox[0], bbox[1]),\n                                 bbox[2]-bbox[0],\n                                 bbox[3]-bbox[1],\n                                 ec='r', fc='none', lw=2.)\nax.add_patch(p)\nplt.show()","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_train_data(idx_list):\n    fig, axs = plt.subplots(1, 3, figsize=(15, 10))\n    fig.subplots_adjust(hspace = .1, wspace=.1)\n    axs = axs.ravel()\n    for i in range(3):\n        image_id = train_data.loc[idx_list[i], 'image_id']\n        data_file = dicom.dcmread(path+'train/'+image_id+'.dicom')\n        img = data_file.pixel_array\n        axs[i].imshow(img, cmap='gray')\n        axs[i].set_title(train_data.loc[idx_list[i], 'class_name'])\n        axs[i].set_xticklabels([])\n        axs[i].set_yticklabels([])\n        if train_data.loc[idx_list[i], 'class_name'] != 'No finding':\n            bbox = [train_data.loc[idx_list[i], 'x_min'],\n                    train_data.loc[idx_list[i], 'y_min'],\n                    train_data.loc[idx_list[i], 'x_max'],\n                    train_data.loc[idx_list[i], 'y_max']]\n            p = matplotlib.patches.Rectangle((bbox[0], bbox[1]),\n                                             bbox[2]-bbox[0],\n                                             bbox[3]-bbox[1],\n                                             ec='r', fc='none', lw=2.)\n            axs[i].add_patch(p)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for num in range(15):\n    idx_list = train_data[train_data['class_id']==num][0:3].index.values\n    plot_train_data(idx_list)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"6. Results\n","metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0"}},{"cell_type":"code","source":"#samp_subm.to_csv('submission1.csv', index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_2class = pd.read_csv(\"../input/vinbigdata-2class-prediction/2-cls test pred.csv\")\nlow_threshold = 0.001\nhigh_threshold = 0.87\npred_2class","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"NORMAL = \"14 1 0 0 1 1\"\n\npred_det_df = pd.read_csv(\"../input/vinbigdatastack/submission_postprocessed.csv\")\nn_normal_before = len(pred_det_df.query(\"PredictionString == @NORMAL\"))\nmerged_df = pd.merge(pred_det_df, pred_2class, on=\"image_id\", how=\"left\")\n\n\nif \"target\" in merged_df.columns:\n    merged_df[\"class0\"] = 1 - merged_df[\"target\"]\n\nc0, c1, c2 = 0, 0, 0\nfor i in range(len(merged_df)):\n    p0 = merged_df.loc[i, \"class0\"]\n    if p0 < low_threshold:\n\n        c0 += 1\n    elif low_threshold <= p0 and p0 < high_threshold:\n\n        merged_df.loc[i, \"PredictionString\"] += f\" 14 {p0} 0 0 1 1\"\n        c1 += 1\n    else:\n\n        merged_df.loc[i, \"PredictionString\"] = NORMAL\n        c2 += 1","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_normal_after = len(merged_df.query(\"PredictionString == @NORMAL\"))\nprint(\n    f\"n_normal: {n_normal_before} -> {n_normal_after} with threshold {low_threshold} & {high_threshold}\"\n)\nprint(f\"Keep {c0} Add {c1} Replace {c2}\")\nsubmission_filepath = str(\"submission.csv\")\nsubmission_df = merged_df[[\"image_id\", \"PredictionString\"]]\nsubmission_df.to_csv(submission_filepath, index=False)\nprint(f\"Saved to {submission_filepath}\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**7. Discussion**","metadata":{}},{"cell_type":"markdown","source":"**8. Furture Improvement**\n","metadata":{}},{"cell_type":"markdown","source":"**9. References**\n\n1. https://www.kaggle.com/kyawkyaw/vinbigdata-chest-x-ray-abnormalities-classifier","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}}]}