{"cells":[{"metadata":{"papermill":{"duration":0.064694,"end_time":"2021-01-19T21:46:41.128094","exception":false,"start_time":"2021-01-19T21:46:41.0634","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Overview"},{"metadata":{"papermill":{"duration":0.061495,"end_time":"2021-01-19T21:46:41.286004","exception":false,"start_time":"2021-01-19T21:46:41.224509","status":"completed"},"tags":[]},"cell_type":"markdown","source":"* This competition classifies chest radiographs into 15 categories, one of which is normal.\n* This is a competition for object detection and disease classification.\n* All images in dataset are DICOM format. So we need to convert data from DICOM to numpy array.[Convert dicom to np.array - the correct way](https://www.kaggle.com/raddar/convert-dicom-to-np-array-the-correct-way) article will be helpful.\n\n* The host of the competition is explained as follows in [this thread](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/207741).\n* In this competition, you’re given a set of training X-ray images in DICOM format, each of which was blindly annotated with bounding boxes of 14 classes by 3 radiologists from a pool of 17, encoded with Rad IDs from R1 to R17. The task is to automatically correctly predict boxes around abnormalities and classify them for the test images, whose ground-truth labels are hidden. Unlike the labels of the training set, those of the test set were already a consensus of 5 radiologists per image.\n* \n# Columns\n\n\n* image_id - unique image identifier\n\n* class_name - the name of the class of detected object (or \"No finding\")\n\n* class_id - the ID of the class of detected object\n\n* rad_id - the ID of the radiologist that made the observation\n\n* x_min - minimum X coordinate of the object's bounding box\n\n* y_min - minimum Y coordinate of the object's bounding box\n\n* x_max - maximum X coordinate of the object's bounding box\n\n* y_max - maximum Y coordinate of the object's bounding box"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","execution":{"iopub.execute_input":"2021-01-19T21:46:41.422794Z","iopub.status.busy":"2021-01-19T21:46:41.422116Z","iopub.status.idle":"2021-01-19T21:46:50.459765Z","shell.execute_reply":"2021-01-19T21:46:50.459204Z"},"papermill":{"duration":9.111842,"end_time":"2021-01-19T21:46:50.459879","exception":false,"start_time":"2021-01-19T21:46:41.348037","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nimport re\nimport pydicom\n\n# import useful tools\nfrom glob import glob\nfrom PIL import Image\nimport cv2\nimport pydicom as dcm\nimport random\nimport matplotlib.patches as patches\nfrom sklearn.model_selection import KFold\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\n\nfrom sklearn.model_selection import StratifiedKFold\nimport warnings\n\n# import data visualization\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nimport seaborn as sns\nimport matplotlib\nimport pydicom as dicom\n\nfrom bokeh.plotting import figure\nfrom bokeh.io import output_notebook, show, output_file\nfrom bokeh.models import ColumnDataSource, HoverTool, Panel\nfrom bokeh.models.widgets import Tabs\n\n\n# import data augmentation\nimport albumentations as albu\n\n# import math module\nimport math\n\n# Libraries\nimport pandas_profiling\nimport xgboost as xgb\nfrom sklearn.metrics import log_loss\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn import preprocessing\nfrom sklearn.model_selection import KFold\nfrom sklearn.tree import DecisionTreeRegressor\nimport matplotlib.patches as patches\nimport plotly.graph_objects as go\nimport plotly.express as px\nimport plotly.figure_factory as ff\n\n# One-hot encoding\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.ensemble import RandomForestRegressor\n\n# Other\nfrom random import randint\nimport warnings\nimport csv\nwarnings.filterwarnings(\"ignore\")","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.061321,"end_time":"2021-01-19T21:46:50.58441","exception":false,"start_time":"2021-01-19T21:46:50.523089","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Config File"},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:46:50.727731Z","iopub.status.busy":"2021-01-19T21:46:50.725881Z","iopub.status.idle":"2021-01-19T21:46:50.728382Z","shell.execute_reply":"2021-01-19T21:46:50.728802Z"},"papermill":{"duration":0.082708,"end_time":"2021-01-19T21:46:50.728933","exception":false,"start_time":"2021-01-19T21:46:50.646225","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"class Cfg(object):\n  \n  def __init__(self):\n    super(Cfg, self).__init__()\n    self.dim = 512\n    self.batch= 8\n    self.steps= 500\n    self.epochs= 10\n    self.train_csv= '../input/vinbigdata-{}-image-dataset/vinbigdata/train.csv'.format(self.dim)\n    self.test_csv= '../input/vinbigdata-{}-image-dataset/vinbigdata/test.csv'.format(self.dim)\n    self.img_dir= '../input/vinbigdata-{}-image-dataset/vinbigdata/train/'.format(self.dim)\n    \n    self.git= 'https://github.com/fizyr/keras-retinanet.git'\n    self.model_url= ['https://github.com/fizyr/keras-retinanet/releases/download/0.5.1/resnet50_coco_best_v2.1.0.h5',\n               'https://github.com/fizyr/keras-retinanet/releases/download/0.5.1/resnet101_oid_v1.0.0.h5',\n               'https://github.com/fizyr/keras-retinanet/releases/download/0.5.1/resnet152_oid_v1.0.0.h5']\n      \n    self.color_code=   {'Cardiomegaly':(124,252,0), 'Aortic enlargement':(135,206,250),\n                        'Pleural thickening':(199,21,133),'ILD':(245,245,220), 'Nodule/Mass':(220,20,60),\n                        'Pulmonary fibrosis':(0,255,255), 'Lung Opacity':(128,128,0), 'Atelectasis':(255,0,255),\n                        'Other lesion':(176,224,230), 'Infiltration':(210,105,30),'Pleural effusion':(105,105,105),\n                        'Calcification':(138,43,226) ,'Consolidation':(250,240,230),'Pneumothorax':(100,149,237)}\n    \ncfg= Cfg()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.062479,"end_time":"2021-01-19T21:46:50.854696","exception":false,"start_time":"2021-01-19T21:46:50.792217","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Loading data"},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:46:50.998059Z","iopub.status.busy":"2021-01-19T21:46:50.996309Z","iopub.status.idle":"2021-01-19T21:46:50.998674Z","shell.execute_reply":"2021-01-19T21:46:50.999081Z"},"papermill":{"duration":0.080121,"end_time":"2021-01-19T21:46:50.99919","exception":false,"start_time":"2021-01-19T21:46:50.919069","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"# Setup the paths to train and test images\nTEST_DIR = \"../input/vinbigdata-chest-xray-abnormalities-detection/test/\"\nTRAIN_DIR = \"../input/vinbigdata-chest-xray-abnormalities-detection/train/\"\ndataset_dir = \"../input/vinbigdata-chest-xray-abnormalities-detection/\"","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:46:51.128309Z","iopub.status.busy":"2021-01-19T21:46:51.127506Z","iopub.status.idle":"2021-01-19T21:46:51.690511Z","shell.execute_reply":"2021-01-19T21:46:51.689535Z"},"papermill":{"duration":0.629283,"end_time":"2021-01-19T21:46:51.690637","exception":false,"start_time":"2021-01-19T21:46:51.061354","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"# Glob the directories and get the lists of train and test images\ntrain_fns = glob(TRAIN_DIR + '*')\ntest_fns = glob(TEST_DIR + '*')\n\n# Loading training data and test data\ntrain_df = pd.read_csv(dataset_dir+'train.csv')\nsample = pd.read_csv(dataset_dir+'sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:46:51.863661Z","iopub.status.busy":"2021-01-19T21:46:51.863062Z","iopub.status.idle":"2021-01-19T21:46:51.871097Z","shell.execute_reply":"2021-01-19T21:46:51.870663Z"},"papermill":{"duration":0.117763,"end_time":"2021-01-19T21:46:51.871189","exception":false,"start_time":"2021-01-19T21:46:51.753426","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"# Images with no abnormal findings will be omitted.\nabnormal_train = train_df[train_df['class_name']!=\"No finding\"]","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:46:52.016606Z","iopub.status.busy":"2021-01-19T21:46:52.015791Z","iopub.status.idle":"2021-01-19T21:47:55.185643Z","shell.execute_reply":"2021-01-19T21:47:55.184145Z"},"papermill":{"duration":63.249992,"end_time":"2021-01-19T21:47:55.18577","exception":false,"start_time":"2021-01-19T21:46:51.935778","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"from tqdm import tqdm\nrows, columns, sex = [], [], []\nids = abnormal_train['image_id'].unique()\nfor i in ids:\n    path = dataset_dir+ 'train/' + i + '.dicom'\n    dicom = dcm.dcmread(path, stop_before_pixels=True)\n    rows.append(dicom.Rows)\n    columns.append(dicom.Columns)\n    sex.append(dicom.PatientSex)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:55.322687Z","iopub.status.busy":"2021-01-19T21:47:55.320926Z","iopub.status.idle":"2021-01-19T21:47:55.323273Z","shell.execute_reply":"2021-01-19T21:47:55.323688Z"},"papermill":{"duration":0.075106,"end_time":"2021-01-19T21:47:55.323796","exception":false,"start_time":"2021-01-19T21:47:55.24869","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"info = pd.DataFrame({'image_id':ids, 'rows':rows, 'columns':columns, 'sex':sex})","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dicom_meta = pd.read_csv('../input/eda-dicom-reading-vinbigdata-chest-x-ray/train_dicom_properties.csv.bz2').rename(columns={'file':'image_id'})","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:55.457081Z","iopub.status.busy":"2021-01-19T21:47:55.456501Z","iopub.status.idle":"2021-01-19T21:47:55.613432Z","shell.execute_reply":"2021-01-19T21:47:55.613867Z"},"papermill":{"duration":0.227949,"end_time":"2021-01-19T21:47:55.614005","exception":false,"start_time":"2021-01-19T21:47:55.386056","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"df= pd.read_csv(cfg.train_csv)\nprint(df.shape)\ndf= df[df.class_name != 'No finding']","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.062632,"end_time":"2021-01-19T21:47:55.740178","exception":false,"start_time":"2021-01-19T21:47:55.677546","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Useful Functions"},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:55.92505Z","iopub.status.busy":"2021-01-19T21:47:55.924297Z","iopub.status.idle":"2021-01-19T21:47:55.928118Z","shell.execute_reply":"2021-01-19T21:47:55.927659Z"},"papermill":{"duration":0.111072,"end_time":"2021-01-19T21:47:55.928211","exception":false,"start_time":"2021-01-19T21:47:55.817139","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"def get_feature_distribution(data, feature):\n    # Get the count for each label\n    label_counts = data[feature].value_counts()\n\n    # Get total number of samples\n    total_samples = len(data)\n\n    # Count the number of items in each class\n    print(\"Feature: {}\".format(feature))\n    for i in range(len(label_counts)):\n        label = label_counts.index[i]\n        count = label_counts.values[i]\n        percent = int((count / total_samples) * 10000) / 100\n        print(\"{:<30s}:{}%\".format(label, count, percent))","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:56.061916Z","iopub.status.busy":"2021-01-19T21:47:56.061152Z","iopub.status.idle":"2021-01-19T21:47:56.064024Z","shell.execute_reply":"2021-01-19T21:47:56.063602Z"},"papermill":{"duration":0.072405,"end_time":"2021-01-19T21:47:56.064111","exception":false,"start_time":"2021-01-19T21:47:55.991706","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"def show_dicom_images(data):\n    img_data = data\n    f, ax = plt.subplots(3,3, figsize=(6,8))\n    for i,data_row in enumerate(img_data):\n        imagePath = data_row\n        data_row_img_data = dcm.read_file(imagePath)\n        data_row_img = dcm.dcmread(imagePath)\n        ax[i//3, i%3].imshow(data_row_img.pixel_array, cmap=plt.cm.bone) \n        ax[i//3, i%3].axis('off')\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:56.20886Z","iopub.status.busy":"2021-01-19T21:47:56.207942Z","iopub.status.idle":"2021-01-19T21:47:56.217323Z","shell.execute_reply":"2021-01-19T21:47:56.216888Z"},"papermill":{"duration":0.079581,"end_time":"2021-01-19T21:47:56.217423","exception":false,"start_time":"2021-01-19T21:47:56.137842","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"def load_img(path):\n    img= cv2.imread(path)\n    img= cv2.resize(img, (cfg.dim, cfg.dim))\n    return img\n\ndef normalize_cod(df):\n    df.x_min= (df.x_min/ df.width)* cfg.dim\n    df.x_max= (df.x_max/ df.width)* cfg.dim\n    \n    df.y_min= (df.y_min/ df.height)* cfg.dim\n    df.y_max= (df.y_max/ df.height)* cfg.dim\n    return df\n\ndf= normalize_cod(df.copy())\ndf= df.reset_index(drop = True)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:56.352412Z","iopub.status.busy":"2021-01-19T21:47:56.350549Z","iopub.status.idle":"2021-01-19T21:47:56.352995Z","shell.execute_reply":"2021-01-19T21:47:56.353419Z"},"papermill":{"duration":0.073077,"end_time":"2021-01-19T21:47:56.353532","exception":false,"start_time":"2021-01-19T21:47:56.280455","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"def read_xray(path, voi_lut = True, fix_monochrome = True):\n    dicom = dcm.read_file(path)\n\n    if voi_lut:\n        data = apply_voi_lut(dicom.pixel_array, dicom)\n    else:\n        data = dicom.pixel_array\n\n    if fix_monochrome and dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n\n    data = data - np.min(data)\n    data = data / np.max(data)\n    return (data * 255).astype(np.uint8)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:56.489069Z","iopub.status.busy":"2021-01-19T21:47:56.488542Z","iopub.status.idle":"2021-01-19T21:47:56.492657Z","shell.execute_reply":"2021-01-19T21:47:56.492175Z"},"papermill":{"duration":0.075679,"end_time":"2021-01-19T21:47:56.492766","exception":false,"start_time":"2021-01-19T21:47:56.417087","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"def plot(name):\n    train_cls = train_df[train_df['class_name'] == name]\n    fig, axes = plt.subplots(4,4, figsize=(10, 10))\n    fig.suptitle(name+\" examples\", fontsize=10)\n    for i in range(4):\n        for j in range(4):\n            row = train_cls.iloc[random.randint(0, len(train_cls))]\n            path = dataset_dir + 'train/' + row['image_id'] + '.dicom'\n            axes[i][j].imshow(read_xray(path), cmap='gray')\n            axes[i][j].add_patch(patches.Rectangle(\n                (row['x_min'], row['y_min']), \n                row['x_max'] - row['x_min'], \n                row['y_max'] - row['y_min'], \n                edgecolor='blue', \n                fill=False)\n            )\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:56.628223Z","iopub.status.busy":"2021-01-19T21:47:56.627302Z","iopub.status.idle":"2021-01-19T21:47:56.635142Z","shell.execute_reply":"2021-01-19T21:47:56.634686Z"},"papermill":{"duration":0.079118,"end_time":"2021-01-19T21:47:56.635235","exception":false,"start_time":"2021-01-19T21:47:56.556117","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"def load_img(path):\n    img= cv2.imread(path)\n    img= cv2.resize(img, (cfg.dim, cfg.dim))\n    return img\n\ndef normalize_cod(df):\n    df.x_min= (df.x_min/ df.width)* cfg.dim\n    df.x_max= (df.x_max/ df.width)* cfg.dim\n    \n    df.y_min= (df.y_min/ df.height)* cfg.dim\n    df.y_max= (df.y_max/ df.height)* cfg.dim\n    return df\n\ndf= normalize_cod(df.copy())\ndf= df.reset_index(drop = True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_width_of__bounding_boxes(data):\n    \n    fig, axes = plt.subplots(7, 2, figsize=(16,20), sharex=True)\n    fig.suptitle(\"width of bounding box for different categories\", fontsize=16)\n    \n    classes = data.class_name.unique()\n    for j, i in enumerate(classes[~np.isin(classes, 'No finding')]):\n        data_ = data[data['class_name']==i]\n        sns.distplot(data_['x_max'] - data_['x_min'], ax=axes[j%7, j//7]);\n        axes[j%7, j//7].title.set_text(i);\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:56.766537Z","iopub.status.busy":"2021-01-19T21:47:56.76578Z","iopub.status.idle":"2021-01-19T21:47:56.770146Z","shell.execute_reply":"2021-01-19T21:47:56.770548Z"},"papermill":{"duration":0.073032,"end_time":"2021-01-19T21:47:56.770658","exception":false,"start_time":"2021-01-19T21:47:56.697626","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"classes= df.class_name.unique()\nind= df.class_id.unique()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:56.912643Z","iopub.status.busy":"2021-01-19T21:47:56.910945Z","iopub.status.idle":"2021-01-19T21:47:56.91338Z","shell.execute_reply":"2021-01-19T21:47:56.913799Z"},"papermill":{"duration":0.0785,"end_time":"2021-01-19T21:47:56.913907","exception":false,"start_time":"2021-01-19T21:47:56.835407","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"##### Color-Code ########\ncolor= cfg.color_code\n\ndef show_bb(i):\n    df_mini= df[df.image_id==df.image_id[i]]\n    path= cfg.img_dir + df.image_id[i] + '.png'\n    img= load_img(path)\n    rep_class=[]\n    font = cv2.FONT_HERSHEY_SIMPLEX \n    for i, row in df_mini.iterrows():\n        class_n= row['class_name']\n        if class_n in rep_class:\n            continue                          # More generalization\n        rep_class.append(class_n)\n        x_min= int(row['x_min']); x_max= int(row['x_max'])\n        y_min= int(row['y_min']); y_max= int(row['y_max'])\n        img= cv2.rectangle(img, (x_min, y_min), (x_max, y_max), color[class_n], 2)\n        fontScale= (x_max- x_min)*2.5/img.shape[1]\n        img= cv2.putText(img, class_n, (x_min, y_min-5), font, fontScale, cv2.LINE_AA)\n    return img","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:57.171788Z","iopub.status.busy":"2021-01-19T21:47:57.171002Z","iopub.status.idle":"2021-01-19T21:47:57.175889Z","shell.execute_reply":"2021-01-19T21:47:57.17548Z"},"papermill":{"duration":0.074997,"end_time":"2021-01-19T21:47:57.175982","exception":false,"start_time":"2021-01-19T21:47:57.100985","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"classes= df.class_name.unique()\nind= df.class_id.unique()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:57.312608Z","iopub.status.busy":"2021-01-19T21:47:57.311836Z","iopub.status.idle":"2021-01-19T21:47:57.315362Z","shell.execute_reply":"2021-01-19T21:47:57.315758Z"},"papermill":{"duration":0.076859,"end_time":"2021-01-19T21:47:57.315879","exception":false,"start_time":"2021-01-19T21:47:57.23902","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"##### Color-Code ########\ncolor= cfg.color_code\n\ndef show_bb(i):\n    df_mini= df[df.image_id==df.image_id[i]]\n    path= cfg.img_dir + df.image_id[i] + '.png'\n    img= load_img(path)\n    rep_class=[]\n    font = cv2.FONT_HERSHEY_SIMPLEX \n    for i, row in df_mini.iterrows():\n        class_n= row['class_name']\n        if class_n in rep_class:\n            continue                          # More generalization\n        rep_class.append(class_n)\n        x_min= int(row['x_min']); x_max= int(row['x_max'])\n        y_min= int(row['y_min']); y_max= int(row['y_max'])\n        img= cv2.rectangle(img, (x_min, y_min), (x_max, y_max), color[class_n], 2)\n        fontScale= (x_max- x_min)*2.5/img.shape[1]\n        img= cv2.putText(img, class_n, (x_min, y_min-5), font, fontScale, cv2.LINE_AA)\n    return img","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.129295,"end_time":"2021-01-19T21:47:58.685846","exception":false,"start_time":"2021-01-19T21:47:58.556551","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Check statistics"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.nunique().to_frame().rename(columns={0:\"Unique Values\"}).style.background_gradient(cmap=\"plasma\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df[['class_id', 'class_name', 'rad_id']].groupby(['class_id', 'class_name']).count().rename(columns={'rad_id': 'Number of records'})","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:58.857987Z","iopub.status.busy":"2021-01-19T21:47:58.857411Z","iopub.status.idle":"2021-01-19T21:47:58.868526Z","shell.execute_reply":"2021-01-19T21:47:58.868032Z"},"papermill":{"duration":0.102225,"end_time":"2021-01-19T21:47:58.868623","exception":false,"start_time":"2021-01-19T21:47:58.766398","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"df.head()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:59.055496Z","iopub.status.busy":"2021-01-19T21:47:59.054679Z","iopub.status.idle":"2021-01-19T21:47:59.16407Z","shell.execute_reply":"2021-01-19T21:47:59.164521Z"},"papermill":{"duration":0.213286,"end_time":"2021-01-19T21:47:59.164647","exception":false,"start_time":"2021-01-19T21:47:58.951361","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"train_rad = train_df['rad_id'].value_counts().reset_index()\nfig = go.Figure(data=[go.Table(header=dict(values=['Radiologist ID', 'Number of Observations'], fill_color='yellow'),\n                 cells=dict(values=[train_rad['index'], train_rad['rad_id']], fill_color='lavender'))\n                     ])\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_distribution_classes(x_values, y_values, title):\n    \n    #colors = ['rgb(26, 118, 255)',] * 15\n    #colors[0] = 'lightslategray'\n\n    fig = go.Figure(data=[go.Bar(\n        x=x_values, \n        y=y_values,\n        text=y_values\n        #marker_color=colors\n    )])\n\n    fig.update_layout(height=400, width=700, title_text=title)\n    fig.update_xaxes(type=\"category\")\n\n    fig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"indexes = train_df.rad_id.unique()\ncounts = train_df.rad_id.value_counts()\n\nsorted_dict = dict(zip(indexes, counts))\nsorted_dict = {k: v for k, v in sorted(sorted_dict.items(), key=lambda item: item[1], reverse = True)}\n\nx = list(sorted_dict.keys())\ny = list(sorted_dict.values())\n\nplot_distribution_classes(x, y, \n                          title=\"Distribution of Annotations by Radioloiest\")","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.081353,"end_time":"2021-01-19T21:47:59.328631","exception":false,"start_time":"2021-01-19T21:47:59.247278","status":"completed"},"tags":[]},"cell_type":"markdown","source":"* Rad IDs from R1 to R17\n* rad_id is the ID of the radiologist that made the observation"},{"metadata":{"trusted":true},"cell_type":"code","source":"whatassigned = train_df[['rad_id', 'class_id', 'image_id']]\\\n    .groupby(['rad_id', 'class_id'])\\\n    .count()\\\n    .reset_index()\\\n    .pivot(index='rad_id', columns='class_id',values='image_id')\\\n    .add_prefix('class')\\\n    .fillna(0)\\\n    .astype(np.int64)\nwhatassigned['Percent with no finding'] = [f'{tmpvar}%' for tmpvar in np.round(100*whatassigned['class14'].values/whatassigned.sum(axis=1).values,2)]\nwhatassigned","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(6, 6))\nsns.countplot(x=\"class_id\", data=train_df)\nplt.title(\"Class ID Distribution\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(6, 6))\nsns.countplot(x=\"rad_id\", data=train_df)\nplt.title(\"RAD ID Distribution\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, axes = plt.subplots(figsize = (20,10), nrows=2, ncols=2)\nax0, ax1, ax2, ax3 = axes.flatten()\n\nax0.hist(train_df.x_min, bins=55, color = \"skyblue\")\nax0.axvline(x=602, color='royalblue', linestyle='dashed', linewidth=2)\nax0.axvline(x=1457, color='royalblue', linestyle='dashed', linewidth=2)\nax0.axvline(x=1014, color='cornflowerblue', linewidth=2)\nax0.set_title('X minimum', fontsize=18)\n\nax1.hist(train_df.x_max, bins=55, color = \"skyblue\")\nax1.axvline(x=1010, color='royalblue', linestyle='dashed', linewidth=2)\nax1.axvline(x=1567, color='cornflowerblue', linewidth=2)\nax1.axvline(x=1947, color='royalblue', linestyle='dashed', linewidth=2)\nax1.set_title('X maximum', fontsize=18)\n\nax2.hist(train_df.y_min, bins=55)\nax2.axvline(x=627, color='orchid', linestyle='dashed', linewidth=2)\nax2.axvline(x=935, color='cornflowerblue', linewidth=2)\nax2.axvline(x=1471, color='orchid', linestyle='dashed', linewidth=2)\nax2.set_title('Y minimum', fontsize=18)\n\nax3.hist(train_df.y_max, bins=55)\nax3.axvline(x=1009, color='orchid', linestyle='dashed', linewidth=2)\nax3.axvline(x=1411, color='cornflowerblue', linewidth=2)\nax3.axvline(x=1911, color='orchid', linestyle='dashed', linewidth=2)\nax3.set_title('Y maximum', fontsize=18)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Submission File\n[The competition page](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/overview/evaluation) has the following to say about the sample.\nImages in the test set may contain more than one object. For each object in a given test image, you must predict a class ID, confidence score, and bounding box in format xmin ymin xmax ymax. If you predict that there are NO objects in a given image, you should predict 14 1.0 0 0 1 1, where 14 is the class ID for \"No finding\", 1.0 is the confidence, and 0 0 1 1 is a one-pixel bounding box.\n\nThe submission file should contain a header and have the following format:\n\n"},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:59.50323Z","iopub.status.busy":"2021-01-19T21:47:59.50268Z","iopub.status.idle":"2021-01-19T21:47:59.509135Z","shell.execute_reply":"2021-01-19T21:47:59.508712Z"},"papermill":{"duration":0.09869,"end_time":"2021-01-19T21:47:59.509231","exception":false,"start_time":"2021-01-19T21:47:59.410541","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"# Confirmation of the format of samples for submission\nsample.head(3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Check the structure of the training data: there are three types of IDs and disease categories, and the maximum and minimum values for x and y, respectively, are listed."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:59.688234Z","iopub.status.busy":"2021-01-19T21:47:59.687408Z","iopub.status.idle":"2021-01-19T21:47:59.690969Z","shell.execute_reply":"2021-01-19T21:47:59.69141Z"},"papermill":{"duration":0.100939,"end_time":"2021-01-19T21:47:59.691526","exception":false,"start_time":"2021-01-19T21:47:59.590587","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"# Display some of the training data\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:47:59.87019Z","iopub.status.busy":"2021-01-19T21:47:59.869532Z","iopub.status.idle":"2021-01-19T21:47:59.878347Z","shell.execute_reply":"2021-01-19T21:47:59.878741Z"},"papermill":{"duration":0.104956,"end_time":"2021-01-19T21:47:59.878862","exception":false,"start_time":"2021-01-19T21:47:59.773906","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"# Display of training data\nprint(train_df)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:48:00.050762Z","iopub.status.busy":"2021-01-19T21:48:00.049978Z","iopub.status.idle":"2021-01-19T21:48:00.075866Z","shell.execute_reply":"2021-01-19T21:48:00.075339Z"},"papermill":{"duration":0.11344,"end_time":"2021-01-19T21:48:00.075958","exception":false,"start_time":"2021-01-19T21:47:59.962518","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"# Check for missing values in the training data\ntrain_df.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:48:00.257512Z","iopub.status.busy":"2021-01-19T21:48:00.256617Z","iopub.status.idle":"2021-01-19T21:48:00.263488Z","shell.execute_reply":"2021-01-19T21:48:00.264004Z"},"papermill":{"duration":0.105369,"end_time":"2021-01-19T21:48:00.264148","exception":false,"start_time":"2021-01-19T21:48:00.158779","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"# Check the unique values of image IDs.\nabnormal_train.image_id.value_counts()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Let's look at the number of diseases in the training data in a graph."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:48:00.442993Z","iopub.status.busy":"2021-01-19T21:48:00.442223Z","iopub.status.idle":"2021-01-19T21:48:00.718181Z","shell.execute_reply":"2021-01-19T21:48:00.718669Z"},"papermill":{"duration":0.37141,"end_time":"2021-01-19T21:48:00.718794","exception":false,"start_time":"2021-01-19T21:48:00.347384","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(6,6))\nsns.countplot(y ='class_name', data=abnormal_train);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Next, let's check the number of diseases in the training data with numbers."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:48:01.088455Z","iopub.status.busy":"2021-01-19T21:48:01.087345Z","iopub.status.idle":"2021-01-19T21:48:01.185953Z","shell.execute_reply":"2021-01-19T21:48:01.186665Z"},"papermill":{"duration":0.189374,"end_time":"2021-01-19T21:48:01.186806","exception":false,"start_time":"2021-01-19T21:48:00.997432","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"train_df[['class_id', 'class_name', 'rad_id']].groupby(['class_id', 'class_name']).count().rename(columns={'rad_id': 'Number of records'}).style.applymap(lambda x: 'background-color:lightsteelblue')","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.085185,"end_time":"2021-01-19T21:48:01.357612","exception":false,"start_time":"2021-01-19T21:48:01.272427","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# No finding"},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:48:01.542823Z","iopub.status.busy":"2021-01-19T21:48:01.541897Z","iopub.status.idle":"2021-01-19T21:48:32.826665Z","shell.execute_reply":"2021-01-19T21:48:32.827145Z"},"papermill":{"duration":31.384659,"end_time":"2021-01-19T21:48:32.827289","exception":false,"start_time":"2021-01-19T21:48:01.44263","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot(\"No finding\")","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.091857,"end_time":"2021-01-19T21:48:33.010887","exception":false,"start_time":"2021-01-19T21:48:32.91903","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Aortic enlargement\n* Aortic enlargement is known as a sign of an aortic aneurysm. This condition often occurs in the ascending aorta."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:48:33.205549Z","iopub.status.busy":"2021-01-19T21:48:33.204739Z","iopub.status.idle":"2021-01-19T21:48:59.553358Z","shell.execute_reply":"2021-01-19T21:48:59.55392Z"},"papermill":{"duration":26.452635,"end_time":"2021-01-19T21:48:59.55408","exception":false,"start_time":"2021-01-19T21:48:33.101445","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot(\"Aortic enlargement\")","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:48:59.760565Z","iopub.status.busy":"2021-01-19T21:48:59.758548Z","iopub.status.idle":"2021-01-19T21:49:00.856661Z","shell.execute_reply":"2021-01-19T21:49:00.856144Z"},"papermill":{"duration":1.203734,"end_time":"2021-01-19T21:49:00.856765","exception":false,"start_time":"2021-01-19T21:48:59.653031","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target1 = abnormal_train[abnormal_train['class_id']==0]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target1['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target1['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target1['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target1['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.098386,"end_time":"2021-01-19T21:49:01.053617","exception":false,"start_time":"2021-01-19T21:49:00.955231","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Cardiomegaly"},{"metadata":{},"cell_type":"markdown","source":"*  Cardiomegaly can be caused by many conditions, including hypertension, coronary artery disease, infections, inherited disorders, and cardiomyopathies.\n* Cardiomegaly is usually diagnosed when the ratio of the heart's width to the width of the chest is more than 50%. This diagnostic criterion may be an essential basis for this competition.\n* Cardiomegaly can be caused by many conditions, including hypertension, coronary artery disease, infections, inherited disorders, and cardiomyopathies.\n* The heart-to-lung ratio criterion for the diagnosis of cardiomegaly is a ratio of greater than 0.5. "},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:49:01.262517Z","iopub.status.busy":"2021-01-19T21:49:01.261707Z","iopub.status.idle":"2021-01-19T21:49:20.636431Z","shell.execute_reply":"2021-01-19T21:49:20.637403Z"},"papermill":{"duration":19.486187,"end_time":"2021-01-19T21:49:20.637611","exception":false,"start_time":"2021-01-19T21:49:01.151424","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot(\"Cardiomegaly\")","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:49:21.231515Z","iopub.status.busy":"2021-01-19T21:49:21.230587Z","iopub.status.idle":"2021-01-19T21:49:22.51123Z","shell.execute_reply":"2021-01-19T21:49:22.511697Z"},"papermill":{"duration":1.398318,"end_time":"2021-01-19T21:49:22.511821","exception":false,"start_time":"2021-01-19T21:49:21.113503","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target２ = abnormal_train[abnormal_train['class_id']==3]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target２['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target２['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target２['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target２['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.10503,"end_time":"2021-01-19T21:49:22.721293","exception":false,"start_time":"2021-01-19T21:49:22.616263","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Pleural thickening"},{"metadata":{},"cell_type":"markdown","source":"* The pleura is the membrane that covers the lungs, and the change in the thickness of the pleura is called pleural thickening.\n* It is often seen in the uppermost part of the lung field (the apex of the lung)."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:49:22.944297Z","iopub.status.busy":"2021-01-19T21:49:22.943466Z","iopub.status.idle":"2021-01-19T21:49:47.108974Z","shell.execute_reply":"2021-01-19T21:49:47.109438Z"},"papermill":{"duration":24.284224,"end_time":"2021-01-19T21:49:47.109584","exception":false,"start_time":"2021-01-19T21:49:22.82536","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot(\"Pleural thickening\")","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:49:47.346506Z","iopub.status.busy":"2021-01-19T21:49:47.345578Z","iopub.status.idle":"2021-01-19T21:49:48.394841Z","shell.execute_reply":"2021-01-19T21:49:48.394402Z"},"papermill":{"duration":1.171638,"end_time":"2021-01-19T21:49:48.394941","exception":false,"start_time":"2021-01-19T21:49:47.223303","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target３ = abnormal_train[abnormal_train['class_id']==11]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target３['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target３['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target３['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target３['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.11238,"end_time":"2021-01-19T21:49:48.619358","exception":false,"start_time":"2021-01-19T21:49:48.506978","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Pulmonary fibrosis"},{"metadata":{},"cell_type":"markdown","source":"* Pulmonary Fibrosis is inflammation of the lung interstitium due to various causes, resulting in thickening and hardening of the walls, fibrosis, and scarring.\n* The fibrotic areas lose their air content, which often results in dense cord shadows or granular shadows."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:49:49.005375Z","iopub.status.busy":"2021-01-19T21:49:49.004521Z","iopub.status.idle":"2021-01-19T21:50:15.560791Z","shell.execute_reply":"2021-01-19T21:50:15.562601Z"},"papermill":{"duration":26.828247,"end_time":"2021-01-19T21:50:15.562788","exception":false,"start_time":"2021-01-19T21:49:48.734541","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot(\"Pulmonary fibrosis\")","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:50:15.987376Z","iopub.status.busy":"2021-01-19T21:50:15.986421Z","iopub.status.idle":"2021-01-19T21:50:17.283125Z","shell.execute_reply":"2021-01-19T21:50:17.283576Z"},"papermill":{"duration":1.49711,"end_time":"2021-01-19T21:50:17.283703","exception":false,"start_time":"2021-01-19T21:50:15.786593","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target４ = abnormal_train[abnormal_train['class_id']==13]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target４['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target４['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target４['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target４['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.119188,"end_time":"2021-01-19T21:50:17.521793","exception":false,"start_time":"2021-01-19T21:50:17.402605","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Nodule/Mass"},{"metadata":{},"cell_type":"markdown","source":"* Nodules and masses are seen primarily in lung cancer, and metastasis from other parts of the body such as colon cancer and kidney cancer, tuberculosis, pulmonary mycosis, non-tuberculous mycobacterium, obsolete pneumonia, and benign tumors.\n* A nodule/mass is a round shade (typically less than 3 cm in diameter – resulting in much smaller than average bounding boxes) that appears on a chest X-ray image."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:50:17.779006Z","iopub.status.busy":"2021-01-19T21:50:17.777764Z","iopub.status.idle":"2021-01-19T21:50:43.501237Z","shell.execute_reply":"2021-01-19T21:50:43.501683Z"},"papermill":{"duration":25.860686,"end_time":"2021-01-19T21:50:43.501808","exception":false,"start_time":"2021-01-19T21:50:17.641122","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot(\"Nodule/Mass\")","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:50:43.767384Z","iopub.status.busy":"2021-01-19T21:50:43.766492Z","iopub.status.idle":"2021-01-19T21:50:44.840078Z","shell.execute_reply":"2021-01-19T21:50:44.839136Z"},"papermill":{"duration":1.209868,"end_time":"2021-01-19T21:50:44.840183","exception":false,"start_time":"2021-01-19T21:50:43.630315","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target５ = abnormal_train[abnormal_train['class_id']==8]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target５['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target５['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target５['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target５['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.127129,"end_time":"2021-01-19T21:50:45.0951","exception":false,"start_time":"2021-01-19T21:50:44.967971","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Lung Opacity"},{"metadata":{},"cell_type":"markdown","source":"* Lung opacity can often be identified as any area in the chest radiograph that is more white than it should be.\n* Please see [this kaggle discussion](https://www.kaggle.com/zahaviguy/what-are-lung-opacities) for more information."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:50:45.364123Z","iopub.status.busy":"2021-01-19T21:50:45.363338Z","iopub.status.idle":"2021-01-19T21:51:18.309089Z","shell.execute_reply":"2021-01-19T21:51:18.309593Z"},"papermill":{"duration":33.088432,"end_time":"2021-01-19T21:51:18.309741","exception":false,"start_time":"2021-01-19T21:50:45.221309","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot(\"Lung Opacity\")","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.132206,"end_time":"2021-01-19T21:51:18.574","exception":false,"start_time":"2021-01-19T21:51:18.441794","status":"completed"},"tags":[]},"cell_type":"markdown","source":"* Please see [What are lung opacities?](https://www.kaggle.com/zahaviguy/what-are-lung-opacities)."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:51:18.849884Z","iopub.status.busy":"2021-01-19T21:51:18.849035Z","iopub.status.idle":"2021-01-19T21:51:19.918933Z","shell.execute_reply":"2021-01-19T21:51:19.919373Z"},"papermill":{"duration":1.213013,"end_time":"2021-01-19T21:51:19.919511","exception":false,"start_time":"2021-01-19T21:51:18.706498","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target６ = abnormal_train[abnormal_train['class_id']==7]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target６['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target６['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target６['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target６['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.133667,"end_time":"2021-01-19T21:51:20.18819","exception":false,"start_time":"2021-01-19T21:51:20.054523","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Pleural effusion"},{"metadata":{},"cell_type":"markdown","source":"* Pleural effusion is the accumulation of water outside the lungs in the chest cavity.\n* The outside of the lungs is covered by a thin membrane consisting of two layers known as the pleura. Fluid accumulation between these two layers (chest-wall/parietal-pleura and the lung-tissue/visceral-pleura) is called pleural effusion.\n* The findings of pleural effusion vary widely and vary depending on whether the radiograph is taken in the upright or supine position.\n* The most common presentation of pleural effusion is elevation of the diaphragm on one side, flattening the diaphragm, or blunting the angle between rib and diaphragm (typically more than 30 degrees)"},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:51:20.473006Z","iopub.status.busy":"2021-01-19T21:51:20.472005Z","iopub.status.idle":"2021-01-19T21:51:46.119459Z","shell.execute_reply":"2021-01-19T21:51:46.119892Z"},"papermill":{"duration":25.796111,"end_time":"2021-01-19T21:51:46.120013","exception":false,"start_time":"2021-01-19T21:51:20.323902","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot(\"Pleural effusion\")","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:51:46.414615Z","iopub.status.busy":"2021-01-19T21:51:46.413754Z","iopub.status.idle":"2021-01-19T21:51:47.44003Z","shell.execute_reply":"2021-01-19T21:51:47.440467Z"},"papermill":{"duration":1.177627,"end_time":"2021-01-19T21:51:47.440603","exception":false,"start_time":"2021-01-19T21:51:46.262976","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target７ = abnormal_train[abnormal_train['class_id']==10]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target７['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target７['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target７['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target７['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.140148,"end_time":"2021-01-19T21:51:47.721564","exception":false,"start_time":"2021-01-19T21:51:47.581416","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Other lesion"},{"metadata":{},"cell_type":"markdown","source":"* Others include all abnormalities that do not fall into any other category. This includes bone penetrating images, fractures, subcutaneous emphysema, etc."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:51:48.018891Z","iopub.status.busy":"2021-01-19T21:51:48.018044Z","iopub.status.idle":"2021-01-19T21:52:19.06957Z","shell.execute_reply":"2021-01-19T21:52:19.070042Z"},"papermill":{"duration":31.208024,"end_time":"2021-01-19T21:52:19.070179","exception":false,"start_time":"2021-01-19T21:51:47.862155","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot(\"Other lesion\")","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:52:19.375381Z","iopub.status.busy":"2021-01-19T21:52:19.374477Z","iopub.status.idle":"2021-01-19T21:52:20.421763Z","shell.execute_reply":"2021-01-19T21:52:20.422153Z"},"papermill":{"duration":1.205104,"end_time":"2021-01-19T21:52:20.422311","exception":false,"start_time":"2021-01-19T21:52:19.217207","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target８ = abnormal_train[abnormal_train['class_id']==9]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target８['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target８['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target８['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target８['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.147633,"end_time":"2021-01-19T21:52:20.716869","exception":false,"start_time":"2021-01-19T21:52:20.569236","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Infiltration"},{"metadata":{},"cell_type":"markdown","source":"* The infiltration of some fluid component into the alveoli causes an infiltrative shadow (Infiltration).\n* It is difficult to distinguish from consolidation and, in some cases, impossible to distinguish. Please see [this link](https://allnurses.com/consolidation-vs-infiltrate-vs-opacity-t483538/) for more information."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:52:21.032991Z","iopub.status.busy":"2021-01-19T21:52:21.032136Z","iopub.status.idle":"2021-01-19T21:52:41.800234Z","shell.execute_reply":"2021-01-19T21:52:41.801767Z"},"papermill":{"duration":20.936976,"end_time":"2021-01-19T21:52:41.80195","exception":false,"start_time":"2021-01-19T21:52:20.864974","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot(\"Infiltration\")","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:52:42.170474Z","iopub.status.busy":"2021-01-19T21:52:42.16957Z","iopub.status.idle":"2021-01-19T21:52:43.23565Z","shell.execute_reply":"2021-01-19T21:52:43.236068Z"},"papermill":{"duration":1.233645,"end_time":"2021-01-19T21:52:43.236188","exception":false,"start_time":"2021-01-19T21:52:42.002543","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target９ = abnormal_train[abnormal_train['class_id']==6]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target９['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target９['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target９['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target９['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.157314,"end_time":"2021-01-19T21:52:43.55537","exception":false,"start_time":"2021-01-19T21:52:43.398056","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# ILD  "},{"metadata":{},"cell_type":"markdown","source":"* ILD stands for \"Interstitial Lung Disease\".\n* Interstitial Lung Disease is a general term for many conditions in which the interstitial space is injured."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:52:43.886541Z","iopub.status.busy":"2021-01-19T21:52:43.875217Z","iopub.status.idle":"2021-01-19T21:53:15.788716Z","shell.execute_reply":"2021-01-19T21:53:15.789276Z"},"papermill":{"duration":32.075781,"end_time":"2021-01-19T21:53:15.789432","exception":false,"start_time":"2021-01-19T21:52:43.713651","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot(\"ILD\")","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.170694,"end_time":"2021-01-19T21:53:16.120721","exception":false,"start_time":"2021-01-19T21:53:15.950027","status":"completed"},"tags":[]},"cell_type":"markdown","source":"* ILD stands for \"Interstitial Lung Disease.\"\n* Interstitial lung disease is a general term for many conditions in which the interstitial space is injured."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:53:16.453536Z","iopub.status.busy":"2021-01-19T21:53:16.452706Z","iopub.status.idle":"2021-01-19T21:53:17.503869Z","shell.execute_reply":"2021-01-19T21:53:17.503412Z"},"papermill":{"duration":1.221861,"end_time":"2021-01-19T21:53:17.503976","exception":false,"start_time":"2021-01-19T21:53:16.282115","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target10 = abnormal_train[abnormal_train['class_id']==5]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target10['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target10['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target10['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target10['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.161204,"end_time":"2021-01-19T21:53:17.831534","exception":false,"start_time":"2021-01-19T21:53:17.67033","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Calcification"},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:53:18.175201Z","iopub.status.busy":"2021-01-19T21:53:18.17397Z","iopub.status.idle":"2021-01-19T21:53:41.449072Z","shell.execute_reply":"2021-01-19T21:53:41.449517Z"},"papermill":{"duration":23.454726,"end_time":"2021-01-19T21:53:41.449643","exception":false,"start_time":"2021-01-19T21:53:17.994917","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot(\"Calcification\")","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.16817,"end_time":"2021-01-19T21:53:41.790357","exception":false,"start_time":"2021-01-19T21:53:41.622187","status":"completed"},"tags":[]},"cell_type":"markdown","source":"* Calcium (calcification) may be deposited in areas where previous inflammation of the lungs or pleura has healed. Calcium may be deposited in the aorta due to atherosclerosis. Or calcification may occur in mediastinal lymph nodes.\n* Many diseases or conditions can cause calcification on chest x-ray.\n* Calcification may occur in the Aorta (as with atherosclerosis) or it may occur in mediastinal lymph nodes (as with previous infection, tuberculosis, or histoplasmosis).\n"},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:53:42.147078Z","iopub.status.busy":"2021-01-19T21:53:42.140305Z","iopub.status.idle":"2021-01-19T21:53:43.189183Z","shell.execute_reply":"2021-01-19T21:53:43.189869Z"},"papermill":{"duration":1.230094,"end_time":"2021-01-19T21:53:43.190037","exception":false,"start_time":"2021-01-19T21:53:41.959943","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target11 = abnormal_train[abnormal_train['class_id']==2]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target11['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target11['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target11['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target11['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.170866,"end_time":"2021-01-19T21:53:43.552957","exception":false,"start_time":"2021-01-19T21:53:43.382091","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Consolidation"},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:53:43.918933Z","iopub.status.busy":"2021-01-19T21:53:43.911606Z","iopub.status.idle":"2021-01-19T21:54:14.721079Z","shell.execute_reply":"2021-01-19T21:54:14.721659Z"},"papermill":{"duration":30.995859,"end_time":"2021-01-19T21:54:14.721795","exception":false,"start_time":"2021-01-19T21:53:43.725936","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot(\"Consolidation\")","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.174766,"end_time":"2021-01-19T21:54:15.088567","exception":false,"start_time":"2021-01-19T21:54:14.913801","status":"completed"},"tags":[]},"cell_type":"markdown","source":"* Consolidation is officially referred to as air space consolidation. It is a decrease in lung permeability due to infiltration of fluid, cells, or tissue replacing the air-containing spaces in the alveoli."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:54:15.452438Z","iopub.status.busy":"2021-01-19T21:54:15.451579Z","iopub.status.idle":"2021-01-19T21:54:16.48561Z","shell.execute_reply":"2021-01-19T21:54:16.486032Z"},"papermill":{"duration":1.220804,"end_time":"2021-01-19T21:54:16.486152","exception":false,"start_time":"2021-01-19T21:54:15.265348","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target12 = abnormal_train[abnormal_train['class_id']==4]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target12['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target12['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target12['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target12['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.312395,"end_time":"2021-01-19T21:54:17.089782","exception":false,"start_time":"2021-01-19T21:54:16.777387","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Atelectasis"},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:54:17.931763Z","iopub.status.busy":"2021-01-19T21:54:17.930879Z","iopub.status.idle":"2021-01-19T21:54:41.992686Z","shell.execute_reply":"2021-01-19T21:54:41.992223Z"},"papermill":{"duration":24.559142,"end_time":"2021-01-19T21:54:41.992803","exception":false,"start_time":"2021-01-19T21:54:17.433661","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"plot(\"Atelectasis\")","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.184611,"end_time":"2021-01-19T21:54:42.366508","exception":false,"start_time":"2021-01-19T21:54:42.181897","status":"completed"},"tags":[]},"cell_type":"markdown","source":"* Atelectasis is a condition where there is no air in part or all of the lungs. And the lungs are collapsed. A common cause of atelectasis is obstruction of the bronchi."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:54:42.75027Z","iopub.status.busy":"2021-01-19T21:54:42.743023Z","iopub.status.idle":"2021-01-19T21:54:44.032919Z","shell.execute_reply":"2021-01-19T21:54:44.032477Z"},"papermill":{"duration":1.484015,"end_time":"2021-01-19T21:54:44.033019","exception":false,"start_time":"2021-01-19T21:54:42.549004","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target13 = abnormal_train[abnormal_train['class_id']==1]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target13['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target13['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target13['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target13['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot(\"Pneumothorax\")","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:55:09.328191Z","iopub.status.busy":"2021-01-19T21:55:09.327297Z","iopub.status.idle":"2021-01-19T21:55:10.365115Z","shell.execute_reply":"2021-01-19T21:55:10.365559Z"},"papermill":{"duration":1.243372,"end_time":"2021-01-19T21:55:10.365687","exception":false,"start_time":"2021-01-19T21:55:09.122315","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target14 = abnormal_train[abnormal_train['class_id']==12]\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target14['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target14['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target14['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target14['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# coding: utf-8\nfrom tqdm import tqdm\nimport time\n\n# Set the total value \nbar = tqdm(total = 1000)\n# Add description\nbar.set_description('Progress rate')\nfor i in range(100):\n    # Set the progress\n    bar.update(25)\n    time.sleep(1)","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.190839,"end_time":"2021-01-19T21:55:10.748866","exception":false,"start_time":"2021-01-19T21:55:10.558027","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Data Visualization"},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:55:11.139133Z","iopub.status.busy":"2021-01-19T21:55:11.138314Z","iopub.status.idle":"2021-01-19T21:55:12.202307Z","shell.execute_reply":"2021-01-19T21:55:12.201383Z"},"papermill":{"duration":1.259666,"end_time":"2021-01-19T21:55:12.202412","exception":false,"start_time":"2021-01-19T21:55:10.942746","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"target1 = abnormal_train[abnormal_train['class_name']=='Aortic enlargement']\nsns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(target1['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(target1['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(target1['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(target1['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.327536,"end_time":"2021-01-19T21:55:12.776197","exception":false,"start_time":"2021-01-19T21:55:12.448661","status":"completed"},"tags":[]},"cell_type":"markdown","source":"* Let's look at the distribution of Rad-IDs.\n* R10, R9 and R8 are prominently high."},{"metadata":{"papermill":{"duration":0.193998,"end_time":"2021-01-19T21:55:13.838338","exception":false,"start_time":"2021-01-19T21:55:13.64434","status":"completed"},"tags":[]},"cell_type":"markdown","source":"* Now let's look at the gender distribution"},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:55:14.242514Z","iopub.status.busy":"2021-01-19T21:55:14.241913Z","iopub.status.idle":"2021-01-19T21:55:14.388875Z","shell.execute_reply":"2021-01-19T21:55:14.388163Z"},"papermill":{"duration":0.354007,"end_time":"2021-01-19T21:55:14.388991","exception":false,"start_time":"2021-01-19T21:55:14.034984","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(6,6))\nsns.countplot(info['sex'], data=train_df)\nplt.title(\"Sex distribution including those with no findings\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.195537,"end_time":"2021-01-19T21:55:14.785553","exception":false,"start_time":"2021-01-19T21:55:14.590016","status":"completed"},"tags":[]},"cell_type":"markdown","source":"* Let's also look at the distribution of sexes with no abnormal findings. There seems to be no particular difference."},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:55:15.189532Z","iopub.status.busy":"2021-01-19T21:55:15.18862Z","iopub.status.idle":"2021-01-19T21:55:15.316154Z","shell.execute_reply":"2021-01-19T21:55:15.316598Z"},"papermill":{"duration":0.334658,"end_time":"2021-01-19T21:55:15.316722","exception":false,"start_time":"2021-01-19T21:55:14.982064","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(6,6))\nsns.countplot(info['sex'], data=abnormal_train)\nplt.title(\"SEX distribution excluding those with no findings\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(6, 6))\nsns.countplot(x=\"rad_id\", data=train_df)\nplt.title(\"RAD ID Distribution including those with no findings\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(6,6))\nsns.countplot(x='rad_id', data=abnormal_train)\nplt.title(\"RAD ID Distribution excluding those with no findings\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Shape Analysis"},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:55:15.729668Z","iopub.status.busy":"2021-01-19T21:55:15.727591Z","iopub.status.idle":"2021-01-19T21:55:15.998712Z","shell.execute_reply":"2021-01-19T21:55:15.998281Z"},"papermill":{"duration":0.481672,"end_time":"2021-01-19T21:55:15.998813","exception":false,"start_time":"2021-01-19T21:55:15.517141","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(6,6))\nax = sns.scatterplot(x='rows', y='columns', data=info, alpha=0.3)\nplt.title(\"row(x) column(x) scatter plot\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:55:16.411303Z","iopub.status.busy":"2021-01-19T21:55:16.410242Z","iopub.status.idle":"2021-01-19T21:55:16.99612Z","shell.execute_reply":"2021-01-19T21:55:16.996568Z"},"papermill":{"duration":0.794797,"end_time":"2021-01-19T21:55:16.996697","exception":false,"start_time":"2021-01-19T21:55:16.2019","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(6,6))\nax = sns.scatterplot(x='x_min', y='y_min', data=abnormal_train, alpha=0.3)\nplt.title(\"min coordinate scatter plot\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:55:17.425761Z","iopub.status.busy":"2021-01-19T21:55:17.424681Z","iopub.status.idle":"2021-01-19T21:55:17.793615Z","shell.execute_reply":"2021-01-19T21:55:17.794035Z"},"papermill":{"duration":0.585473,"end_time":"2021-01-19T21:55:17.794153","exception":false,"start_time":"2021-01-19T21:55:17.20868","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(6,6))\nax = sns.scatterplot(x='x_max', y='y_max', data=abnormal_train, alpha=0.3)\nplt.title(\"max coordinate scatter plot\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.set_style('whitegrid')\nplt.figure()\nfig, ax = plt.subplots(2,2,figsize=(6,6))\nsns.distplot(abnormal_train['x_max'],kde=True,bins=50, color=\"red\", ax=ax[0,0])\nsns.distplot(abnormal_train['y_max'],kde=True,bins=50, color=\"blue\", ax=ax[0,1])\nsns.distplot(abnormal_train['x_min'],kde=True,bins=50, color=\"green\", ax=ax[1,0])\nsns.distplot(abnormal_train['y_min'],kde=True,bins=50, color=\"magenta\", ax=ax[1,1])\nlocs, labels = plt.xticks()\nplt.tick_params(axis='both', which='major', labelsize=12)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize = (6, 6))\nx = info[\"rows\"]\ny = info[\"columns\"]\nsns.distplot(x * y, kde = True, color = \"brown\")\nplt.xlabel(\"pixel count\", fontsize = 16)\nplt.title(\"Pixel Count Analysis\", fontsize = 18)\nplt.grid(True)\nplt.axis(\"on\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Features based on where bounding boxes for a class tend to be"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Create a list of classes and a dictionary\nclasses = train_df[['class_id', 'class_name', 'rad_id']]\\\n    .groupby(['class_id', 'class_name'])\\\n    .count()\\\n    .rename(columns={'rad_id': 'Number of records'})\\\n    .reset_index()\n\nfor index, row in classes.iterrows():\n    if index==0:\n        label_dict = {row['class_id']: row['class_name']}\n    else:\n        label_dict.update({row['class_id']: row['class_name']})\n\ntrain_df = pd.merge(train_df, dicom_meta, on='image_id', how='left')\ntrain_df['x_max'] = train_df['x_max']/train_df['Columns']\ntrain_df['x_min'] = train_df['x_min']/train_df['Columns']\ntrain_df['y_max'] = train_df['y_max']/train_df['Rows']\ntrain_df['y_min'] = train_df['y_min']/train_df['Rows']\ntrain_df['width'] = (train_df['x_max']-train_df['x_min'])\ntrain_df['height'] = (train_df['y_max']-train_df['y_min'])\ntrain_df['area'] = train_df['height']*train_df['width']\ntrain_df['x_center'] = (train_df['x_max']+train_df['x_min'])/2\ntrain_df['y_center'] = (train_df['y_max']+train_df['y_min'])/2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# In this cell, we create a feature for how many other classes a radiologist has already assigned \n# for the same image.\n\ntmp1 = pd.merge(train_df, train_df, on=['image_id', 'rad_id'], how='left').fillna(0)\n#tmp1 = tmp1[ (tmp1['class_id_x']!=tmp1['class_id_y']) | (tmp1['x_min_x']!=tmp1['x_min_y']) | (tmp1['x_max_x']!=tmp1['x_max_y']) | (tmp1['y_min_x']!=tmp1['y_min_y']) | (tmp1['y_max_x']!=tmp1['y_max_y'])]\n\n\ntmp_cols = ['image_id', 'rad_id', 'class_id_x', 'x_min_x', 'y_min_x', 'x_max_x', 'y_max_x']\ntmp1 = tmp1[ tmp_cols + ['class_id_y', 'x_max_y']]\\\n    .groupby(tmp_cols+['class_id_y'])\\\n    .count()\\\n    .reset_index()\\\n    .pivot(index=tmp_cols,\n           columns='class_id_y', values='x_max_y')\\\n    .add_prefix('other_')\\\n    .reset_index()\\\n    .rename(columns={'class_id_x':'class_id',\n                     'class_id_x': 'class_id',\n                     'x_min_x': 'x_min',\n                     'y_min_x': 'y_min',\n                     'x_max_x': 'x_max',\n                     'y_max_x': 'y_max'})\n\n\ntrain_df = pd.merge(train_df, tmp1, \n                 on=['image_id', 'rad_id', 'class_id', 'x_min', 'y_min', 'x_max', 'y_max'],\n                 how='left')\ntrain_df[['other_'+str(i) for i in range(15)]] = train_df[['other_'+str(i) for i in range(15)]]\\\n    .fillna(0)\\\n    .astype(np.int)\n\n# Finally, we subtract the extra count of +1 for each label itself (when we want to predict it, we\n# do not want to have a feature that leaks the label, which it otherwise would).\nfor idx, row in train_df.iterrows():\n    if row['class_id']<14:        \n        train_df['other_' + str(row['class_id'])].values[idx] += -1\n    ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"locations = np.zeros((14, 1000, 1000))\nfor index, row in tqdm(train_df.iterrows(), total=train_df.shape[0]):\n    if row['class_id']<14:\n        locations[row['class_id'], \n                  ((np.round(row['y_min'],3)*1000).astype(np.int)):((np.round(row['y_max'],3)*1000).astype(np.int)), \n                  ((np.round(row['x_min'],3)*1000).astype(np.int)):((np.round(row['x_max'],3)*1000).astype(np.int))] += 1\n        \nclasscounts = train_df[['image_id', 'rad_id', 'class_id','class_name']]\\\n    .groupby(['image_id', 'rad_id', 'class_id'])\\\n    .count()\\\n    .reset_index()\\\n    .pivot(index=['image_id', 'rad_id'], columns='class_id', values='class_name')\\\n    .rename(columns={i:'n_class'+str(i) for i in range(15)})\\\n    .fillna(0)\n   \nclassareas = train_df[['image_id', 'rad_id', 'class_id','area']]\\\n    .groupby(['image_id', 'rad_id', 'class_id'])\\\n    .sum()\\\n    .reset_index()\\\n    .pivot(index=['image_id', 'rad_id'], columns='class_id', values='area')\\\n    .rename(columns={i:'area_class'+str(i) for i in range(15)})\\\n    .fillna(0)\n\ntrain_df = pd.merge( pd.merge( train_df, classcounts, on=['image_id', 'rad_id'], how='left'), \n                  classareas, on=['image_id', 'rad_id'], how='left')\ntrain_df = train_df[train_df['class_id']!=14]\n\nclasses = train_df[['class_id', 'class_name', 'rad_id']]\\\n    .groupby(['class_id', 'class_name'])\\\n    .count()\\\n    .rename(columns={'rad_id': 'Number of records'})\\\n    .reset_index()\n    \nfor index, row in classes.iterrows():\n    if index==0:\n        label_dict = {row['class_id']: row['class_name']}\n    else:\n        label_dict.update({row['class_id']: row['class_name']})\n        \nf, axs = plt.subplots(5, 3, sharey=True, sharex=True, figsize=(16,28));\n\nfor class_id in range(14):\n    axs[class_id // 3, class_id - 3*(class_id // 3)].imshow(locations[class_id], cmap='inferno', interpolation='nearest');\n    axs[class_id // 3, class_id - 3*(class_id // 3)].set_title(str(class_id) + ': ' + label_dict[class_id])\n    \nplt.show();    ","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.219217,"end_time":"2021-01-19T21:55:21.440372","exception":false,"start_time":"2021-01-19T21:55:21.221155","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Trends in bounding boxes differences"},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_width_of__bounding_boxes(abnormal_train)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:55:22.208757Z","iopub.status.busy":"2021-01-19T21:55:22.205733Z","iopub.status.idle":"2021-01-19T21:55:22.244468Z","shell.execute_reply":"2021-01-19T21:55:22.245621Z"},"papermill":{"duration":0.435403,"end_time":"2021-01-19T21:55:22.245798","exception":false,"start_time":"2021-01-19T21:55:21.810395","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"classes = train_df[['class_id', 'class_name', 'rad_id']].groupby(['class_id', 'class_name']).count().rename(columns={'rad_id': 'Number of records'}).reset_index()\n\nfor index, row in classes.iterrows():\n    if index==0:\n        label_dict = {row['class_id']: row['class_name']}\n    else:\n        label_dict.update({row['class_id']: row['class_name']})","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:55:22.875714Z","iopub.status.busy":"2021-01-19T21:55:22.874667Z","iopub.status.idle":"2021-01-19T21:55:22.877938Z","shell.execute_reply":"2021-01-19T21:55:22.877465Z"},"papermill":{"duration":0.23467,"end_time":"2021-01-19T21:55:22.878041","exception":false,"start_time":"2021-01-19T21:55:22.643371","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"cols = ['#e41a1c', '#377eb8','#4daf4a','#984ea3','#ff7f00','#ffff33','#a65628','#f781bf','#999999', '#000000', '#1b9e77', '#d95f02', '#7570b3', '#e7298a']","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:55:23.389797Z","iopub.status.busy":"2021-01-19T21:55:23.388644Z","iopub.status.idle":"2021-01-19T21:55:23.860348Z","shell.execute_reply":"2021-01-19T21:55:23.861409Z"},"papermill":{"duration":0.751663,"end_time":"2021-01-19T21:55:23.861585","exception":false,"start_time":"2021-01-19T21:55:23.109922","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"train_df['size'] = (train_df['x_max']-train_df['x_min'])*(train_df['y_max']-train_df['y_min'])\nsizes = train_df.loc[train_df['class_id']<14, ['class_id', 'size']].groupby('class_id').mean().reset_index()\n\nplt.figure(figsize=(6, 6));\nplt.bar(sizes['class_id'], sizes['size'], \n        tick_label=[str(i) + ': ' + label_dict[i] for i in range(14)],\n        color=cols);\nplt.xticks(rotation='vertical');","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:55:24.393558Z","iopub.status.busy":"2021-01-19T21:55:24.392532Z","iopub.status.idle":"2021-01-19T21:55:24.668886Z","shell.execute_reply":"2021-01-19T21:55:24.669337Z"},"papermill":{"duration":0.519153,"end_time":"2021-01-19T21:55:24.669464","exception":false,"start_time":"2021-01-19T21:55:24.150311","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"numbers = train_df.loc[train_df['class_id']<14, ['image_id', 'rad_id', 'class_id', 'size']].groupby(['image_id', 'rad_id', 'class_id']).count().reset_index().groupby('class_id').mean('size').reset_index()\n\nplt.figure(figsize=(6, 6));\nplt.bar(numbers['class_id'], numbers['size'], tick_label=[str(i) + ': ' + label_dict[i] for i in range(14)], color=cols);\nplt.xticks(rotation='vertical');","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:55:25.222384Z","iopub.status.busy":"2021-01-19T21:55:25.221427Z","iopub.status.idle":"2021-01-19T21:55:25.363593Z","shell.execute_reply":"2021-01-19T21:55:25.364138Z"},"papermill":{"duration":0.463767,"end_time":"2021-01-19T21:55:25.364295","exception":false,"start_time":"2021-01-19T21:55:24.900528","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"tmpdf = train_df[['class_id', 'image_id', 'rad_id']].groupby(['class_id', 'image_id']).count().reset_index()\ntmpdf['rad_id'] = np.minimum(tmpdf['rad_id'].values, 1)\ncorr  = tmpdf.pivot(index='image_id', columns='class_id', values='rad_id').fillna(0).reset_index(drop=True).corr()\ncorr.style.background_gradient(cmap='coolwarm', vmin=-1.0, vmax=1.0).set_precision(2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Let's see how it correlates with the different classes\n* The following correlations were found to be strong.\n* '0: Aortic enlargement' and '3: Cardiomegaly'.\n* '１0: Pleural effusion' and '１１: Pleural thickening'.\n* '１１: Pleural thickening' and '１３: Pulmonary fibrosis'."},{"metadata":{},"cell_type":"markdown","source":"There are 15 different radiographic observations which correspond to:\n\n* 0 - Aortic enlargement\n* 1 - Atelectasis\n* 2 - Calcification\n* 3 - Cardiomegaly\n* 4 - Consolidation\n* 5 - ILD\n* 6 - Infiltration\n* 7 - Lung Opacity\n* 8 - Nodule/Mass\n* 9 - Other lesion\n* 10 - Pleural effusion\n* 11 - Pleural thickening\n* 12 - Pneumothorax\n* 13 - Pulmonary fibrosis\n* 14 - No finding"},{"metadata":{},"cell_type":"markdown","source":"* View the distribution by object's bounding box and class ID"},{"metadata":{},"cell_type":"markdown","source":"# Plot bounding box"},{"metadata":{"trusted":true},"cell_type":"code","source":"def dicom2array(path, voi_lut=True, fix_monochrome=True):\n    dicom = pydicom.read_file(path)\n    # VOI LUT (if available by DICOM device) is used to\n    # transform raw DICOM data to \"human-friendly\" view\n    if voi_lut:\n        data = apply_voi_lut(dicom.pixel_array, dicom)\n    else:\n        data = dicom.pixel_array\n    # depending on this value, X-ray may look inverted - fix that:\n    if fix_monochrome and dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n    data = data - np.min(data)\n    data = data / np.max(data)\n    data = (data * 255).astype(np.uint8)\n    return data\n        \n    \ndef plot_img(img, size=(7, 7), is_rgb=True, title=\"\", cmap='gray'):\n    plt.figure(figsize=size)\n    plt.imshow(img, cmap=cmap)\n    plt.suptitle(title)\n    plt.show()\n    \n\ndef plot_imgs(imgs, cols=4, size=7, is_rgb=True, title=\"\", cmap='gray', img_size=(500,500)):\n    rows = len(imgs)//cols + 1\n    fig = plt.figure(figsize=(cols*size, rows*size))\n    for i, img in enumerate(imgs):\n        if img_size is not None:\n            img = cv2.resize(img, img_size)\n        fig.add_subplot(rows, cols, i+1)\n        plt.imshow(img, cmap=cmap)\n    plt.suptitle(title)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"imgs = []\nimg_ids = abnormal_train['image_id'].values\nclass_ids = abnormal_train['class_id'].unique()\n\n# map label_id to specify color\nlabel2color = {class_id:[randint(0,255) for i in range(3)] for class_id in class_ids}\nthickness = 3\nscale = 5\n\n\n\nfor i in range(8):\n    img_id = random.choice(img_ids)\n    img_path = f'{dataset_dir}/train/{img_id}.dicom'\n    img = dicom2array(path=img_path)\n    img = cv2.resize(img, None, fx=1/scale, fy=1/scale)\n    img = np.stack([img, img, img], axis=-1)\n    \n    boxes = abnormal_train.loc[abnormal_train['image_id'] == img_id, ['x_min', 'y_min', 'x_max', 'y_max']].values/scale\n    labels = abnormal_train.loc[abnormal_train['image_id'] == img_id, ['class_id']].values.squeeze()\n    \n    for label_id, box in zip(labels, boxes):\n        color = label2color[label_id]\n        img = cv2.rectangle(\n            img,\n            (int(box[0]), int(box[1])),\n            (int(box[2]), int(box[3])),\n            color, thickness\n    )\n    img = cv2.resize(img, (500,500))\n    imgs.append(img)\n    \nplot_imgs(imgs, cmap=None)","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.24783,"end_time":"2021-01-19T21:57:06.549579","exception":false,"start_time":"2021-01-19T21:57:06.301749","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Modeling"},{"metadata":{"execution":{"iopub.execute_input":"2021-01-19T21:57:07.064885Z","iopub.status.busy":"2021-01-19T21:57:07.063937Z","iopub.status.idle":"2021-01-19T21:57:07.067941Z","shell.execute_reply":"2021-01-19T21:57:07.067416Z"},"papermill":{"duration":0.270792,"end_time":"2021-01-19T21:57:07.068033","exception":false,"start_time":"2021-01-19T21:57:06.797241","status":"completed"},"tags":[],"trusted":true},"cell_type":"code","source":"# training dataset\nfeatures = ['image_id' ,'class_id', 'rad_id', 'x_min', 'y_min', 'x_max', 'y_max']\ntrain_ftr = train_df[features]\ntrain_ftr.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# coding: utf-8\nfrom tqdm import tqdm\nimport time\n\n# Set the total value \nbar = tqdm(total = 1000)\n# Add description\nbar.set_description('Progress rate')\nfor i in range(100):\n    # Set the progress\n    bar.update(25)\n    time.sleep(1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Submission"},{"metadata":{"trusted":true},"cell_type":"code","source":"# predictions.to_csv('submission.csv',index= False)","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.250913,"end_time":"2021-01-19T21:57:07.566115","exception":false,"start_time":"2021-01-19T21:57:07.315202","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Acknowledgements\n* [EDA - VinBigData Chest X-ray Abnormalities](https://www.kaggle.com/trungthanhnguyen0502/eda-vinbigdata-chest-x-ray-abnormalities)\n* [VinBigData: EDA All You need to know](https://www.kaggle.com/dhananjay3/vinbigdata-eda-all-you-need-to-know)\n[*Chest_X-ray: Knowledges for the 14 abnormalities*](https://www.kaggle.com/sakuraandblackcat/chest-x-ray-knowledges-for-the-14-abnormalities)\n* [Convert dicom to np.array - the correct way](https://www.kaggle.com/raddar/convert-dicom-to-np-array-the-correct-way)\n* [VinBigData Chest X-ray Abnormalities Detection](https://www.kaggle.com/hamditarek/vinbigdata-chest-x-ray-abnormalities-detection)\n* [EDA & .dicom reading: VinBigData Chest X-ray](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray)\n* [Chest_X-ray_Starter](https://www.kaggle.com/drcapa/chest-x-ray-starter)\n* [VinBigData Chest X-ray EDA with Plotly](https://www.kaggle.com/debarshichanda/vinbigdata-chest-x-ray-eda-with-plotly)\n* [VinBigData Retinanet-Detection [Training] ](https://www.kaggle.com/akhileshdkapse/vinbigdata-retinanet-detection-training/data)\n* [VinBigData: EDA All You need to know](https://www.kaggle.com/dhananjay3/vinbigdata-eda-all-you-need-to-know)\n* [VBD Chest X-ray Abnormalities Detection | EDA📊🔴](https://www.kaggle.com/mrutyunjaybiswal/vbd-chest-x-ray-abnormalities-detection-eda)\n* [Finding data issues and mislabeled bounding boxes](https://www.kaggle.com/bjoernholzhauer/finding-data-issues-and-mislabeled-bounding-boxes)\n* [Chest X-ray Abnormalities Doctor-EDA](https://www.kaggle.com/anantgupt/chest-x-ray-abnormalities-doctor-eda)\n* [All you need to know about DICOM](https://www.kaggle.com/asimzahid/all-you-need-to-know-about-dicom)\n* [EDA_train_csv](https://www.kaggle.com/soudainchat/eda-train-csv)"},{"metadata":{"papermill":{"duration":0.25386,"end_time":"2021-01-19T21:57:08.069806","exception":false,"start_time":"2021-01-19T21:57:07.815946","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Work in progress…"},{"metadata":{},"cell_type":"markdown","source":"# Your upvote is my motivation"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}