{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Team Kenya: Kaggle RSNA competition.\n<div style=\"color:white;display:fill;\n            background-color:#9e4c4c9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>Team Members</b></p>\n</div>\n\n> 1. Munyala Eliud - Team Lead (Data Eng.)\n2. Wanjiru Kariuki - MLOps\n3. Mukiri Mwirigi - (Blockchain|Crypto|AI)\n4. Paul Ndirangu - (MLOps)\n5. Ken Kuria - (MLOps)\n```","metadata":{}},{"cell_type":"markdown","source":"# 1. Introduction\n\n<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>1.1 Background</b></p>\n</div>\n\n**Context**\n\nRSNA has teamed with the American Society of Neuroradiology (ASNR) and the American Society of Spine Radiology (ASSR) to conduct an AI challenge competition exploring whether artificial intelligence can be used to aid in the detection and localization of cervical spine fractures.\n\nTo create the ground truth dataset, the challenge planning task force collected imaging data sourced from twelve sites on six continents, including approximately 3,000 CT studies. Spine radiology specialists from the ASNR and ASSR provided expert image level annotations these studies to indicate the presence, vertebral level and location of any cervical spine fractures.\n\nIn this challenge competition, you will try to develop machine learning models that match the radiologists' performance in detecting and localizing fractures to the seven vertebrae that comprise the cervical spine. Winners will be recognized at an event during the RSNA 2022 annual meeting.\n\n<center>\n<img src=\"https://www.holisticbodyworks.com.au/wp-content/uploads/2018/05/Thoracic-Spine.jpg\" width=400>\n</center>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>1.2 Data Understanding</b></p>\n</div>\n\n**Dataset origin**\n\nThe dataset we are using is made up of roughly **3000 CT studies**, [from twelve locations and across six continents](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/overview/acknowledgements). Spine **radiology specialists** have provided **annotations** to indicate the presence, vertebral level and location of any cervical spine fractures.\n\nSpecial thanks to the competition **hosts** for providing such a comprehensive dataset:\n* *Radiological Society of North America (RSNA)*\n* *American Society of Neuroradiology (ASNR)*\n* *American Society of Spine Radiology (ASSR)*","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>1.3 Evaluation metric</b></p>\n</div>\n\nWe need to predict the **probability of fracture** for each of the **seven cervical vertebrae** denoted by C1, C2, C3, C4, C5, C6 and C7 as well as an **overall probability** of any fractures in the cervical spine. This means there will be **8 rows per image id** in the submission file. Note that fractures in the skull base, thoracic spine, ribs, and clavicles are **ignored**.\n    \nThe **competition metric** is a **weighted multi-label logarithmic loss** (averaged across all patients)\n","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>1.4 Code requirements</b></p>\n</div>\n\nThis is a **code competition**, which means that **submissions** are made through **notebooks**. Furthermore, the submission notebook is subject to these conditions:\n\n* **run-time** (CPU/GPU) **<= 9 hours**\n* **internet** access **disabled**\n* **external data is allowed**, including pre-trained models\n* the submission file must be named **submission.csv**\n              \nNote that the **test set is hidden**, and will populated when you submit your notebook. \n\n<hr>\n\n*Competition timeline:*\n    \n* Start date - 28th July 2022\n* Finish date - 27th October 2022","metadata":{}},{"cell_type":"markdown","source":"## 2. Getting Started\n\n<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>2.1 Import Libraries</b></p>\n</div>","metadata":{}},{"cell_type":"code","source":"# To read compressed dicom files:\n\n'''Online'''\n! pip install python-gdcm\n! pip install pylibjpeg pylibjpeg-libjpeg pydicom","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:55:33.339227Z","iopub.execute_input":"2022-09-04T07:55:33.339648Z","iopub.status.idle":"2022-09-04T07:55:58.131322Z","shell.execute_reply.started":"2022-09-04T07:55:33.339603Z","shell.execute_reply":"2022-09-04T07:55:58.129705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# libraries\ntry:\n    import numpy as np\n    import pandas as pd\n    import matplotlib.pyplot as plt\n    %matplotlib inline\n    import matplotlib.patches as patches\n    import seaborn as sns\n    sns.set(style='darkgrid', font_scale=1.6)\n    import cv2\n    import os\n    from os import listdir\n    import re\n    import gc\n    import pydicom\n    import os\n    from glob import glob\n    from pydicom.pixel_data_handlers.util import apply_voi_lut\n    from tqdm import tqdm\n    from pprint import pprint\n    from time import time\n    import itertools\n    from skimage import measure \n    from mpl_toolkits.mplot3d.art3d import Poly3DCollection\n    import nibabel as nib\n    from glob import glob\n    import warnings\n    #warnings.filterwarnings(\"ignore\", category=DeprecationWarning)\n    #warnings.filterwarnings(\"ignore\", category=UserWarning)\n    #warnings.filterwarnings(\"ignore\", category=FutureWarning)\nexcept:\n    print(\"Import errors\")    \n    \n\n#.......\n","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:55:58.134533Z","iopub.execute_input":"2022-09-04T07:55:58.1352Z","iopub.status.idle":"2022-09-04T07:55:59.446264Z","shell.execute_reply.started":"2022-09-04T07:55:58.135158Z","shell.execute_reply":"2022-09-04T07:55:59.444995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>2.2 Loading the Data</b></p>\n</div>","metadata":{}},{"cell_type":"code","source":"# Load metadata\ntrain_df = pd.read_csv(\"../input/rsna-2022-cervical-spine-fracture-detection/train.csv\")\ntrain_bbox = pd.read_csv(\"../input/rsna-2022-cervical-spine-fracture-detection/train_bounding_boxes.csv\")\ntest_df = pd.read_csv(\"../input/rsna-2022-cervical-spine-fracture-detection/test.csv\")\nss = pd.read_csv(\"../input/rsna-2022-cervical-spine-fracture-detection/sample_submission.csv\")\n\n# Print dataframe shapes\nprint('train shape:', train_df.shape)\nprint('train bbox shape:', train_bbox.shape)\nprint('test shape:', test_df.shape)\nprint('ss shape:', ss.shape)\nprint('')\n\n# Show first few entries\ntrain_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:55:59.448174Z","iopub.execute_input":"2022-09-04T07:55:59.448682Z","iopub.status.idle":"2022-09-04T07:55:59.522265Z","shell.execute_reply.started":"2022-09-04T07:55:59.448636Z","shell.execute_reply":"2022-09-04T07:55:59.521147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATA_DIR = \"../input/rsna-2022-cervical-spine-fracture-detection\"\nTRAIN_DIR = os.path.join(DATA_DIR, \"train_images\")\nSEGM_DIR = os.path.join(DATA_DIR, \"segmentations\")","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:55:59.524783Z","iopub.execute_input":"2022-09-04T07:55:59.525123Z","iopub.status.idle":"2022-09-04T07:55:59.532094Z","shell.execute_reply.started":"2022-09-04T07:55:59.525083Z","shell.execute_reply":"2022-09-04T07:55:59.530886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> **train.csv** - contains metadata for train_images.\n> * StudyInstanceUID - The study ID. There is one unique study ID for each patient scan.\n> * patient_overall - The patient level outcome, i.e. if any of the vertebrae are fractured.\n> * C[1-7] - Whether the given vertebrae is fractured.","metadata":{}},{"cell_type":"code","source":"train_bbox.head(3)","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:55:59.533859Z","iopub.execute_input":"2022-09-04T07:55:59.534328Z","iopub.status.idle":"2022-09-04T07:55:59.552471Z","shell.execute_reply.started":"2022-09-04T07:55:59.534286Z","shell.execute_reply":"2022-09-04T07:55:59.550272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> **train_bounding_boxes.csv** - contains bounding boxes of where fractures occured for a subset of the training set.\n> * StudyInstanceUID - The study ID. There is one unique study ID for each patient scan.\n> * x -> x-coordinate of bounding box bottom left corner\n> * y -> y-coordinate of bounding box bottom left corner\n> * width -> width of bounding box\n> * height -> height of bounding box\n> * slice_number -> slice number of scan\n\n**Note:** We only have bounding boxes for a [subset](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/343105) of the training set. We'll explore the exact proportion later on.","metadata":{}},{"cell_type":"code","source":"test_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:55:59.553857Z","iopub.execute_input":"2022-09-04T07:55:59.554197Z","iopub.status.idle":"2022-09-04T07:55:59.565383Z","shell.execute_reply.started":"2022-09-04T07:55:59.554166Z","shell.execute_reply":"2022-09-04T07:55:59.564389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> **test.csv** - contains metadata for test_images.\n> * row_id - The row ID. This will match the same column in the sample submission file.\n> * StudyInstanceUID - The study ID.\n> * prediction_type - Which one of the eight target columns needs a prediction in this row.\n\n**Note:** The full test set will be **populated at inference time**.","metadata":{}},{"cell_type":"code","source":"ss.head()","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:55:59.566766Z","iopub.execute_input":"2022-09-04T07:55:59.567153Z","iopub.status.idle":"2022-09-04T07:55:59.581906Z","shell.execute_reply.started":"2022-09-04T07:55:59.567122Z","shell.execute_reply":"2022-09-04T07:55:59.580668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Exploratory Data Analysis\n<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>3.1 Fracture distributions</b></p>\n</div>\n\n> **Insights:**\n* The overall target is roughly **balanced** (52/48 split). \n* **C7** has the **highest proportion of fractures** (19%) whereas **C3** has the **lowest** (4%). \n* Several patients have **more than one** fracture.\n* If **multiple fractures** occur on a single patient, they tend to occur in vertebrae **close together**, e.g. C4 & C5 as opposed to C1 & C7.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12,7))\nplt.plot()\nax1 = sns.countplot(data=train_df, x='patient_overall')\nfor container in ax1.containers:\n    ax1.bar_label(container)\nplt.title('Fractures by patient')\nplt.ylim([0,1300])\n","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:55:59.583093Z","iopub.execute_input":"2022-09-04T07:55:59.583599Z","iopub.status.idle":"2022-09-04T07:55:59.859723Z","shell.execute_reply.started":"2022-09-04T07:55:59.583567Z","shell.execute_reply":"2022-09-04T07:55:59.858604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# Unpivot train_df for plotting\ntrain_melt = pd.melt(train_df, id_vars = ['StudyInstanceUID', 'patient_overall'],\n             value_vars = ['C1','C2','C3','C4','C5','C6','C7'],\n             var_name=\"Vertebrae\",\n             value_name=\"Fractured\")\nplt.figure(figsize=(12,7))\nplt.plot()\nax2 = sns.countplot(data=train_melt, x='Vertebrae', hue='Fractured')\nfor container in ax2.containers:\n    ax2.bar_label(container)\nplt.title('Fractures by vertebrae')\nplt.ylim([0,2800])","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:55:59.861306Z","iopub.execute_input":"2022-09-04T07:55:59.86171Z","iopub.status.idle":"2022-09-04T07:56:00.235779Z","shell.execute_reply.started":"2022-09-04T07:55:59.861676Z","shell.execute_reply":"2022-09-04T07:56:00.234517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12,7))\nax = sns.countplot(x = train_df[['C1','C2','C3','C4','C5','C6','C7']].sum(axis=1))\nfor container in ax.containers:\n    ax.bar_label(container)\nplt.title('Number of fractures by patient')\nplt.xlabel('Number of fractures')\nplt.ylim([0,1300])","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:00.24066Z","iopub.execute_input":"2022-09-04T07:56:00.24101Z","iopub.status.idle":"2022-09-04T07:56:00.525741Z","shell.execute_reply.started":"2022-09-04T07:56:00.24098Z","shell.execute_reply":"2022-09-04T07:56:00.524645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Heatmap of correlations\nplt.figure(figsize=(8,6))\nsns.heatmap(train_df[['C1','C2','C3','C4','C5','C6','C7']].corr(), cmap='rocket', vmin=-1, vmax=1)\nplt.title('Correlations')","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:00.526934Z","iopub.execute_input":"2022-09-04T07:56:00.527773Z","iopub.status.idle":"2022-09-04T07:56:00.845483Z","shell.execute_reply.started":"2022-09-04T07:56:00.527741Z","shell.execute_reply":"2022-09-04T07:56:00.844208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>3.2 Study Id's</b></p>\n</div>\n\nThe cases in the dataset have **unique id's** like '1.2.826.0.1.3680043.6200'. It turns out only the **number after the last full stop** is important. ","metadata":{}},{"cell_type":"code","source":"# Example\nsample = train_df['StudyInstanceUID'][0]\nsample","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:00.847028Z","iopub.execute_input":"2022-09-04T07:56:00.847407Z","iopub.status.idle":"2022-09-04T07:56:00.853642Z","shell.execute_reply.started":"2022-09-04T07:56:00.84734Z","shell.execute_reply":"2022-09-04T07:56:00.852595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Find unique numbers in study id's\nfor i in range(7):\n    print(train_df['StudyInstanceUID'].map(lambda x : x.split('.')[i]).unique())","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:00.855561Z","iopub.execute_input":"2022-09-04T07:56:00.855968Z","iopub.status.idle":"2022-09-04T07:56:00.877673Z","shell.execute_reply.started":"2022-09-04T07:56:00.855934Z","shell.execute_reply":"2022-09-04T07:56:00.876301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>3.3 Dicom Files</b></p>\n</div>\nA **.dcm** file follows the **Digital Imaging and Communications in Medicine** (DICOM) format. It is the standard format used for storing **medical images** and **related metadata**. It dates back to 1983, although it has been revised many times. \n\nWe can use the [pydicom library](https://pydicom.github.io/) to open and explore these files.","metadata":{}},{"cell_type":"code","source":"example = \"../input/rsna-2022-cervical-spine-fracture-detection/train_images/1.2.826.0.1.3680043.10001/1.dcm\"\nexample_ds = pydicom.dcmread(example)\nexample_ds","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:00.87912Z","iopub.execute_input":"2022-09-04T07:56:00.879796Z","iopub.status.idle":"2022-09-04T07:56:00.906668Z","shell.execute_reply.started":"2022-09-04T07:56:00.879761Z","shell.execute_reply":"2022-09-04T07:56:00.905419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The **image data** is stored in an array under **'Pixel Data'**. Everything else is **metadata**.\n* The **'Rows'** and **'Columns'** values tell us the **image size**.\n* The **'Pixel Spacing'** and **'Slice Thickness'** tell us the **pixel size** and **thickness**.\n* The **'Window Center'** and **'Window Width'** give information about the **brightness** and **contrast** of the image respectively.\n* The **'Rescale Intercept'** and **'Rescale Slope'** determine the range of pixel values. ([ref](https://stackoverflow.com/questions/10193971/rescale-slope-and-rescale-intercept)).\n* **'ImagePositionPatient'** tells us the x, y, and z coordinates of the top left corner of each image in mm\n* **InstanceNumber** is the slice number.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>4.2 Exploring the images</b></p>\n</div>\n\nLet's look at the what the **image data** from the dcm files look like.","metadata":{}},{"cell_type":"code","source":"# Adapted from https://www.kaggle.com/code/andradaolteanu/rsna-fracture-detection-dicom-images-explore\nbase_path = \"../input/rsna-2022-cervical-spine-fracture-detection\"\npatient_id = '1.2.826.0.1.3680043.12281'\ndcm_paths = glob(f\"{base_path}/train_images/{patient_id}/*\")\ndef atoi(text):\n    return int(text) if text.isdigit() else text\ndef natural_keys(text):\n    return [atoi(c) for c in re.split(r'(\\d+)', text)]\ndcm_paths.sort(key=natural_keys)\n\n# Get images\nfiles = [pydicom.dcmread(path) for path in dcm_paths]\nimages = [apply_voi_lut(file.pixel_array, file) for file in files]\n\n# Plot images\nfig, axes = plt.subplots(nrows=3, ncols=6, figsize=(24,12))\nfig.suptitle(f'ID: {patient_id}', weight=\"bold\", size=20)\n\nstart = 110\nfor i in range(start,start+18):\n    img = images[i]\n    file = files[i]\n    slice_no = i\n\n    # Plot the image\n    x = (i-start) // 6\n    y = (i-start) % 6\n\n    axes[x, y].imshow(img, cmap=\"bone\")\n    axes[x, y].set_title(f\"Slice: {slice_no}\", fontsize=14, weight='bold')\n    axes[x, y].axis('off')","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:00.908231Z","iopub.execute_input":"2022-09-04T07:56:00.908996Z","iopub.status.idle":"2022-09-04T07:56:08.804037Z","shell.execute_reply.started":"2022-09-04T07:56:00.908953Z","shell.execute_reply":"2022-09-04T07:56:08.802563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is great, but at the moment we don't know which **vertebrae** is being shown in each image. One way to work this out is by using the **segmentations**.","metadata":{}},{"cell_type":"markdown","source":"# 5. Segmentations\n\n<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>5.1 What is NIfTI? </b></p>\n</div>\n\nA **.nii** file follows the **Neuroimaging Informatics Technology Initiative** (NIfTI) format. Compared to the DICOM, NIfTI is **simpler** and **easier** to support. \n\nTo open .nii files we can use the [nibabel library](https://nipy.org/nibabel/gettingstarted.html).","metadata":{}},{"cell_type":"code","source":"ex_path2 = f\"../input/rsna-2022-cervical-spine-fracture-detection/segmentations/{patient_id}.nii\"\nnii_example = nib.load(ex_path2)\n\n# Convert to numpy array\nseg = nii_example.get_fdata()\nseg.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:08.806402Z","iopub.execute_input":"2022-09-04T07:56:08.806793Z","iopub.status.idle":"2022-09-04T07:56:10.376425Z","shell.execute_reply.started":"2022-09-04T07:56:08.806757Z","shell.execute_reply":"2022-09-04T07:56:10.374978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Align orientation with images\nseg = seg[:, ::-1, ::-1].transpose(2, 1, 0)\nseg.shape","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:10.37783Z","iopub.execute_input":"2022-09-04T07:56:10.378152Z","iopub.status.idle":"2022-09-04T07:56:10.385672Z","shell.execute_reply.started":"2022-09-04T07:56:10.378123Z","shell.execute_reply":"2022-09-04T07:56:10.384517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each nifti file contains segmentations for **all slices** in a scan. However, we need to be careful about the **orientation** of the segmentations.\n\n> NIFTI files consist of segmentation in the sagittal plane, while the DICOM files are in the axial plane.\n\nDealing with NIFTI files [discussion post](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/340612).","metadata":{}},{"cell_type":"code","source":"data_folder = '../input/rsna-2022-cervical-spine-fracture-detection/'\nnii_dir = data_folder + 'segmentations/'","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:10.387453Z","iopub.execute_input":"2022-09-04T07:56:10.388343Z","iopub.status.idle":"2022-09-04T07:56:10.400228Z","shell.execute_reply.started":"2022-09-04T07:56:10.388288Z","shell.execute_reply":"2022-09-04T07:56:10.397818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> [**<em>glob</em>** (short for global)](https://towardsdatascience.com/the-python-glob-module-47d82f4cbd2d) is used to return all file paths that match a specific pattern. We can use glob to search for a specific file pattern, or perhaps more usefully, search for files where the filename matches a certain pattern by using wildcard characters.","metadata":{}},{"cell_type":"code","source":"nii_paths = glob(nii_dir + '*.nii')\nprint(f'counts: {len(nii_paths)}')\nprint(f'sample: {nii_paths[0]}')","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:10.401823Z","iopub.execute_input":"2022-09-04T07:56:10.402219Z","iopub.status.idle":"2022-09-04T07:56:10.435825Z","shell.execute_reply.started":"2022-09-04T07:56:10.402188Z","shell.execute_reply":"2022-09-04T07:56:10.434882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>5.2 Exploring the masks</b></p>\n</div>","metadata":{}},{"cell_type":"code","source":"# Plot images\nfig, axes = plt.subplots(nrows=3, ncols=6, figsize=(24,12))\nfig.suptitle(f'ID: {patient_id}', weight=\"bold\", size=20)\n\nfor i in range(start,start+18):\n    mask = seg[i]\n    slice_no = i\n\n    # Plot the image\n    x = (i-start) // 6\n    y = (i-start) % 6\n\n    axes[x, y].imshow(mask, cmap='inferno')\n    axes[x, y].set_title(f\"Slice: {slice_no}\", fontsize=14, weight='bold')\n    axes[x, y].axis('off')","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:10.438614Z","iopub.execute_input":"2022-09-04T07:56:10.440776Z","iopub.status.idle":"2022-09-04T07:56:12.615677Z","shell.execute_reply.started":"2022-09-04T07:56:10.44071Z","shell.execute_reply":"2022-09-04T07:56:12.614278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Compare these with the previous images. These masks give us the **location** of the vertebrae, which is very helpful because we know the fractures can only occur in these regions.\n\nThey also tells us which **vertebrae** are in the images. By looking at the unique values in each slice, we find a 0 for the background and another number like 2 for vertebrae C2. ","metadata":{}},{"cell_type":"code","source":"np.unique(seg[116])","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:12.617668Z","iopub.execute_input":"2022-09-04T07:56:12.618104Z","iopub.status.idle":"2022-09-04T07:56:12.635055Z","shell.execute_reply.started":"2022-09-04T07:56:12.618063Z","shell.execute_reply":"2022-09-04T07:56:12.633899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Unfortunately, we don't have segmentations for the whole train set. ","metadata":{}},{"cell_type":"code","source":"# Number of cases with masks\nseg_paths = glob(f\"{base_path}/segmentations/*\")\nprint(f'Number of cases with segmentations: {len(seg_paths)}, ({np.round(100*len(seg_paths)/len(train_df),1)}%)')","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:12.637102Z","iopub.execute_input":"2022-09-04T07:56:12.637569Z","iopub.status.idle":"2022-09-04T07:56:12.644865Z","shell.execute_reply.started":"2022-09-04T07:56:12.637526Z","shell.execute_reply":"2022-09-04T07:56:12.643622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Extract metadata\n\n<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>6.1 Save metadata </b></p>\n</div>","metadata":{}},{"cell_type":"code","source":"ex_path = \"../input/rsna-2022-cervical-spine-fracture-detection/train_images/1.2.826.0.1.3680043.10001/101.dcm\"\ndcm_example = pydicom.dcmread(ex_path)\ndcm_example","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:12.646893Z","iopub.execute_input":"2022-09-04T07:56:12.647269Z","iopub.status.idle":"2022-09-04T07:56:12.678583Z","shell.execute_reply.started":"2022-09-04T07:56:12.647239Z","shell.execute_reply":"2022-09-04T07:56:12.677632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# From https://www.kaggle.com/code/andradaolteanu/rsna-fracture-detection-dicom-images-explore\ndef get_observation_data(path):\n    '''\n    Get information from the .dcm files\n    '''\n\n    dataset = pydicom.read_file(path)\n    \n    # Dictionary to store the information from the image\n    observation_data = {\n        \"Rows\" : dataset.get(\"Rows\"),\n        \"Columns\" : dataset.get(\"Columns\"),\n        \"SOPInstanceUID\" : dataset.get(\"SOPInstanceUID\"),\n        \"ContentDate\" : dataset.get(\"ContentDate\"),\n        \"SliceThickness\" : dataset.get(\"SliceThickness\"),\n        \"InstanceNumber\" : dataset.get(\"InstanceNumber\"),\n        \"ImagePositionPatient\" : dataset.get(\"ImagePositionPatient\"),\n        \"ImageOrientationPatient\" : dataset.get(\"ImageOrientationPatient\"),\n    }\n\n    # String columns\n    str_columns = [\"SOPInstanceUID\", \"ContentDate\", \n                   \"SliceThickness\", \"InstanceNumber\"]\n    for k in str_columns:\n        observation_data[k] = str(dataset.get(k)) if k in dataset else None\n\n    return observation_data\n\ndef get_metadata():\n    '''\n    Retrieves the desired metadata from the .dcm files and saves it into dataframe.\n    '''\n    \n    exceptions = 0\n    dicts = []\n\n    for k in tqdm(range(len(train_df))):\n        if (k % 100)==0:\n            print(f'Iteration: {k}')\n            \n        dt = train_df.iloc[k, :]\n\n        # Get all .dcm paths for this Instance\n        dcm_paths = glob(f\"{base_path}/train_images/{dt.StudyInstanceUID}/*\")\n\n        for path in dcm_paths:\n            try:\n                # Get datasets\n                dataset = get_observation_data(path)\n                dicts.append(dataset)\n            except Exception as e:\n                exceptions += 1\n                continue\n\n    # Convert into df\n    meta_train_data = pd.DataFrame(data=dicts, columns=md_example.keys())\n    \n    # Export information\n    meta_train_data.to_csv(\"meta_train.csv\", index=False)\n    \n    print(f\"Metadata created. Number of total fails: {exceptions}.\")\n","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:12.680611Z","iopub.execute_input":"2022-09-04T07:56:12.681221Z","iopub.status.idle":"2022-09-04T07:56:12.696627Z","shell.execute_reply.started":"2022-09-04T07:56:12.681175Z","shell.execute_reply":"2022-09-04T07:56:12.695181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Example\nmd_example = get_observation_data(ex_path)\npprint(md_example)","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:12.698478Z","iopub.execute_input":"2022-09-04T07:56:12.698943Z","iopub.status.idle":"2022-09-04T07:56:12.715172Z","shell.execute_reply.started":"2022-09-04T07:56:12.698892Z","shell.execute_reply":"2022-09-04T07:56:12.713961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create and save the metadata (~ 2 hours)\n#get_metadata()","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:12.719535Z","iopub.execute_input":"2022-09-04T07:56:12.720547Z","iopub.status.idle":"2022-09-04T07:56:12.728623Z","shell.execute_reply.started":"2022-09-04T07:56:12.720494Z","shell.execute_reply":"2022-09-04T07:56:12.727459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You can find the dataset containing the metadata [here](https://www.kaggle.com/datasets/samuelcortinhas/rsna-2022-spine-fracture-detection-metadata).","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>6.2 Explore metadata </b></p>\n</div>","metadata":{}},{"cell_type":"code","source":"# Read in saved metadata\nmeta_train = pd.read_csv(\"../input/meta-data/meta_train.csv\")\nmeta_train[\"StudyInstanceUID\"] = meta_train[\"SOPInstanceUID\"].apply(lambda x: \".\".join(x.split(\".\")[:-2]))\nprint('meta_train shape:', meta_train.shape)\nmeta_train.head(3)","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:12.732439Z","iopub.execute_input":"2022-09-04T07:56:12.73311Z","iopub.status.idle":"2022-09-04T07:56:17.132131Z","shell.execute_reply.started":"2022-09-04T07:56:12.733076Z","shell.execute_reply":"2022-09-04T07:56:17.13108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Image size\nmeta_train[\"ImageSize\"] = meta_train[\"Rows\"].astype(str) + \" x \" + meta_train[\"Columns\"].astype(str)\n\n# Plot image sizes\nplt.figure(figsize=(18, 5))\nax = sns.countplot(data=meta_train, x=\"ImageSize\")\nfor container in ax.containers:\n    ax.bar_label(container)\nplt.ylim([0,800000])\nplt.title('Image sizes in train images', fontsize=25, y=1.02)","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:17.139738Z","iopub.execute_input":"2022-09-04T07:56:17.140155Z","iopub.status.idle":"2022-09-04T07:56:19.840993Z","shell.execute_reply.started":"2022-09-04T07:56:17.140105Z","shell.execute_reply":"2022-09-04T07:56:19.839401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Almost all images have size 512x512. We should therefore resize the other images to size 512x512.","metadata":{}},{"cell_type":"markdown","source":"**Content Date**","metadata":{}},{"cell_type":"code","source":"# Unique values\nmeta_train['ContentDate'].unique()","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:19.843492Z","iopub.execute_input":"2022-09-04T07:56:19.843971Z","iopub.status.idle":"2022-09-04T07:56:19.858946Z","shell.execute_reply.started":"2022-09-04T07:56:19.843938Z","shell.execute_reply":"2022-09-04T07:56:19.857403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can drop this feature as it is constant.","metadata":{}},{"cell_type":"code","source":"meta_train.drop('ContentDate', axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:19.860544Z","iopub.execute_input":"2022-09-04T07:56:19.861707Z","iopub.status.idle":"2022-09-04T07:56:20.470115Z","shell.execute_reply.started":"2022-09-04T07:56:19.86166Z","shell.execute_reply":"2022-09-04T07:56:20.46622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Slice Thickness**","metadata":{}},{"cell_type":"code","source":"meta_train[\"SliceThickness\"].unique()","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:20.474046Z","iopub.execute_input":"2022-09-04T07:56:20.47624Z","iopub.status.idle":"2022-09-04T07:56:20.517152Z","shell.execute_reply.started":"2022-09-04T07:56:20.476128Z","shell.execute_reply":"2022-09-04T07:56:20.514468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot slice thickness\nplt.figure(figsize=(18, 5))\nax = sns.countplot(data=meta_train, x=\"SliceThickness\")\nfor container in ax.containers:\n    ax.bar_label(container)\nax.set_xticklabels(['0.488','0.500','0.600','0.60...2','0.625','0.664','0.670','0.75','0.800','0.900','1.000'])\nplt.ylim([0,420000])\nplt.title('Slice thickness distribution', fontsize=25, y=1.02)","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:20.523171Z","iopub.execute_input":"2022-09-04T07:56:20.525588Z","iopub.status.idle":"2022-09-04T07:56:20.946822Z","shell.execute_reply.started":"2022-09-04T07:56:20.525451Z","shell.execute_reply":"2022-09-04T07:56:20.945817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Number of slices per scan**","metadata":{}},{"cell_type":"code","source":"# Slice counts\nslice_counts = meta_train[\"StudyInstanceUID\"].value_counts().reset_index()\nslice_counts.columns = [\"StudyInstanceUID\", \"count\"]\n\n# Distribution of slices counts\nplt.figure(figsize=(18, 5))\nsns.histplot(data=slice_counts, x=\"count\", kde=True, bins=40)\nplt.title(\"Number of slices by scan\", size=25, y=1.02)\nplt.xlabel(\"Number of Slices\", size = 18)","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:20.947987Z","iopub.execute_input":"2022-09-04T07:56:20.948542Z","iopub.status.idle":"2022-09-04T07:56:21.357243Z","shell.execute_reply.started":"2022-09-04T07:56:20.948511Z","shell.execute_reply":"2022-09-04T07:56:21.355995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Image Position Patient**","metadata":{}},{"cell_type":"code","source":"# Extract x, y, z coordinates of position vector\nmeta_train['ImagePositionPatient_x'] = meta_train['ImagePositionPatient'].apply(lambda x: float(x.replace(',','').replace(']','').replace('[','').split()[0]))\nmeta_train['ImagePositionPatient_y'] = meta_train['ImagePositionPatient'].apply(lambda x: float(x.replace(',','').replace(']','').replace('[','').split()[1]))\nmeta_train['ImagePositionPatient_z'] = meta_train['ImagePositionPatient'].apply(lambda x: float(x.replace(',','').replace(']','').replace('[','').split()[2]))","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:21.358836Z","iopub.execute_input":"2022-09-04T07:56:21.359453Z","iopub.status.idle":"2022-09-04T07:56:24.253769Z","shell.execute_reply.started":"2022-09-04T07:56:21.359419Z","shell.execute_reply":"2022-09-04T07:56:24.252422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot position coordinates\nplt.figure(figsize=(20,5))\nplt.subplot(1,3,1)\nsns.histplot(meta_train['ImagePositionPatient_x'])\nplt.ylim([0,20000])\nplt.title('x-coordinate', fontsize=25, y=1.02)\n\nplt.subplot(1,3,2)\nsns.histplot(meta_train['ImagePositionPatient_y'], color='C1')\nplt.ylabel('')\nplt.ylim([0,20000])\nplt.title('y-coordinate', fontsize=25, y=1.02)\n\nplt.subplot(1,3,3)\nsns.histplot(meta_train['ImagePositionPatient_z'], color='C2')\nplt.ylabel('')\nplt.ylim([0,20000])\nplt.title('z-coordinate', fontsize=25, y=1.02)","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:24.255354Z","iopub.execute_input":"2022-09-04T07:56:24.255844Z","iopub.status.idle":"2022-09-04T07:56:29.045281Z","shell.execute_reply.started":"2022-09-04T07:56:24.255794Z","shell.execute_reply":"2022-09-04T07:56:29.044017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can use the 'ImagePositionPatient_z' feature to infer the vertebrae present in each slice.","metadata":{}},{"cell_type":"markdown","source":"**Image Orientation Patient**\n\nThis feature tells us coordinates for the position of the patient in the scanner. It could tell us if the images are slightly rotated for example.","metadata":{}},{"cell_type":"code","source":"meta_train['ImageOrientationPatient'].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:29.046762Z","iopub.execute_input":"2022-09-04T07:56:29.047104Z","iopub.status.idle":"2022-09-04T07:56:29.143581Z","shell.execute_reply.started":"2022-09-04T07:56:29.047073Z","shell.execute_reply":"2022-09-04T07:56:29.142507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"meta_train['ImageOrientationPatient'].unique()[:5]","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:29.145107Z","iopub.execute_input":"2022-09-04T07:56:29.145447Z","iopub.status.idle":"2022-09-04T07:56:29.238711Z","shell.execute_reply.started":"2022-09-04T07:56:29.145417Z","shell.execute_reply":"2022-09-04T07:56:29.237439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Clean meta data**","metadata":{}},{"cell_type":"code","source":"# Clean metadata\nmeta_train_clean = meta_train.drop(['SOPInstanceUID','ImagePositionPatient','ImageOrientationPatient','ImageSize'], axis=1)\nmeta_train_clean.rename(columns={\"Rows\": \"ImageHeight\", \"Columns\": \"ImageWidth\",\"InstanceNumber\": \"Slice\"}, inplace=True)\nmeta_train_clean = meta_train_clean[['StudyInstanceUID','Slice','ImageHeight','ImageWidth','SliceThickness','ImagePositionPatient_x','ImagePositionPatient_y','ImagePositionPatient_z']]\nmeta_train_clean.sort_values(by=['StudyInstanceUID','Slice'], inplace=True)\nmeta_train_clean.reset_index(drop=True, inplace=True)\n\n# Export information\n#meta_train_clean.to_csv(\"meta_train_clean.csv\", index=False)\n\n# Preview first few columns\nmeta_train_clean.head(3)","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:29.240245Z","iopub.execute_input":"2022-09-04T07:56:29.240822Z","iopub.status.idle":"2022-09-04T07:56:29.613075Z","shell.execute_reply.started":"2022-09-04T07:56:29.240778Z","shell.execute_reply":"2022-09-04T07:56:29.611968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 7. Bounding boxes\n\n<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b> 7.1 Box distributions </b></p>\n</div>\n\nWe are only given bounding boxes for a **subset** of the data. In particular, only **12%** of patients in the train set have any bounding box measurements.\n\nThis information is useful in telling us exactly where the fractures have occured. We could consider training an **object localisation** algorithm to provide bounding boxes for the whole train set. ","metadata":{}},{"cell_type":"code","source":"print(f'Patients with bounding box measurements: {train_bbox[\"StudyInstanceUID\"].nunique()} ({np.round(100*train_bbox[\"StudyInstanceUID\"].nunique()/len(train_df),1)} %)')","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:29.614436Z","iopub.execute_input":"2022-09-04T07:56:29.614936Z","iopub.status.idle":"2022-09-04T07:56:29.622352Z","shell.execute_reply.started":"2022-09-04T07:56:29.614905Z","shell.execute_reply":"2022-09-04T07:56:29.62113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# From https://www.kaggle.com/code/leventelippenszky/rsna-eda-dicom-segmentations-bboxes-3d-plot\ntrain_df_bbox = train_df[train_df[\"StudyInstanceUID\"].isin(train_bbox[\"StudyInstanceUID\"])]\n\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(16, 5))\nsns.countplot(x=\"patient_overall\", data=train_df_bbox, ax=ax1)\nax1.set_title(\"Fracture overall (patients with bboxes)\")\n\ntrain_df_bbox_melt = pd.melt(train_df_bbox, id_vars=[\"StudyInstanceUID\", \"patient_overall\"], var_name=\"cervical_vertebrae\", value_name=\"fracture\")\nsns.countplot(x=\"cervical_vertebrae\", hue=\"fracture\", data=train_df_bbox_melt, ax=ax2)\nax2.set_title(\"Fracture of cervical vertebrae (patients with bboxes)\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:29.624426Z","iopub.execute_input":"2022-09-04T07:56:29.624891Z","iopub.status.idle":"2022-09-04T07:56:30.008769Z","shell.execute_reply.started":"2022-09-04T07:56:29.624846Z","shell.execute_reply":"2022-09-04T07:56:30.007924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(16,5))\nsns.histplot(train_bbox[\"StudyInstanceUID\"].value_counts().values, kde=True, bins=40)\nplt.title('Number of slices with bounding boxes per patient')\nplt.xlabel('Number of bboxes')","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:30.00988Z","iopub.execute_input":"2022-09-04T07:56:30.010798Z","iopub.status.idle":"2022-09-04T07:56:30.379804Z","shell.execute_reply.started":"2022-09-04T07:56:30.010763Z","shell.execute_reply":"2022-09-04T07:56:30.378968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(16,6))\nplt.subplot(1,2,1)\nsns.scatterplot(data=train_bbox, x='x', y='y')\nplt.title('Anchor coordinates')\n\nplt.subplot(1,2,2)\nsns.scatterplot(data=train_bbox, x='width', y='height')\nplt.title('Width and heights')","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:30.380895Z","iopub.execute_input":"2022-09-04T07:56:30.382078Z","iopub.status.idle":"2022-09-04T07:56:30.887286Z","shell.execute_reply.started":"2022-09-04T07:56:30.382032Z","shell.execute_reply":"2022-09-04T07:56:30.886038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b> 7.2 Examples of fractures </b></p>\n</div>","metadata":{}},{"cell_type":"code","source":"def plot_fracture(slice_num,bbox_id,ax_id1,ax_id2):\n    file = pydicom.dcmread(f\"{base_path}/train_images/{bbox_id}/{slice_num}.dcm\")\n    img = apply_voi_lut(file.pixel_array, file)\n    info = train_bbox[(train_bbox['StudyInstanceUID']==bbox_id)&(train_bbox['slice_number']==slice_num)]\n    rect = patches.Rectangle((float(info.x), float(info.y)), float(info.width), float(info.height), linewidth=3, edgecolor='r', facecolor='none')\n\n    axes[ax_id1,ax_id2].imshow(img, cmap=\"bone\")\n    axes[ax_id1,ax_id2].add_patch(rect)\n    axes[ax_id1,ax_id2].set_title(f\"ID:{bbox_id}, Slice: {slice_num}\", fontsize=20, weight='bold',y=1.02)\n    axes[ax_id1,ax_id2].axis('off')\n\nfig, axes = plt.subplots(nrows=2, ncols=2, figsize=(24,24))\nplot_fracture(119,'1.2.826.0.1.3680043.25651',0,0)\nplot_fracture(156,'1.2.826.0.1.3680043.23817',0,1)\nplot_fracture(325,'1.2.826.0.1.3680043.12031',1,0)\nplot_fracture(151,'1.2.826.0.1.3680043.11899',1,1)","metadata":{"execution":{"iopub.status.busy":"2022-09-04T07:56:30.888349Z","iopub.execute_input":"2022-09-04T07:56:30.888682Z","iopub.status.idle":"2022-09-04T07:56:32.23638Z","shell.execute_reply.started":"2022-09-04T07:56:30.888652Z","shell.execute_reply":"2022-09-04T07:56:32.235005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>8. Modeling</b></p>\n</div>\n","metadata":{}},{"cell_type":"code","source":"# libjpeg & gdcm without internet access\n# src: https://www.kaggle.com/code/awsaf49/pydicom-conda-helper\n\n!wget 'https://anaconda.org/conda-forge/libjpeg-turbo/2.1.0/download/linux-64/libjpeg-turbo-2.1.0-h7f98852_0.tar.bz2' -q\n!wget 'https://anaconda.org/conda-forge/libgcc-ng/9.3.0/download/linux-64/libgcc-ng-9.3.0-h2828fa1_19.tar.bz2' -q\n!wget 'https://anaconda.org/conda-forge/gdcm/2.8.9/download/linux-64/gdcm-2.8.9-py37h500ead1_1.tar.bz2' -q\n!wget 'https://anaconda.org/conda-forge/conda/4.10.1/download/linux-64/conda-4.10.1-py37h89c1867_0.tar.bz2' -q\n!wget 'https://anaconda.org/conda-forge/certifi/2020.12.5/download/linux-64/certifi-2020.12.5-py37h89c1867_1.tar.bz2' -q\n!wget 'https://anaconda.org/conda-forge/openssl/1.1.1k/download/linux-64/openssl-1.1.1k-h7f98852_0.tar.bz2' -q\n\n!conda install 'libjpeg-turbo-2.1.0-h7f98852_0.tar.bz2' -c conda-forge -y\n!conda install 'libgcc-ng-9.3.0-h2828fa1_19.tar.bz2' -c conda-forge -y\n!conda install 'gdcm-2.8.9-py37h500ead1_1.tar.bz2' -c conda-forge -y\n!conda install 'conda-4.10.1-py37h89c1867_0.tar.bz2' -c conda-forge -y\n!conda install 'certifi-2020.12.5-py37h89c1867_1.tar.bz2' -c conda-forge -y\n!conda install 'openssl-1.1.1k-h7f98852_0.tar.bz2' -c conda-forge -y\n\n# MONAI 3D model\n!pip install -q monai\n!pip install -q git+https://github.com/ildoonet/pytorch-gradual-warmup-lr.git","metadata":{"execution":{"iopub.status.busy":"2022-09-04T16:12:13.314766Z","iopub.execute_input":"2022-09-04T16:12:13.315306Z","iopub.status.idle":"2022-09-04T16:14:11.832743Z","shell.execute_reply.started":"2022-09-04T16:12:13.315205Z","shell.execute_reply":"2022-09-04T16:14:11.830824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Libraries\nimport os\nimport re\nimport gc\nimport cv2\nimport wandb\nfrom PIL import Image\nimport random\nimport math\nimport shutil\nfrom glob import glob\nfrom tqdm import tqdm\nfrom pprint import pprint\nfrom time import time\nimport warnings\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib as mpl\nfrom matplotlib import cm\nimport matplotlib.patches as patches\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nfrom matplotlib.offsetbox import AnnotationBbox, OffsetImage\nfrom matplotlib.colors import ListedColormap, LinearSegmentedColormap\nfrom matplotlib.patches import Rectangle\nfrom IPython.display import display_html\nplt.rcParams.update({'font.size': 16})\n\n# Environment check\nwarnings.filterwarnings(\"ignore\")\nos.environ[\"WANDB_SILENT\"] = \"true\"\nCONFIG = {'competition': 'RSNA_SpineFructure', '_wandb_kernel': 'aot'}\n\n# Custom colors\nclass clr:\n    S = '\\033[1m' + '\\033[94m'\n    E = '\\033[0m'\n    \nmy_colors = [\"#5EAFD9\", \"#449DD1\", \"#3977BB\", \n             \"#2D51A5\", \"#5C4C8F\", \"#8B4679\",\n             \"#C53D4C\", \"#E23836\", \"#FF4633\", \"#FF5746\"]\nCMAP1 = ListedColormap(my_colors)\n\nprint(clr.S+\"Notebook Color Schemes:\"+clr.E)\nsns.palplot(sns.color_palette(my_colors))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-04T16:14:24.68202Z","iopub.execute_input":"2022-09-04T16:14:24.682458Z","iopub.status.idle":"2022-09-04T16:14:26.57523Z","shell.execute_reply.started":"2022-09-04T16:14:24.682423Z","shell.execute_reply":"2022-09-04T16:14:26.573595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# PyTorch\nimport torch\nfrom torch.utils.data import TensorDataset, DataLoader, Dataset\nimport torch.nn as nn\nimport torch.nn.functional as F\nimport torch.optim as optim\nfrom torch.optim import lr_scheduler\nfrom torch.utils.data.sampler import SubsetRandomSampler, RandomSampler, SequentialSampler\nfrom torch.optim.lr_scheduler import StepLR, ReduceLROnPlateau, CosineAnnealingLR\nimport torchvision\nimport torchvision.transforms as transforms\nfrom warmup_scheduler import GradualWarmupScheduler\nimport albumentations\n\nfrom sklearn.model_selection import GroupKFold, train_test_split, StratifiedKFold\nfrom sklearn.metrics import roc_auc_score, cohen_kappa_score, confusion_matrix\n\n# MONAI 3D\nfrom monai.transforms import Randomizable, apply_transform\nfrom monai.transforms import Compose, Resize, ScaleIntensity, ToTensor, RandAffine\nfrom monai.networks.nets import densenet","metadata":{"execution":{"iopub.status.busy":"2022-09-04T16:16:59.298413Z","iopub.execute_input":"2022-09-04T16:16:59.299705Z","iopub.status.idle":"2022-09-04T16:17:04.400025Z","shell.execute_reply.started":"2022-09-04T16:16:59.299632Z","shell.execute_reply":"2022-09-04T16:17:04.39855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**[🐝 W&B Fork & Run](\"https://wandb.ai/site\")**","metadata":{}},{"cell_type":"code","source":"# 🐝 Secrets\n!pip install wandb\n\nimport wandb\nwandb.login()","metadata":{"execution":{"iopub.status.busy":"2022-09-04T16:28:13.938827Z","iopub.execute_input":"2022-09-04T16:28:13.939259Z","iopub.status.idle":"2022-09-04T16:30:18.439911Z","shell.execute_reply.started":"2022-09-04T16:28:13.939224Z","shell.execute_reply":"2022-09-04T16:30:18.43851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# run = wandb.init(project=\"vinbigdata\", entity=\"paul-ndirangu\")","metadata":{"execution":{"iopub.status.busy":"2022-09-04T16:31:07.696714Z","iopub.execute_input":"2022-09-04T16:31:07.69715Z","iopub.status.idle":"2022-09-04T16:31:07.702839Z","shell.execute_reply.started":"2022-09-04T16:31:07.697115Z","shell.execute_reply":"2022-09-04T16:31:07.701616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# try:\n#     import pylibjpeg\n# except:\n#     # Offline dependencies:\n#     !mkdir -p /root/.cache/torch/hub/checkpoints/\n#     !cp ../input/rsna-2022-whl/efficientnet_v2_s-dd5fe13b.pth  /root/.cache/torch/hub/checkpoints/\n\n#     !pip install /kaggle/input/rsna-2022-whl/{pydicom-2.3.0-py3-none-any.whl,pylibjpeg-1.4.0-py3-none-any.whl,python_gdcm-3.0.15-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl}\n#     !pip install /kaggle/input/rsna-2022-whl/{torch-1.12.1-cp37-cp37m-manylinux1_x86_64.whl,torchvision-0.13.1-cp37-cp37m-manylinux1_x86_64.whl}\n\n","metadata":{"execution":{"iopub.status.busy":"2022-09-04T16:48:18.761133Z","iopub.execute_input":"2022-09-04T16:48:18.761539Z","iopub.status.idle":"2022-09-04T16:48:18.767521Z","shell.execute_reply.started":"2022-09-04T16:48:18.761508Z","shell.execute_reply":"2022-09-04T16:48:18.766514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#9e4c4c;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>Research</b></p>\n</div>\n\nLet's look at the what the **image data** from the dcm files look like.\n\n\n> Research blog links\n\n>---\n\n[EDA image classification_1](https://towardsdatascience.com/exploratory-data-analysis-ideas-for-image-classification-d3fc6bbfb2d2#:~:text=Exploratory%20data%20analysis%20comprises%20of,of%20predictors%20across%20different%20classes.)\n<br>\n[EDA image classification_2](https://medium.com/geekculture/eda-for-image-classification-dcada9f2567a)\n<br>\n[EDA for images kaggle](https://www.kaggle.com/code/jpmiller/basic-eda-with-images)\n<br>\n[image-classification-tips-and-tricks-from-13-kaggle-competitions](https://neptune.ai/blog/image-classification-tips-and-tricks-from-13-kaggle-competitions)\n\n### Notebook Copied With Edits from:\n> [SAMUEL CORTINHAS](https://www.kaggle.com/code/samuelcortinhas/rsna-fracture-detection-in-depth-eda)\n\n> [_LEV_LIPINSKI](https://www.kaggle.com/code/leventelippenszky/rsna-eda-dicom-segmentations-bboxes-3d-plot)","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}