{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **EDA Simplified: RSNA Cervical Spine Fracture Detection**","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"## Introduction\nFrom [UW-M GI Image Tract Segmentation](https://www.kaggle.com/code/dinowun/eda-simplified-uwm-gi-tract-segmentation-w-w-b) to [HuBMAP's Hacking the Human Body](https://www.kaggle.com/code/dinowun/eda-simplified-hubmap-hpa-organ-segmentation?scriptVersionId=100803535) and then [Mayo Clinic Strip AI](https://www.kaggle.com/code/dinowun/eda-simplified-mayo-clinic-strip-ai-b), we've covered most data over segmenting gastrointestinal tracts with Weights & Biases, finding the tissues of 5 human body parts, and analyzing the images of the large artery atherosclerosis and cardioembrolic strokes. But now, we've gotten to this competition, that we're identifying the cervical spine fractures from scans. Because of this, there were **overwhelmingly** 1.5 million spine fracture cases that occur annually in the United States, resulting in over 17,730 spinal cord injuries. In fact, the most common site of spine fracture is the cervical spine. Furthermore, fractures can be more difficult to detect on imaging due to superimposed degenerative disease and osteoporosis. Imaging diagnosis of adult spine fractures is now almost exclusively performed with computed tomography (CT) instead of radiographs (x-rays). So, the whole thing over this is to spot a single fracture from a CT scan, and on the top of that, a blend of our EDA. With all that being said, let's dive deep to scavenge over a fracture on a cervical spine!\n\nBefore we proceed, be sure to read our previous \"EDA Simplified\" series about human body and medical stuff:\n\n* UW-M GI Tract Segmentation w/ W&B: https://www.kaggle.com/code/dinowun/eda-simplified-uwm-gi-tract-segmentation-w-w-b\n* HuBMAP + HPA Organ Segmentation: https://www.kaggle.com/code/dinowun/eda-simplified-hubmap-hpa-organ-segmentation\n* Mayo Clinic STRIP AI: https://www.kaggle.com/code/dinowun/eda-simplified-mayo-clinic-strip-ai","metadata":{}},{"cell_type":"markdown","source":"## Imports & Setup\nIn order to get started with EDA on the RSNA competition, we import the three modules for data science, linear algebra, and machine learning, which is numpy as np, pandas as pd, and cv2 (OpenCV). Next, we import the modules used for plotting and reading images, which is plotly with the graph_objects and express submodules as go and px and matplotlib with the pyplot submodule as plt. We then import the important modules used for the utility of miscellaneous stuff which is the os module and glob from the glob module for operating system stuff, the nibabel as nib for neuroimaging, and tqdm from the tqdm module for progress stuff. Finally, we import the pydicom module and apply_voi_lut from pydicom module again but with the pixel_data_handlers submodule along with the util submodule for reading out dcm image files unlike tiff, jpg, or png files in the previous competition. However, some required installation for the pydicom modules before importing it, in which the data section stated that:\n\n> The DICOM image files are ≤ 1 mm slice thickness, axial orientation, and bone kernel. Note that some of the DICOM files are JPEG compressed. You may require additional resources to read the pixel array of these files, such as GDCM and pylibjpeg.\n\nThus, [we noticed some failed runs after a lot of attempts over a single commit](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/341412), so, we installed the python-gdcm, pylibjpeg, pylibjpeg-libjpeg, and pydicom via pip when the internet option is on.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport cv2\n\nimport matplotlib.pyplot as plt\nimport plotly.graph_objects as go\nimport plotly.express as px\n\nimport os\nfrom glob import glob\nimport nibabel as nib\nfrom tqdm import tqdm\n\n# Be sure to install some dicom files first! \n! pip install python-gdcm\n! pip install pylibjpeg pylibjpeg-libjpeg pydicom\n\nimport pydicom\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:35:57.347278Z","iopub.execute_input":"2022-10-07T10:35:57.347696Z","iopub.status.idle":"2022-10-07T10:36:20.984056Z","shell.execute_reply.started":"2022-10-07T10:35:57.347661Z","shell.execute_reply":"2022-10-07T10:36:20.98264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After importing the modules, let's create two dataframes! For the first dataframe, we define the train_df dataframe to the the pd module with the read_csv function to read out the path leading to train.csv file. And for the second dataframe, we define another dataframe, train_bound_box_df, to the pd module with the read_csv function again, but this time, the train_bounding_boxes dataframe from the same path. And finally, let's display the first rows of the two dataframes separately with the head function!","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(\"../input/rsna-2022-cervical-spine-fracture-detection/train.csv\")\ntrain_bound_box_df = pd.read_csv(\"../input/rsna-2022-cervical-spine-fracture-detection/train_bounding_boxes.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:20.986885Z","iopub.execute_input":"2022-10-07T10:36:20.987358Z","iopub.status.idle":"2022-10-07T10:36:21.020006Z","shell.execute_reply.started":"2022-10-07T10:36:20.987316Z","shell.execute_reply":"2022-10-07T10:36:21.018842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.022238Z","iopub.execute_input":"2022-10-07T10:36:21.022676Z","iopub.status.idle":"2022-10-07T10:36:21.038813Z","shell.execute_reply.started":"2022-10-07T10:36:21.022641Z","shell.execute_reply":"2022-10-07T10:36:21.037747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_bound_box_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.040015Z","iopub.execute_input":"2022-10-07T10:36:21.040516Z","iopub.status.idle":"2022-10-07T10:36:21.063062Z","shell.execute_reply.started":"2022-10-07T10:36:21.040384Z","shell.execute_reply":"2022-10-07T10:36:21.061448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we generate and preview the two dataframes, let's move on to analyzing the basics of our two dataframes!","metadata":{}},{"cell_type":"markdown","source":"## Chapter 1: The Basic Analysis of the Two Dataframes\nTo get blasting through it, we first find out how many data entities are there in each of the two dataframes (train_df and train_bound_box_df) with the len function.","metadata":{}},{"cell_type":"code","source":"len(train_df)","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.066106Z","iopub.execute_input":"2022-10-07T10:36:21.06658Z","iopub.status.idle":"2022-10-07T10:36:21.077052Z","shell.execute_reply.started":"2022-10-07T10:36:21.066542Z","shell.execute_reply":"2022-10-07T10:36:21.075559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(train_bound_box_df)","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.078578Z","iopub.execute_input":"2022-10-07T10:36:21.07895Z","iopub.status.idle":"2022-10-07T10:36:21.091523Z","shell.execute_reply.started":"2022-10-07T10:36:21.078919Z","shell.execute_reply":"2022-10-07T10:36:21.090087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see, the data in each two dataframes is mostly big. The number of data entities overall in the train_df dataframe is 2019 while the number of them in the train_bound_box_df dataframe is 7217.","metadata":{}},{"cell_type":"markdown","source":"Now let's find out whether there are NaN values in the two dataframes. To do that, we plug in the isna function to determine how many NaN values are there in each one, thus plugging in the sum function to count how many NaN values in all.","metadata":{}},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.092626Z","iopub.execute_input":"2022-10-07T10:36:21.09307Z","iopub.status.idle":"2022-10-07T10:36:21.10953Z","shell.execute_reply.started":"2022-10-07T10:36:21.092948Z","shell.execute_reply":"2022-10-07T10:36:21.108491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_bound_box_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.111156Z","iopub.execute_input":"2022-10-07T10:36:21.111934Z","iopub.status.idle":"2022-10-07T10:36:21.124964Z","shell.execute_reply.started":"2022-10-07T10:36:21.111869Z","shell.execute_reply":"2022-10-07T10:36:21.12372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Perfect! There are no NaN values shown in both dataframes! Each data count of nan values on both train_df and train_bound_box_df dataframes showed 0 of them, meaning that the two dataframes is clear from NaN values.","metadata":{}},{"cell_type":"markdown","source":"Finally, let's generate and observe the shape of the two dataframes with the shape attribute!","metadata":{}},{"cell_type":"code","source":"train_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.126618Z","iopub.execute_input":"2022-10-07T10:36:21.127268Z","iopub.status.idle":"2022-10-07T10:36:21.14165Z","shell.execute_reply.started":"2022-10-07T10:36:21.127227Z","shell.execute_reply":"2022-10-07T10:36:21.139237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_bound_box_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.144376Z","iopub.execute_input":"2022-10-07T10:36:21.148223Z","iopub.status.idle":"2022-10-07T10:36:21.158019Z","shell.execute_reply.started":"2022-10-07T10:36:21.148128Z","shell.execute_reply":"2022-10-07T10:36:21.156881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After running the last two cells of this chapter, we observed that the shape of the train_df dataframe contained 9 columns and 2019 rows of data while the train_bound_box_df contained 6 columns and 7217 rows of data. Now that we've finished going through the basic analysis of the two dataframes, let's go on to analyzing two dataframes on each chapter!","metadata":{}},{"cell_type":"markdown","source":"## Chapter 2: The Analysis of train_df\nBefore we begin our analysis of our train_df dataframe, here's what facts we need to know: the cervical spine was counted from 1 to 7 from top to bottom. This cervical spine structuring in the train_df dataframe indicates that C1 to C7 may or may not be fractured.\n\nAnd with that, let's plot over the data from the train_df dataframe!","metadata":{}},{"cell_type":"markdown","source":"First, let's find out how many counts of data in the patient_overall data! First, we define the patient_overall variable to the train_df dataframe with the patient_overall data index and count the values of it with the value_counts variable, setting the normalize (normalizing the data) parameter to True. We then define the fig variable figure to the px module with the pie function, setting the patient_overall variable inside, thus setting the values (values for the pie chart) to the patient_overall variable's values with the values attribute, the names (names for the pie chart) attribute to the patient_overall variable's index with the index attribute, and the title (title name for pie figure) to \"Counts of Patient Overall\". Finally, we show the fig variable figure with the show function.","metadata":{}},{"cell_type":"code","source":"patient_overall = train_df[\"patient_overall\"].value_counts(normalize=True)\n\nfig = px.pie(patient_overall, values=patient_overall.values, labels=patient_overall.index, title=\"Counts of Patient Overall\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.159385Z","iopub.execute_input":"2022-10-07T10:36:21.159889Z","iopub.status.idle":"2022-10-07T10:36:21.225578Z","shell.execute_reply.started":"2022-10-07T10:36:21.159831Z","shell.execute_reply":"2022-10-07T10:36:21.224301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Though our pie graph has no labels about the sorted counts of the patient_overall data, we see that there are 52.4% of patients didn't have their cervical spine fractured (which is indicated as 0) and the other 47.6% of patients did have their cervical spine fractured (which is indicated as 1).","metadata":{}},{"cell_type":"markdown","source":"And for sure, let's create 7 bar charts based on the segments of the cervical spines! Per each data from the structure of C1 to C7, We define 7 variables, cervical_1 to cervical_7, to the train_df's C1 to C7 data index and count their values with the value_counts function, setting the normalize parameter to True also. And finally for each of them, we use the px module with the bar function, setting the C1 to C7 variables as the dataframe input, thus setting the x parameter to the index of the 7 variables with the index attribute and the y parameter to the values of the 7 variables with the values attribute. We then update our x and y axes of our figure with the update_xaxes and update_yaxes functions, setting the title parameter to each of the x and y labels as \"boolean\" and \"counts overall\".\nFinally, we show 7 figures in each fig variable with the show function.","metadata":{}},{"cell_type":"code","source":"c1 = train_df[\"C1\"].value_counts(normalize=True)\n\nfig = px.bar(c1, x=c1.index, y=c1.values)\nfig.update_xaxes(title=\"boolean\")\nfig.update_yaxes(title=\"counts overall\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.227152Z","iopub.execute_input":"2022-10-07T10:36:21.227746Z","iopub.status.idle":"2022-10-07T10:36:21.295902Z","shell.execute_reply.started":"2022-10-07T10:36:21.227708Z","shell.execute_reply":"2022-10-07T10:36:21.294634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"c2 = train_df[\"C2\"].value_counts(normalize=True)\n\nfig = px.bar(c2, x=c2.index, y=c2.values)\nfig.update_xaxes(title=\"boolean\")\nfig.update_yaxes(title=\"counts overall\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.297687Z","iopub.execute_input":"2022-10-07T10:36:21.298333Z","iopub.status.idle":"2022-10-07T10:36:21.37728Z","shell.execute_reply.started":"2022-10-07T10:36:21.298161Z","shell.execute_reply":"2022-10-07T10:36:21.376273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"c3 = train_df[\"C3\"].value_counts(normalize=True)\n\nfig = px.bar(c3, x=c3.index, y=c3.values)\nfig.update_xaxes(title=\"boolean\")\nfig.update_yaxes(title=\"counts overall\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.382565Z","iopub.execute_input":"2022-10-07T10:36:21.383655Z","iopub.status.idle":"2022-10-07T10:36:21.452441Z","shell.execute_reply.started":"2022-10-07T10:36:21.3836Z","shell.execute_reply":"2022-10-07T10:36:21.451183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"c4 = train_df[\"C4\"].value_counts(normalize=True)\n\nfig = px.bar(c4, x=c4.index, y=c4.values)\nfig.update_xaxes(title=\"boolean\")\nfig.update_yaxes(title=\"counts overall\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.454351Z","iopub.execute_input":"2022-10-07T10:36:21.455602Z","iopub.status.idle":"2022-10-07T10:36:21.52096Z","shell.execute_reply.started":"2022-10-07T10:36:21.455555Z","shell.execute_reply":"2022-10-07T10:36:21.5196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"c5 = train_df[\"C5\"].value_counts(normalize=True)\n\nfig = px.bar(c5, x=c5.index, y=c5.values)\nfig.update_xaxes(title=\"boolean\")\nfig.update_yaxes(title=\"counts overall\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.522727Z","iopub.execute_input":"2022-10-07T10:36:21.523198Z","iopub.status.idle":"2022-10-07T10:36:21.594288Z","shell.execute_reply.started":"2022-10-07T10:36:21.523149Z","shell.execute_reply":"2022-10-07T10:36:21.593278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"c6 = train_df[\"C6\"].value_counts(normalize=True)\n\nfig = px.bar(c6, x=c6.index, y=c6.values)\nfig.update_xaxes(title=\"boolean\")\nfig.update_yaxes(title=\"counts overall\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.596045Z","iopub.execute_input":"2022-10-07T10:36:21.59646Z","iopub.status.idle":"2022-10-07T10:36:21.668048Z","shell.execute_reply.started":"2022-10-07T10:36:21.596424Z","shell.execute_reply":"2022-10-07T10:36:21.666614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"c7 = train_df[\"C7\"].value_counts(normalize=True)\n\nfig = px.bar(c7, x=c7.index, y=c7.values)\nfig.update_xaxes(title=\"boolean\")\nfig.update_yaxes(title=\"counts overall\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.671109Z","iopub.execute_input":"2022-10-07T10:36:21.672255Z","iopub.status.idle":"2022-10-07T10:36:21.740903Z","shell.execute_reply.started":"2022-10-07T10:36:21.672193Z","shell.execute_reply":"2022-10-07T10:36:21.739634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After running 7 code cells, We see that from C1 to C7, most of them weren't fractured while some of them were fractured. But, if we examine the train_df dataframe closely, we see that one or more sections from C1 to C7 were fractured. So, let's go on to our next chapter of analyzing the bounding boxes in another dataframe, which is train_bound_box_df.","metadata":{}},{"cell_type":"markdown","source":"## Chapter 3: The Analysis of train_bound_box_df\nSince we're on chapter 2 now, let's start analyzing the x values in a histogram based out of the data in slice_number! First, we define the fig variable figure to the px module with the histogram function, setting the train_bound_box_df dataframe as the dataframe input, and then setting the x (x-axis) parameter to slice_number as the data input, along with the y (y-axis) parameter to x as the data input thus the marginal (distribution visualization) parameter to violin. Finally, we display the fig variable figure with the show function.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_bound_box_df, x=\"slice_number\", y=\"x\", marginal=\"violin\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.743167Z","iopub.execute_input":"2022-10-07T10:36:21.743698Z","iopub.status.idle":"2022-10-07T10:36:21.83871Z","shell.execute_reply.started":"2022-10-07T10:36:21.743648Z","shell.execute_reply":"2022-10-07T10:36:21.837158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After running this code above and displaying the histogram chart, we analyzed that the highest sum of x data by slice_number is between 140 to 149, with their data counted as approximately 89.4k, while the least is 480 to 489 with roughly 528 data entities of it. Furthermore, The minimum observation is 26, along with the first and third quartile, is 128 and 234, the median is 173, and the max is 483.","metadata":{}},{"cell_type":"markdown","source":"Then, we are going to analyze the y data values by slice_number in a histogram! It's like the same as what we did on our previous analysis on the x data values, but we set the y (y-axis) parameter to the y data index.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_bound_box_df, x=\"slice_number\", y=\"y\", marginal=\"violin\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.840581Z","iopub.execute_input":"2022-10-07T10:36:21.841074Z","iopub.status.idle":"2022-10-07T10:36:21.9368Z","shell.execute_reply.started":"2022-10-07T10:36:21.841026Z","shell.execute_reply":"2022-10-07T10:36:21.935221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Like (and unlike) the previous graph consisting of our analysis over the x data by the slice_number data, we noticed that the highest sum of y data by the slice_number data is 140 to 149 while the lowest is 480 to 489. Furthermore, the median, min, max, 1st, and 3rd interquartile stayed the same as the previous graph above.","metadata":{}},{"cell_type":"markdown","source":"Now, let's correlate the x data by y data, in a scatter plot! All we need to do is defining the fig variable to the px module with the scatter function to create a scatter plot graph figure, setting the train_bound_box_df dataframe as the dataframe input, along with the x (x-axis) parameter to x and y (y-axis) parameter to y. Finally, we show the fig variable figure.","metadata":{}},{"cell_type":"code","source":"fig = px.scatter(train_bound_box_df, x=\"x\", y=\"y\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:21.939016Z","iopub.execute_input":"2022-10-07T10:36:21.939626Z","iopub.status.idle":"2022-10-07T10:36:22.008633Z","shell.execute_reply.started":"2022-10-07T10:36:21.939571Z","shell.execute_reply":"2022-10-07T10:36:22.00747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After compiling this code cell above, we see that a large number of plots in this figure were clamped up together that we need to zoom in with Plotly's \"zoom-in\" option to see the plots closely.","metadata":{}},{"cell_type":"markdown","source":"Now let's find the overall distribution of the width data in the train_bound_box_df dataframe! To do this, we are creating a histogram figure in Plotly by defining the fig variable to the px module with the histogram function, setting the train_bound_box_df dataframe inside of it, along with the x parameter set to slice_number, the y parameter to width, and the marginal parameter to violin. Finally, we show out our fig figure with the show function.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_bound_box_df, x=\"slice_number\", y=\"width\", marginal=\"violin\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:22.010828Z","iopub.execute_input":"2022-10-07T10:36:22.011373Z","iopub.status.idle":"2022-10-07T10:36:22.103825Z","shell.execute_reply.started":"2022-10-07T10:36:22.011319Z","shell.execute_reply":"2022-10-07T10:36:22.102321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After compiling this cell to display that graph above, we see that the highest sum of the width data by the slice_number data is 170 to 179 with 42.57k pverall data entities rounded while the lowest sum of the width data by the slice_number data is 480 to 489 with 464 overall entities rounded. Furthermore, the min data is 26, the first and third quantile is 128 and 234, the median is 173, the upper fence is 393, and the max is 483.","metadata":{}},{"cell_type":"markdown","source":"Now let's find the height data by slice_number data in a histogram! It's like the same as what we did to analyzing the width data, but we set up the y (y-axis) parameter to height inside the histogram function from the px module.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_bound_box_df, x=\"slice_number\", y=\"height\", marginal=\"violin\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:22.10594Z","iopub.execute_input":"2022-10-07T10:36:22.106337Z","iopub.status.idle":"2022-10-07T10:36:22.19792Z","shell.execute_reply.started":"2022-10-07T10:36:22.1063Z","shell.execute_reply":"2022-10-07T10:36:22.196368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Unlike the previous graph visualizing the distributed width data in the train_bound_box, the highest sum of the height data by slice_number is 150 to 159 with about 35.5 thousand entities of it. However, the least sum of this graph overall is similar to the histogram graph consisting the sum of width, in which it is 480 to 489 though this graph contained about 427 entities. Furthermore, the min, q1, q3, median, upper fence, and max in this height data distribution is same as the width data distribution.","metadata":{}},{"cell_type":"markdown","source":"Now let's create a scatter plot based on the width and height data! To do that, we define the fig variable to the px module with the scatter function to indicate that we're going to create a scatter plot figure, setting the train_bound_box_df dataframe as the data input, the x (x-axis) parameter to the width data, and the y (y-axis) parameter to the height data. Finally, we display our fig figure variable with the show function!","metadata":{}},{"cell_type":"code","source":"fig = px.scatter(train_bound_box_df, x=\"width\", y=\"height\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:22.199782Z","iopub.execute_input":"2022-10-07T10:36:22.200319Z","iopub.status.idle":"2022-10-07T10:36:22.267972Z","shell.execute_reply.started":"2022-10-07T10:36:22.200268Z","shell.execute_reply":"2022-10-07T10:36:22.266602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What we saw on this graph is that the plots are scattered from together to separately as the width data increases. However, there are some outliers in this graph, thus having the plots showing the downward and upward slope.","metadata":{}},{"cell_type":"markdown","source":"And for all the data in the train_bound_box_df dataframe has been all visualized, let's go onto visualizing the cervical spine images!","metadata":{}},{"cell_type":"markdown","source":"## Chapter 4: Cervical Spine Image Visualization\nNow here's the fun and final part we're going into, let's visualize over the images given in this competition! First, let's define the segmentation_masks variable to the sorted function to sort the values of the glob function to return all the file paths of the RSNA competition directory leading to the nii files in the segmentations folder that match a specific pattern. Next, we print out the number of entities with the len function to the segmentation_masks variable.","metadata":{}},{"cell_type":"code","source":"segmentation_masks = sorted(\n    glob(\n        \"../input/rsna-2022-cervical-spine-fracture-detection/segmentations/*.nii\"\n    )\n)\n\nlen(segmentation_masks)","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:22.270008Z","iopub.execute_input":"2022-10-07T10:36:22.270725Z","iopub.status.idle":"2022-10-07T10:36:22.279817Z","shell.execute_reply.started":"2022-10-07T10:36:22.270684Z","shell.execute_reply":"2022-10-07T10:36:22.278467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After the part where we find the number of data in the segmentation_masks variable with the len function, we see that there are 87 data entities in the segmentation folder.","metadata":{}},{"cell_type":"markdown","source":"Next, we define a function, load_nibabel, that contains the path variable input. And inside of this function, we return the nib module with the load function to load the path variable function input and then get the data from it with the get_fdata function method.","metadata":{}},{"cell_type":"code","source":"def load_nibabel(path):\n    return nib.load(path).get_fdata()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:22.281084Z","iopub.execute_input":"2022-10-07T10:36:22.281483Z","iopub.status.idle":"2022-10-07T10:36:22.290384Z","shell.execute_reply.started":"2022-10-07T10:36:22.281449Z","shell.execute_reply":"2022-10-07T10:36:22.28892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After defining the load_nibabel function, we also define additional two functions, multi_dim_plots with the multi_dim_array_input and id parameter along with the num_slices parameter that is set to 64, and plot_batches with the samples parameter and the cmap option set to none.\n\nFor the multi_dim_plots function, we define the fig variable to the plt module with the figure function to setup our new figure, setting the figsize parameter to 30 by 30. And after that, we setup our title figure with the title function from the plt module, containing the formatted string of \"Plotting first {} slices of {}\", with the num_slices and id parameters inside, thus setting the fontdict parameter to a dictionary that has the fontsize key set to 20. We then eliminate the x and y ticks by defining an empty list inside the xticks and yticks functions under the plt module. And moving on to counting the images, we define the xy variable to the int form (with the int function) of the square root of the num_slices parameter provided by the np module's sqrt function. Furthermore, a for loop is defined, looping the i variable to the range (with the range function) of the num_slices parameter. Inside the for loop, the ax was defined to the fig variable figure that has the add_subplot function to add the subplots of the figure, setting the two inputs as all of xy variables, along with the i variable that is incremented by 1. We then use the imshow function from the plt module to show the images of the multi_dim_array_input parameter with the slice_index of ... (no methods), with each step of the num_slices parameter thus having another slice index of no method (...) and the step of the i variable. For the finishing touch of the for loop, we turn off the axis with axis function from the plt module again, containing the string, \"off\". Finally, outside of the for loop inside the function, we show our figure with the plt module with the show function.\n\nMeanwhile in the plot_batches function, we define the fig variable to what we did while we were in multi_dim_plots function. Next, we define the xy function just like how we did in the multi_dim_plots function but this time inside the int function of converting an entity to int, we install the the shape attribute with the slice index of 0 to the samples parameter inside to find the first value of it. Now, just like the same as the multi_dim_plots function, we create a for loop, but this time, looping the i and sample variables to enumerating the samples parameter with the enumerate function. Inside the for loop, we test out the code inside the for loop with the try statement. And inside of this, we define the ax parameter to the subplot function for creating a subplot, setting to what we did when we were in the multi_dim_plots function. And looking for the except statement, the ax variable is defined to similar of what it was defined to for the try statement, but the first two inputs were filled by 8. And outside the try and except statements, the plt module with the axis parameter set the axis to \"off\" and then show out the images of the sample variable, setting the cmap parameter to the cmap function parameter with the imshow function. Finally, we show off the figure with the show function provided by the plt module.","metadata":{}},{"cell_type":"code","source":"def multi_dim_plots(multi_dim_array_input, id, num_slices=64):\n    fig = plt.figure(figsize=(30,30))\n    plt.title(f\"Plotting First {num_slices} Slices of {id}\", fontdict={'fontsize': 20})\n    plt.yticks([])\n    plt.xticks([])\n    \n    xy = int(np.sqrt(num_slices))\n    for i in range(num_slices):\n        ax = fig.add_subplot(xy, xy, i + 1)\n        plt.imshow(multi_dim_array_input[..., :num_slices][..., i])\n        plt.axis(\"off\")\n    plt.show()\n    \ndef plot_batches(samples, cmap=None):\n    plt.figure(figsize=(30,30))\n    \n    xy = int(np.sqrt(samples.shape[0]))\n    for i, sample in enumerate(samples):\n        try:\n            ax = plt.subplot(xy, xy, i+1)\n        except:\n            ax = plt.subplot(8, 8, i+1)\n        plt.axis(\"off\")\n        plt.imshow(sample, cmap=cmap)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:22.292237Z","iopub.execute_input":"2022-10-07T10:36:22.292735Z","iopub.status.idle":"2022-10-07T10:36:22.30518Z","shell.execute_reply.started":"2022-10-07T10:36:22.292687Z","shell.execute_reply":"2022-10-07T10:36:22.303949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now here's the best part of this chapter after tirelessly setup the two functions, let's display off the images of the segmentation masks! All we need to do is to create a for loop, looping the msk variable to the sample function to sample out the segmentation_masks variable, dividing them into 3 values. Inside of that, we define the variable m to the load_nibabel function that contained the msk variable. Finally, the multi_dim_plots function has been called, setting the m function as the \"multi_dim_array_input\" thus setting the id function to the msk variable that splitted off the slashes and extract the values of it with the split function along with the slice index of last element (-1).","metadata":{}},{"cell_type":"code","source":"from random import sample\n\nfor msk in sample(segmentation_masks, 3):\n    m = load_nibabel(msk)\n    multi_dim_plots(m, id=msk.split('/')[-1])","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:22.306749Z","iopub.execute_input":"2022-10-07T10:36:22.307106Z","iopub.status.idle":"2022-10-07T10:36:42.073203Z","shell.execute_reply.started":"2022-10-07T10:36:22.307075Z","shell.execute_reply":"2022-10-07T10:36:42.071736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After looping and plotting the batch of images three times, we observed that each specific image from a batch shows annotations over each cervical spine. Furthermore, we also observed some little fissures in some cervical spine images, meaning that some cervical spines were fractured.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the dicom images! Before we start visualizing the dicom images, we defined a function called read_dicom_file, setting the path variable as our function parameter, along setting the voi_lut and fix_monochrome parameters to True. Inside the read_dicom_file function, the dicom_input variable is defined to the pydicom module with the read_file module to read the specified dicom file input of the path parameter. Then, an if-statement is defined whether the voi_lut parameter is set to True to indicate that the VOI LUT is used to transform raw DICOM data to \"human-friendly\" view, then the data variable parameter is defined to the apply_voi_lut function to apply the voi lut of the pixel array of the dicom_input variable with the pixel_array attribute and the dicom_input variable itself. Otherwise, the data variable is defined to the pixel array of the dicom_input variable with the pixel array only. And depending on this value of the dicom images, the x-ray images over that maybe looking inverted. To fix that, we made another if-statement, whether fix_monochrome and the interpretation of the photometric images over the dicom_input variable with the PhotometricInterpretation is set to \"MONOCHROME1\", then the data variable is defined to the subtraction of the data variable from the maximum of the data variable array with the np module with the amax method then on the outside of the if statement, subtracted the minimum of the data variable with the np module with the min function from the data variable then dividing the data variable by the maximum of the data with the np module with the max function and then multiplied by 255 with the data variable again and convert the type of it with the astype function to an eight-bit unsigned integer by the uint8 attribute from the np module and finally return the data and dicom_input variables.","metadata":{}},{"cell_type":"code","source":"def read_dicom_file(path, voi_lut=True, fix_monochrome=True):\n    dicom_input = pydicom.read_file(path)\n    \n    if voi_lut:\n        data = apply_voi_lut(dicom_input.pixel_array, dicom_input)\n    else:\n        data = dicom_input.pixel_array\n        \n    if fix_monochrome and dicom_input.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n        \n    data = data - np.min(data)\n    data = data / np.max(data)\n    data = (data * 255).astype(np.uint8)\n    return data, dicom_input","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:42.074962Z","iopub.execute_input":"2022-10-07T10:36:42.075369Z","iopub.status.idle":"2022-10-07T10:36:42.083553Z","shell.execute_reply.started":"2022-10-07T10:36:42.075334Z","shell.execute_reply":"2022-10-07T10:36:42.08197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And after that, we define the dicom_files variable to the glob function, containing any file directory leading to .dcm files to indicate that we are returning all file paths that matched the dicom files and then print out the number of files in the dicom_files variable with the len function.","metadata":{}},{"cell_type":"code","source":"dicom_files = glob(\n    \"../input/rsna-2022-cervical-spine-fracture-detection/train_images/1.2.826.0.1.3680043.1062/*.dcm\"\n)\n\nlen(dicom_files)","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:42.085345Z","iopub.execute_input":"2022-10-07T10:36:42.086091Z","iopub.status.idle":"2022-10-07T10:36:42.109343Z","shell.execute_reply.started":"2022-10-07T10:36:42.086051Z","shell.execute_reply":"2022-10-07T10:36:42.107992Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this file directory we're in for our analysis, there are 201 dicom files in this set. However, the number of dicom files can be few or more depending on which file directory we're in each set of the folders under the train_images directory.","metadata":{}},{"cell_type":"markdown","source":"Now let's read and plot the first 16 dicom files! We define the temp variable to an array containing the function call of the read_dicom_file function with the i variable that is looped in the dicom_files file directory variable with the slice index of 16 after the colon to indicate the first 16 images along with the slice index of 0 followed by the slice index of None with the ellipsis notation for slicing and then defined to concatenating the previous temp variable with the axis parameter set to 0 in the concatenate function from the np module. And with the final touch of it, we return the shape of the temp variable with the shape attribute.","metadata":{}},{"cell_type":"code","source":"temp = [read_dicom_file(i)[0][None, ...] for i in dicom_files[:16]]\ntemp = np.concatenate(temp, axis=0)\ntemp.shape","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:42.110727Z","iopub.execute_input":"2022-10-07T10:36:42.11109Z","iopub.status.idle":"2022-10-07T10:36:42.353352Z","shell.execute_reply.started":"2022-10-07T10:36:42.111058Z","shell.execute_reply":"2022-10-07T10:36:42.352223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For the shape of the temp variable, there are 16 layers of images, 512 rows for the width, and 512 columns for the height inside of it.\n\nAnd time for plotting! We use the plot_batches function, inputting the temp variable and setting the cmap parameter to gray.","metadata":{}},{"cell_type":"code","source":"plot_batches(temp, cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:42.355161Z","iopub.execute_input":"2022-10-07T10:36:42.355771Z","iopub.status.idle":"2022-10-07T10:36:44.426771Z","shell.execute_reply.started":"2022-10-07T10:36:42.355735Z","shell.execute_reply":"2022-10-07T10:36:44.424055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After the result of this code cell above is plotted, we see 16 dicom images about the cervical spines in x-ray scan, rarely noticing the cracks in each one of them. However, we can change the color of the dicom images when plotting each batch of them.","metadata":{}},{"cell_type":"markdown","source":"So, let's get started on changing colors to the dicom images! We use the plot_batches function again, just like the same as we did previously from the above code cell, but this time, change the cmap parameter (color map) to bone.","metadata":{}},{"cell_type":"code","source":"plot_batches(temp, cmap=\"bone\")","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:44.428262Z","iopub.execute_input":"2022-10-07T10:36:44.428625Z","iopub.status.idle":"2022-10-07T10:36:47.224759Z","shell.execute_reply.started":"2022-10-07T10:36:44.428592Z","shell.execute_reply":"2022-10-07T10:36:47.223639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we ran this code cell, it's like the same output of the dicom images like our previous analysis, but however, the color of the dicom images is bone colored instead of the standard gray colors. Furthermore, it kinda looks like an x-ray scan.","metadata":{}},{"cell_type":"markdown","source":"Now it's time to visualize the bounding boxes! Before that, let's find the number of unique ids with bbox labels with the len function, containing the train_bound_box_df dataframe we created and visualized previously with the StudyInstanceUID data index for the unique identifier for the study of an individual patient, along finding the number of unique values in the same thing but outside from the len function.","metadata":{}},{"cell_type":"code","source":"len(train_bound_box_df.StudyInstanceUID), train_bound_box_df.StudyInstanceUID.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:47.227404Z","iopub.execute_input":"2022-10-07T10:36:47.22892Z","iopub.status.idle":"2022-10-07T10:36:47.239683Z","shell.execute_reply.started":"2022-10-07T10:36:47.228862Z","shell.execute_reply":"2022-10-07T10:36:47.238262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we compiled the code above, there are 7217 entities, along with 235 unique values in the train_bound_box_df dataframe.","metadata":{}},{"cell_type":"markdown","source":"Now let's go into the bounding box preprocessing! First, let's define the train_images variable to the file directory leading to the train images directory of the competition data.","metadata":{}},{"cell_type":"code","source":"train_images = '../input/rsna-2022-cervical-spine-fracture-detection/train_images'","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:47.241423Z","iopub.execute_input":"2022-10-07T10:36:47.241915Z","iopub.status.idle":"2022-10-07T10:36:47.255213Z","shell.execute_reply.started":"2022-10-07T10:36:47.241877Z","shell.execute_reply":"2022-10-07T10:36:47.253456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we setup the two lists, named five_twelve_temp and seven_sixty_eight_temp to each blank list. We then define the for loop, looping the index and the row variables in the tqdm function, containing the train_bound_box_df dataframe with the iterrows function to iterate every row of it, followed by the setup of total parameter (overall data) to the number of entities of the train_bound_box_df dataframe with the len function. \n\nInside the for loop, the dicom_bbox_path variable is created and defined to compromising the file directory of the train_images variable, the string conversion of the row variable that has the data index of 'StudyInstanceUID' with the str function, the string conversion of the row variable to the slice index of the 'slice_number' with the str function again, and the '.dcm' suffix by using the join function from the os module with the path submodule. Next, the dicom_bbox_read variable and the underscore is defined to the read_dicom_file function to read the dicom file of the dicom_bbox_path variable. Then, the part in where the processing applies to the bounding boxes happens when the xmin, ymin is defined to integer conversion of the row variable's x and y data index with the int function and the xmax and ymax variables defined to the xmin and ymin varaibles plus the integer conversion of the row variable's width and height data index with the int variable and then generate a rectangle with the rectangle function from the cv2 module, with the dicom_bbox_read variable, along with two arrays containing two variables: xmin, ymin and xmax and ymax, and another array containing 0, 255, 255, along with 3 on the outside of the array. After the bounding boxes were preprocessed, an if-statement is defined whether first slice index of the dicom_bbox_read variable's shape with the shape function is equal to 512, then the dicom_bbox_read variable with the slice index of None, with a single instance access with an ellipsis is appended to five_twelve_temp variable list, otherwise appended to seven_sixty_eight_temp variable list. \n\nAnd outside of the for loop, the five_twelve_temp and seven_sixty_eight_temp variable lists is redefined to being concatenated with no axis (axis parameter set to 0) with the concatenate functions by the pd module. Finally, we display the shape of each of them (five_twelve_temp and seven_sixty_eight temp lists) with the shape attribute.","metadata":{}},{"cell_type":"code","source":"five_twelve_temp = []\nseven_sixty_eight_temp = []\n\n# Remember the format of the dicom setup: [either train/test]_images/[StudyInstanceUID]/[slice_number].dcm\nfor index, row in tqdm(train_bound_box_df.iterrows(), total=len(train_bound_box_df)):\n    dicom_bbox_path = os.path.join(train_images, str(row['StudyInstanceUID']), str(row['slice_number']) + '.dcm')\n    dicom_bbox_read, _ = read_dicom_file(dicom_bbox_path)\n    \n    xmin = int(row['x'])\n    ymin = int(row['y'])\n    xmax = xmin + int(row['width'])\n    ymax = ymin + int(row['height'])\n    cv2.rectangle(dicom_bbox_read, (xmin, ymin), (xmax, ymax), (0, 255, 255), 3)\n    \n    if dicom_bbox_read.shape[0] == 512:\n        five_twelve_temp.append(dicom_bbox_read[None, ...])\n    else:\n        seven_sixty_eight_temp.append(dicom_bbox_read[None, ...])\n        \nfive_twelve_temp = np.concatenate(five_twelve_temp, axis=0)\nseven_sixty_eight_temp = np.concatenate(seven_sixty_eight_temp, axis=0)\nfive_twelve_temp.shape, seven_sixty_eight_temp.shape","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:36:47.25721Z","iopub.execute_input":"2022-10-07T10:36:47.257686Z","iopub.status.idle":"2022-10-07T10:38:38.396207Z","shell.execute_reply.started":"2022-10-07T10:36:47.257638Z","shell.execute_reply":"2022-10-07T10:38:38.394748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After this long run, we see that most dicom images has temp 512, while a few has temp 768. If we calculate the percentage between the total number of dicom files in this competition data and the number of images over temp 512 and temp 768, that'll make up about 99% of this data for temp 512 images and 0.1% of data for temp 768 images. \n\n⚠️ WARNING: this notebook cell above may potentially cause some abnormalties depending whether the internet option is off or on, making some notebook runs failing because of this error below:\n![](https://i.imgur.com/TXna0RW.png)\nTo fix this problem, all you need to do is to click the \"Restart & Clear Cell Outputs\" button from the three dotted icon in the top right and run all code cells again.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the first sixteen images of the dicom segmentations of each, single \"512 temp\" image! We use the plot_batches function, setting the five_twelve_temp with the slice index of 16 after the colon to indicate the first 16 images inside of it and setting the cmap to 'jet'.","metadata":{}},{"cell_type":"code","source":"plot_batches(five_twelve_temp[:16], cmap='jet')","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:38:38.398185Z","iopub.execute_input":"2022-10-07T10:38:38.398661Z","iopub.status.idle":"2022-10-07T10:38:40.769272Z","shell.execute_reply.started":"2022-10-07T10:38:38.398617Z","shell.execute_reply":"2022-10-07T10:38:40.767627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With the bounding boxes on the dicom images, we can see the fractures of the cervical spine from the first sixteen images. However, like our dicom image visualization without the bounding boxes, we can visualize the dicom images with bounding boxes in another color.","metadata":{}},{"cell_type":"markdown","source":"For that, we use the plot_batches function and apply the same setup previously, but for the cmap parameter, we set it to the bone and then the inferno color.","metadata":{}},{"cell_type":"code","source":"plot_batches(five_twelve_temp[:16], cmap=\"bone\")","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:38:40.770988Z","iopub.execute_input":"2022-10-07T10:38:40.771398Z","iopub.status.idle":"2022-10-07T10:38:43.432Z","shell.execute_reply.started":"2022-10-07T10:38:40.771359Z","shell.execute_reply":"2022-10-07T10:38:43.430468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_batches(five_twelve_temp[:16], cmap=\"inferno\")","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:38:43.434816Z","iopub.execute_input":"2022-10-07T10:38:43.435358Z","iopub.status.idle":"2022-10-07T10:38:46.099794Z","shell.execute_reply.started":"2022-10-07T10:38:43.435311Z","shell.execute_reply":"2022-10-07T10:38:46.097308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After running two code cells in two graphs, We saw them in the way of how the images and bounding boxes were alike, but the color of it, is inverted with bone and inferno color instead of the jet color. In other words, the color of the images and boxes changed, while the bounding boxes and the each look of the x-ray dicom images remained intact.","metadata":{}},{"cell_type":"markdown","source":"And finally for our image and bounding boxes visualization section, we'll visualize the 768 temp images! All we need to do is to use the plot_batches function, setting the seven_sixty_eight_temp dicom image variables with the slice index of 4 after the colon to plot 4 images in a batch since we understand that there are less images mentioning 768 temp, and setting the cmap parameter to \"jet\".","metadata":{}},{"cell_type":"code","source":"plot_batches(seven_sixty_eight_temp[:4], cmap=\"jet\")","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:38:46.101516Z","iopub.execute_input":"2022-10-07T10:38:46.101908Z","iopub.status.idle":"2022-10-07T10:38:48.095723Z","shell.execute_reply.started":"2022-10-07T10:38:46.101872Z","shell.execute_reply":"2022-10-07T10:38:48.093539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After the four \"768 temp\" images is displayed, we observed that all four of the dicom images shown were scanned on the C1 section, notably on the back of the neck part from the human body. Furthermore, just like the dicom images from the five_twelve_temp variable over the \"512 temp\", we analyzed some fractures in the C1 section, in which we saw the mini crack, along with the bounding box itself.","metadata":{}},{"cell_type":"markdown","source":"But let's visualize the \"768 temp\" dicom images in a different color! We use the plot_batches function we created again, setting up the same way we did but however, setting the cmap parameter to inferno and then bone.","metadata":{}},{"cell_type":"code","source":"plot_batches(seven_sixty_eight_temp[:4], cmap=\"inferno\")","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:38:48.102907Z","iopub.execute_input":"2022-10-07T10:38:48.103627Z","iopub.status.idle":"2022-10-07T10:38:50.779471Z","shell.execute_reply.started":"2022-10-07T10:38:48.103586Z","shell.execute_reply":"2022-10-07T10:38:50.777279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_batches(seven_sixty_eight_temp[:4], cmap=\"bone\")","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:38:50.781057Z","iopub.execute_input":"2022-10-07T10:38:50.781469Z","iopub.status.idle":"2022-10-07T10:38:52.87178Z","shell.execute_reply.started":"2022-10-07T10:38:50.781433Z","shell.execute_reply":"2022-10-07T10:38:52.869269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After changing the \"768 temp\" color from jet to inferno and bone, we only just saw the first part of the cervical spine and seeing the crack at the same time, but however, the background surrounding the cervical spine part that has the fracture turned from blue (jet) to dark (inferno and bone), along with the bounding boxes that spotted the fracture.\n\nAnd with all the analysis by graph and by visualzing the nibabel and dicom images, we've completely done our adventure over analyzing the cracks and fractures in this competition!","metadata":{}},{"cell_type":"markdown","source":"## Conclusion\nFrom all of our re-analysis from our graphs we made and the dicom and nibabel images we analyzed, it seems like there are nearly most patients didn't got their cervical spine broken, while nearly other most of them got their cervical spine broken. But however, we doubt that the fractured cervical spine cases may grow from now on, which is almost devastating to see each of them having neck pain day by day. And what we all know about fractured cervical spines, is that they were mostly result from a lot of high trauma and ground level falls from elderly people. And next time, always exercise, manage your medications, and talk to a therapist or a counselor over high level trauma, so that you'll not get your cervical spine fractured.","metadata":{}}]}