{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **EDA Simplified: Mayo Clinic STRIP AI**\n![](https://storage.googleapis.com/kaggle-competitions/kaggle/37333/logos/header.png?t=2022-06-29-00-47-20)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"## Introduction\nSlurred speech. Numbnesses. Muscle weakness. These are some of the symptoms mentioned about this second infamous cause-of-death medical disturbance, which is a stroke, sometimes called a brain attack.\n\n![](https://www.mayoclinic.org/-/media/kcms/gbs/patient-consumer/images/2014/10/29/13/33/r7_ischemicstroke.jpg)\n\nFrom this image provided by Mayo Clinic (the competition host), we can see that a stroke blood clot happens if a clot lodges in cerebral artery, breaks off, and then travels. Fortunately, Mayo Clinic shows us the solution to this health problem, which is the STRIP AI aka Stroke Thromboembolism Registry of Imaging and Pathology. Speaking of which, it is a uniquely large multicenter project led by Mayo Clinic Neurovascular Lab with the aim of histopathologic characterization of thromboemboli of various etiologies and examining clot composition and its relation to mechanical thrombectomy revascularization. And that's how we are going to use image classification of stroke blood clots, in an excellent way of doing EDA! So, let's begin to unclog the clots with data analysis!","metadata":{}},{"cell_type":"markdown","source":"## Imports and CSV File Setup\nLike some EDA notebooks, we can start with importing the necessary modules for linear algebra and data processing, which is numpy as np and pandas as pd. We then import out the matplotlib module with the pyplot submodule as plt for analyzing the images, the plotly module with the express submodule as px for plotting out data, along with the seaborn module as sns, and last but not least, the tifffile module for reading out the tiff images, as mentioned in the [HuBMAP competition](https://www.kaggle.com/code/dinowun/eda-simplified-hubmap-hpa-organ-segmentation).","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\n\nimport tifffile","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:10.548443Z","iopub.execute_input":"2022-08-07T07:08:10.549386Z","iopub.status.idle":"2022-08-07T07:08:13.206598Z","shell.execute_reply.started":"2022-08-07T07:08:10.549248Z","shell.execute_reply":"2022-08-07T07:08:13.205411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After importing the important modules used for doing EDA in this STRIP AI competition, let's setup our two dataframes, train_df and other_df by defining them to read out the csv with the read_csv function provided by the pd module, containing a file directory path leading to the two csv files. Finally, we'll display our two newly created dataframes with the head function.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(\"../input/mayo-clinic-strip-ai/train.csv\")\nother_df = pd.read_csv(\"../input/mayo-clinic-strip-ai/other.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:13.209389Z","iopub.execute_input":"2022-08-07T07:08:13.210241Z","iopub.status.idle":"2022-08-07T07:08:13.238514Z","shell.execute_reply.started":"2022-08-07T07:08:13.210193Z","shell.execute_reply":"2022-08-07T07:08:13.236621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:13.242769Z","iopub.execute_input":"2022-08-07T07:08:13.244373Z","iopub.status.idle":"2022-08-07T07:08:13.271898Z","shell.execute_reply.started":"2022-08-07T07:08:13.244321Z","shell.execute_reply":"2022-08-07T07:08:13.270639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"other_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:13.275775Z","iopub.execute_input":"2022-08-07T07:08:13.276301Z","iopub.status.idle":"2022-08-07T07:08:13.293921Z","shell.execute_reply.started":"2022-08-07T07:08:13.276243Z","shell.execute_reply":"2022-08-07T07:08:13.291564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we create and peek at the two dataframes, let's use the basic EDA on the two dataframes!","metadata":{}},{"cell_type":"markdown","source":"## Chapter 1: Double Basics, Double Dataframes\nSince we have two dataframes setup, let's analyze the basics of the two dataframes! First, let's find out how many data entities are there in train_df and other_df dataframes.","metadata":{}},{"cell_type":"code","source":"len(train_df)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:13.295622Z","iopub.execute_input":"2022-08-07T07:08:13.296832Z","iopub.status.idle":"2022-08-07T07:08:13.304528Z","shell.execute_reply.started":"2022-08-07T07:08:13.29678Z","shell.execute_reply":"2022-08-07T07:08:13.303635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(other_df)","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:13.306868Z","iopub.execute_input":"2022-08-07T07:08:13.307967Z","iopub.status.idle":"2022-08-07T07:08:13.318462Z","shell.execute_reply.started":"2022-08-07T07:08:13.307915Z","shell.execute_reply":"2022-08-07T07:08:13.316033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see, the number of entities in the train_df dataframe is 754, while the number of entities in the other_df dataframe is 396.","metadata":{}},{"cell_type":"markdown","source":"Let's move on to finding the nan values of the two dataframes. All we need to do is to use the isna function along with the sum function.","metadata":{}},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:13.3202Z","iopub.execute_input":"2022-08-07T07:08:13.321314Z","iopub.status.idle":"2022-08-07T07:08:13.333891Z","shell.execute_reply.started":"2022-08-07T07:08:13.321255Z","shell.execute_reply":"2022-08-07T07:08:13.332307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"other_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:13.335669Z","iopub.execute_input":"2022-08-07T07:08:13.336464Z","iopub.status.idle":"2022-08-07T07:08:13.347134Z","shell.execute_reply.started":"2022-08-07T07:08:13.33641Z","shell.execute_reply":"2022-08-07T07:08:13.345896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we ran the two code cells above, we noticed that there are no nan values in the train_df dataframe. However, if we analyzed the sum of nan values in the other_df dataframe, then we found out that there are 334 nan values in the data in other_specified. ","metadata":{}},{"cell_type":"markdown","source":"Last but not least, let's find the shape of the two dataframes with the shape attribute!","metadata":{}},{"cell_type":"code","source":"train_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:13.348969Z","iopub.execute_input":"2022-08-07T07:08:13.349704Z","iopub.status.idle":"2022-08-07T07:08:13.357686Z","shell.execute_reply.started":"2022-08-07T07:08:13.349649Z","shell.execute_reply":"2022-08-07T07:08:13.356121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"other_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:13.365019Z","iopub.execute_input":"2022-08-07T07:08:13.365752Z","iopub.status.idle":"2022-08-07T07:08:13.37257Z","shell.execute_reply.started":"2022-08-07T07:08:13.365705Z","shell.execute_reply":"2022-08-07T07:08:13.371456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As a result, we made a good comparison that both train_df and other_df dataframes has 5 data columns. However, the differences between them is that there are 754 rows of data entities in the train_df dataframe whilst the number the other_df dataframe has 396 rows of data entities.\n\nAnd with that, we completed the first chapter of analyzing the two dataframes! Now, time to move on to the next two chapters of data analysis on two dataframes separately!","metadata":{}},{"cell_type":"markdown","source":"## Chapter 2: Data Analysis (train_df)\nWe're now on analyzing the train_df dataframe! Before we begin, let's go over the data on the train_df dataframe one by one!\n* **image_id**: ID specified for the image data.\n* **center_id**: ID specified for the center the patient is in.\n* **patient_id**: ID specified for each patient.\n* **image_num**: Enumerates images of clots obtained by the same patient.\n* **label**: An important data for classifying either Cardioembolic (CE) or Large Artery Atherosclerosis (LAA).\n\nWith all of that explained, let's analyze the center data in the train_df dataframe! First, with Plotly, we define fig to the px module that has the histogram function to indicate that we are creating a histogram, setting the data_frame parameter (dataframe specification) to our train_df dataframe, the x parameter (x-axis specification) set to the center_id data index, the marginal parameter (margin plot) to violin or other specified, and the nbin parameter (number of bins) to 400. Optionally, you can update layout with the update_layout function to the fig variable, setting the template to any built-in template Plotly provides. Finally, we show out the fig variable with the show function.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(data_frame=train_df, x=\"center_id\", marginal=\"box\", nbins=400)\nfig.update_layout(template=\"plotly_dark\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:13.374928Z","iopub.execute_input":"2022-08-07T07:08:13.375371Z","iopub.status.idle":"2022-08-07T07:08:14.800922Z","shell.execute_reply.started":"2022-08-07T07:08:13.375338Z","shell.execute_reply":"2022-08-07T07:08:14.799718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When the graph was generated after running this code cell above, we see that there are tiny lines of data counts in this histogram, ranging from 1 to 11, in which the highest number of data in this graph is 11, while the lowest is 8-9. Above from our histogram, we created a box plot, in which the median number of data is 7, the first quartile is 4, and the last quartile is 11.","metadata":{}},{"cell_type":"markdown","source":"Now, let's move on to analyzing image_num! It's like the same as what we did for analyzing the center_id data, but we set the marginal parameter to violin to see which number range has the most data.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(data_frame=train_df, x=\"image_num\", marginal=\"violin\", nbins=400)\nfig.update_layout(template=\"plotly_dark\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:14.803176Z","iopub.execute_input":"2022-08-07T07:08:14.803546Z","iopub.status.idle":"2022-08-07T07:08:14.947562Z","shell.execute_reply.started":"2022-08-07T07:08:14.803516Z","shell.execute_reply":"2022-08-07T07:08:14.946049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Unlike the previous graph, the highest number of data for counts of image_num in the train_df dataframe is 0, while the lowest is 4. Thus, the median of this graph is also 0.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the overall data about the center_id and image_num in Seaborn! All we need to do is to create a categorical plot by using the catplot function from sns, setting the data (dataframe) parameter to the train_df module, setting the x (x-axis) parameter to image_num data index, the y (y-axis) parameter to center_id.","metadata":{}},{"cell_type":"code","source":"sns.catplot(data=train_df, x=\"image_num\", y=\"center_id\")","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:14.949842Z","iopub.execute_input":"2022-08-07T07:08:14.950346Z","iopub.status.idle":"2022-08-07T07:08:15.303526Z","shell.execute_reply.started":"2022-08-07T07:08:14.9503Z","shell.execute_reply":"2022-08-07T07:08:15.301934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And after running this code cell, we can see that the line plots were ranged under the image_num labels, with 4 having the least data.","metadata":{}},{"cell_type":"markdown","source":"Finally, let's create a pie chart visualizing the labels in Matplotlib (since it was used for plotting)! Before that, let's examine the values of the label data index in the train_df dataframe with the value_counts function.","metadata":{}},{"cell_type":"code","source":"train_df[\"label\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:15.305181Z","iopub.execute_input":"2022-08-07T07:08:15.305644Z","iopub.status.idle":"2022-08-07T07:08:15.319564Z","shell.execute_reply.started":"2022-08-07T07:08:15.305609Z","shell.execute_reply":"2022-08-07T07:08:15.318396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, let's plot the values mentioned above to the pie chart (547 in Cardioembolic, 207 in Large Artery Atherosclerosis)! We define a list variable, counts, to a list containing two numbers mentioned in the previous sentence in parentheses along with the labels variable to \"CE\" and \"LAA\". Then, we define fig and ax variables to creating subplots with the plt module that has the subplots function. We then use the pie function to the ax variable to create a pie chart, inputting the counts list inside of it and setting four parameters: labels (label for the pie chart) set to the labels variable, autopct (percentage format) set to \"%1.1f%%\", shadow (shadow for the bar chart) set to True, and startangle (starting angle) set to 90. Furthermore, we set the axis to equal with the axis function in the ax variable to ensure that our pie chart is drawn in a circle format. Finally, we show our figure with the show function to the plt module.","metadata":{}},{"cell_type":"code","source":"labels = \"CE\", \"LAA\"\ncounts = [547, 207]\n\nfig, ax = plt.subplots()\nax.pie(counts, labels=labels, autopct=\"%1.1f%%\", shadow=True, startangle=90)\nax.axis(\"off\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:15.321191Z","iopub.execute_input":"2022-08-07T07:08:15.321548Z","iopub.status.idle":"2022-08-07T07:08:15.43888Z","shell.execute_reply.started":"2022-08-07T07:08:15.321516Z","shell.execute_reply":"2022-08-07T07:08:15.437502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What we observed from our pie chart we created, we can see that there's 72.5% of data containing Cardioembrolic label in the train_df dataframe, while 27.5% of data contains Large Artery Atherosclerosis. So, that means the Cardioembrolic label covers most data than the Large Artery Atheroscelerosis.","metadata":{}},{"cell_type":"markdown","source":"Now that we've done chapter 1 analyzing the train_df dataframe, let's move on to chapter 2 of other_df dataframe!","metadata":{}},{"cell_type":"markdown","source":"## Chapter 2: Data Analysis (other_df)\nSince we are at analyzing the other_df dataframe, let's go over the data given by this dataframe one by one, just like how we did when analyzing the train_df dataframe!\n* **image_id**: ID specified for the image data\n* **patient_id**: ID specified for each patient\n* **image_num**: Enumerates images of clots obtained by the same patient.\n* **other_specified**: The specific etiology, when known, in case the etiology is labeled as `Other`.\n* **label**: Unlike the train_df dataframe's label data index, it specified either Unknown or Other of the etiology of the clot.","metadata":{}},{"cell_type":"markdown","source":"With all that being said, let's analyze the data given in the other_df dataframe! First, let's analyze the image_num data by creating a bar chart in Plotly. To do that, we define a variable called fig, to the px module with the bar function, setting the x (x-axis) parameter to listing the unique values with the unique function by the np module to the image_num data index from the other_df dataframe and the y (y-axis) parameter to counting the list of the image_num data index of the other_df dataframe listed by the the list function along with the count function, containing the i variable looping in the unique values of the image_num data given by the other_df dataframe. We then update our x and y axes of our figure, setting the title to \"Image Num\" (x-axis) and \"Number of Them\" (y-axis). Finally, we show the fig figure with the show function.","metadata":{}},{"cell_type":"code","source":"fig = px.bar(x=np.unique(other_df[\"image_num\"]), y=[list(other_df[\"image_num\"]).count(i) for i in np.unique(other_df[\"image_num\"])])\nfig.update_xaxes(title=\"Image Num\")\nfig.update_yaxes(title=\"Number of Them\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:15.441043Z","iopub.execute_input":"2022-08-07T07:08:15.442038Z","iopub.status.idle":"2022-08-07T07:08:15.562074Z","shell.execute_reply.started":"2022-08-07T07:08:15.441955Z","shell.execute_reply":"2022-08-07T07:08:15.560833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As a result of this, we can see that the image_num 0 has the most counts with 336 entities of them, while the least of them is the image_num 4 with a single entity.","metadata":{}},{"cell_type":"markdown","source":"Let's proceed to take a look at the other_specified data in here! First, we define a variable, other_specified_data to count the values of the \"other_specified\" data from the other_df dataframe with the value_counts function, setting the normalize (normalizing an object) to True, the ascending (least to most data) parameter to True, and dropna (dropping the NaN values) to False.\n\nNext, we define our figure variable, fig, to the bar function from the px module again, but this time, we're going to setup two parameters differently: the x one to renaming the other_specified_data variable's NaN (provided by the np module that has the nan attribute) values to \"Not Specified\" in a dictionary with the rename function and then indexing it with the index attribute, and the y parameter to the values of the other_specified_data variable with the values attribute times 100. We then update our x and y axis with the update_xaxes and update_yaxes functions to the fig variable, setting the title parameter to \"Other Diagnoses\" and \"Number of Them\". Finally, we display our fig figure with the show function.","metadata":{}},{"cell_type":"code","source":"other_specified_data = other_df[\"other_specified\"].value_counts(\n    normalize=True, ascending=True, dropna=False\n)\n\nfig = px.bar(x=other_specified_data.rename({np.nan: \"Not Specified\"}).index, y=other_specified_data.values * 100)\nfig.update_xaxes(title=\"Other Diagnoses\")\nfig.update_yaxes(title=\"Number of Them\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:15.563849Z","iopub.execute_input":"2022-08-07T07:08:15.564548Z","iopub.status.idle":"2022-08-07T07:08:15.630957Z","shell.execute_reply.started":"2022-08-07T07:08:15.564501Z","shell.execute_reply":"2022-08-07T07:08:15.630089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After running this cell above, we see that most other_specified data in the other_df dataframe is NaN values aka \"Not Specified\", while the least is tumor embolization.","metadata":{}},{"cell_type":"markdown","source":"We now proceed to find the values of the label data in the other_df dataframe by creating a pie chart in Plotly! First, we define a variable, other_diagnoses, to the other_df dataframe's label data index and count it with the value_counts function, setting the normalize parameter (for normalizing) to True.\n\nNext, we define the fig variable to the px module with the pie function to create a pie chart, putting in the other_diagnoses variable thus setting the values to the values (with the values attribute) of the other_diagnoses variable, the names set to the index (with the index attribute) of the other_diagnoses variable, and the title to \"Label Distribution for Others\". Finally, we show the fig variable with the show function.","metadata":{}},{"cell_type":"code","source":"other_diagnoses = other_df[\"label\"].value_counts(normalize=True)\n\nfig = px.pie(other_diagnoses, values=other_diagnoses.values, names=other_diagnoses.index, title=\"Label Distribution for Others\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:15.632909Z","iopub.execute_input":"2022-08-07T07:08:15.633321Z","iopub.status.idle":"2022-08-07T07:08:15.711773Z","shell.execute_reply.started":"2022-08-07T07:08:15.633286Z","shell.execute_reply":"2022-08-07T07:08:15.710633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Per the Label Distribution for Others pie chart after running the code cell above, we observed that 83.6% of the label data in the other_df dataframe classified as Unknown, while 16.4% of it classified as Other. Furthermore, we made a good correlation that a few data from our previous analysis of the other_specified data is mostly similar with the 16.4% classified as \"Other\" in label data.","metadata":{}},{"cell_type":"markdown","source":"Last but not least, let's find out the number of patient records in the other_df dataframe! We first define a variable called other_patient_records to grouping the patient_id data with the label data in the other_df dataframe with the groupby function and count it with the count function and then count the values of it with the value_counts function, setting the normalize parameter (normalizing option) to True.\n\nNext, we define the fig variable one more time to the px module with the bar function to create a bar chart with Plotly, setting the x parameter to index of the other_patient_records variable (with the index attribute) and the y (y-axis) parameter to the values (with the values attribute) of the other_patient_values variable again but multiplied by 100. We then proceed to update our x and y axes with the update_xaxes and update_yaxes functions to the fig variable, setting the title to \"No. of records\" and \"The Percentage\". Lastly, let's show our fig figure with the show function!","metadata":{}},{"cell_type":"code","source":"other_patient_records = other_df.groupby(\"patient_id\")[\"label\"].count().value_counts(normalize=True)\n\nfig = px.bar(x=other_patient_records.index, y=other_patient_records.values * 100)\nfig.update_xaxes(title=\"No. of records\")\nfig.update_yaxes(title=\"The Percentage\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:15.713425Z","iopub.execute_input":"2022-08-07T07:08:15.713837Z","iopub.status.idle":"2022-08-07T07:08:15.778156Z","shell.execute_reply.started":"2022-08-07T07:08:15.71379Z","shell.execute_reply":"2022-08-07T07:08:15.776854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Alternatively, we can create a pie chart just like the one we did previously when analyzing the label data in the other_df, but with the number of records in the patient_id data.","metadata":{}},{"cell_type":"code","source":"fig = px.pie(other_diagnoses, values=other_patient_records.values * 100, names=other_patient_records.index, title=\"Records per Patient\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:15.779956Z","iopub.execute_input":"2022-08-07T07:08:15.780353Z","iopub.status.idle":"2022-08-07T07:08:15.832578Z","shell.execute_reply.started":"2022-08-07T07:08:15.780318Z","shell.execute_reply":"2022-08-07T07:08:15.831235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After running the separate two cells for creating a pie and a bar chart, we can clearly see that 86 percent of the data records in the patient_id and label data is mostly 1 record, while 0.3ish percent of them is 5 records. So, we can infer that most patients specified in the other_df data has **exactly** 1 record.\n\nIn summary, we've completed our data analysis of the train_df and other_df's dataframes. And now, let's look forward to analyzing the tiff images given in this competition!","metadata":{}},{"cell_type":"markdown","source":"## Chapter 4: Image Preprocessing & Analysis\nAfter we done some EDA over the train_df and other_df dataframes, let's preprocess the images given in this competition and analyze it! Before that, we need to import additional modules for our chapter, which is cv2 for machine learning (aka OpenCV), random for randomness, os for operating system stuff, matplotlib module again, but with the image submodule as mpimg for reading out images, Image from PIL for images (aka Python Image Library), OpenSlide from the openslide module for reading whole-slide images, defaultdict from the collections module for specialized container datatypes in a default dictionary, and finally, resize from skimage module with the transform submodule for resizing the image. Thus, we also need to import glob for file directory stuff. And finally, we need to set the MAX_IMAGE_PIXELS to None for the Image module","metadata":{}},{"cell_type":"code","source":"import cv2\nimport random\nimport os\n\nimport matplotlib.image as mpimg\nfrom PIL import Image\n\nfrom openslide import OpenSlide\nfrom collections import defaultdict\n\nfrom skimage.transform import resize \nimport glob\n\nImage.MAX_IMAGE_PIXELS = None","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:15.834333Z","iopub.execute_input":"2022-08-07T07:08:15.83503Z","iopub.status.idle":"2022-08-07T07:08:16.32547Z","shell.execute_reply.started":"2022-08-07T07:08:15.834963Z","shell.execute_reply":"2022-08-07T07:08:16.324347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The next part is going to be fun. We are going to load out the images from this competiton data! First, we define three variables: train, test, and other, to the glob module with the glob function to list all of the images from the three separate image directories. And finally, we print out the number of images in each directory by applying the len function to each three variables.","metadata":{}},{"cell_type":"code","source":"train = glob.glob(\"../input/mayo-clinic-strip-ai/train/*\")\ntest = glob.glob(\"../input/mayo-clinic-strip-ai/test/*\")\nother = glob.glob(\"../input/mayo-clinic-strip-ai/other/*\")\n\nprint(\"No. of images in train: \", len(train))\nprint(\"No. of images in test: \", len(test))\nprint(\"No. of images in other: \", len(other))","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:16.326836Z","iopub.execute_input":"2022-08-07T07:08:16.327187Z","iopub.status.idle":"2022-08-07T07:08:16.436529Z","shell.execute_reply.started":"2022-08-07T07:08:16.327157Z","shell.execute_reply":"2022-08-07T07:08:16.43531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After running this code cell above, we can see that there are 754 images in train, 396 images in other, and 4 images in test.","metadata":{}},{"cell_type":"markdown","source":"With all of that being said, let's create a dataframe for analyzing images in the train directory! To do that, we define a variable called img_prop to the dictionary class of a list. Next, we create a for loop, looping i and path variables to counting with the enumerate function of the train variable.\n\nInside of this for loop, the img_path is defined to the train variable with the slice index of the i value variable. Then, the slide variable is defined to generating the individual deep zoom titles from the img_path variable. And now, the img_prop with the data indexes of image_id, width, height, size, and path is defined to appending with the append function to:\n* The image_path with the slice index of the range between the last 12 rows to the last 4 rows to extract the id of the image (image_id).\n* The dimensions of the slide variable slides with the dimensions attribute, with the slice index of 0 to indicate the width (width).\n* The dimensions of the slide variable slides with the dimensions attribute again, but with the slice index of 1 to indicate the height (height).\n* The rounded number by the nearest hundredth by 2 with the round function of getting the size of the img_path variable with the os module that has the path submodule followed by the getsize function divided by 1e-6 (size).\n* The whole img_path variable (path).\n\nOutside of that for loop, we define a new dataframe, image_df, to the pd module with the DataFrame function to create a new dataframe out of the img_prop data variables. Furthermore, we create our new data index in the image_df dataframe called img_aspect_ratio to the division of the train_df dataframe's width and height data indexes. We now then sort the values with the sort_values function, setting the by (sort by which data) parameter to image_id, and the inplace (change default behavior) parameter to True and then reset the image_df dataframe's index with the reset_index function, setting the inplace (change default behavior) parameter to True, and drop (drop some specified values) to True.\n\nThus, we redefine image_df to merge two dataframes with the merge function to the train_df dataframe, setting on (on which data) parameter to image_id. Finally, we'll display the first rows of the image_df dataframe.","metadata":{}},{"cell_type":"code","source":"img_prop = defaultdict(list)\n\nfor i, path in enumerate(train):\n    img_path = train[i]\n    slide = OpenSlide(img_path)\n    img_prop[\"image_id\"].append(img_path[-12:-4])\n    img_prop[\"width\"].append(slide.dimensions[0])\n    img_prop[\"height\"].append(slide.dimensions[1])\n    img_prop[\"size\"].append(round(os.path.getsize(img_path) / 1e-6, 2))\n    img_prop[\"path\"].append(img_path)\n    \nimage_df = pd.DataFrame(img_prop)\nimage_df[\"img_aspect_ratio\"] = image_df[\"width\"] / image_df[\"height\"]\nimage_df.sort_values(by=\"image_id\", inplace=True)\nimage_df.reset_index(inplace=True, drop=True)\n\nimage_df = image_df.merge(train_df, on=\"image_id\")\nimage_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:16.437993Z","iopub.execute_input":"2022-08-07T07:08:16.438317Z","iopub.status.idle":"2022-08-07T07:08:34.749367Z","shell.execute_reply.started":"2022-08-07T07:08:16.438288Z","shell.execute_reply":"2022-08-07T07:08:34.748034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After creating the image_df image dataframe, let's go on to sorting the CE and LAA images! To get started, we define the CE and LAA variables to the image_df image dataframe that has the location with the loc attribute of the label data index of the image_df image dataframe equal to CE and LAA in the path data index.","metadata":{}},{"cell_type":"code","source":"CE = image_df.loc[image_df['label'] == 'CE', \"path\"]\nLAA = image_df.loc[image_df['label'] == 'LAA', \"path\"]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:34.750724Z","iopub.execute_input":"2022-08-07T07:08:34.751077Z","iopub.status.idle":"2022-08-07T07:08:34.758539Z","shell.execute_reply.started":"2022-08-07T07:08:34.751047Z","shell.execute_reply":"2022-08-07T07:08:34.757302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And here's the part we're going have fun with, plotting down the images!","metadata":{}},{"cell_type":"markdown","source":"### **CE Data**\nLet's plot the CE images! Before we begin, we set the default style in the plt module with the style attribute along with the use function. And after that, we define the fig and axes variables to the plt module with the subplots function, setting the two values to 1, 5, and the figsize (figure size) parameter to 16 by 16. Then, we create a for loop that loops the ax variable into the axes variable with the reshape function to reshape it to -1.\n\nInside this for loop, we define the img_path variable to the np module with the random attribute along with the choice function to display a random set of the CE images in the CE variable. Thus, the img variable was defined to the Image module with the open function to open the images from the img_path variable and then setting up the thumbnail with the thumbnail function by 300 by 300, and then resample with the Lanczos algorithm by using the Image module with the Resampling module alongside with the LANCZOS attribute. After that, the images were shown by the img variable to the ax variable with the imshow function, alongside setting the title to CE with the set_title function. Finally, we show the figure with the plt module that has the show function.","metadata":{}},{"cell_type":"code","source":"plt.style.use('default')\nfig, axes = plt.subplots(1, 5, figsize=(16,16))\n\nfor ax in axes.reshape(-1):\n    img_path = np.random.choice(CE)\n    img = Image.open(img_path)\n    img.thumbnail((300,300), Image.Resampling.LANCZOS)\n    ax.imshow(img)\n    ax.set_title(\"CE\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:08:34.759795Z","iopub.execute_input":"2022-08-07T07:08:34.760162Z","iopub.status.idle":"2022-08-07T07:09:33.036538Z","shell.execute_reply.started":"2022-08-07T07:08:34.76013Z","shell.execute_reply":"2022-08-07T07:09:33.035419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After running the code cell for a very long time, we see that there are 5 random-picked images were labeled as CE (Cardioembolic). This happens when there is a clot in the heart, and then travels to the brain and lodges the clot to the brain.","metadata":{}},{"cell_type":"markdown","source":"### **LAA Data**\nNow, let's analyze the LAA images! It's like the same as what we did to analyzing CE images but we use the LAA variable and label it LAA instead of the CE variable and labeling CE.","metadata":{}},{"cell_type":"code","source":"plt.style.use('default')\nfig, axes = plt.subplots(1, 5, figsize=(16,16))\n\nfor ax in axes.reshape(-1):\n    img_path = np.random.choice(LAA)\n    img = Image.open(img_path)\n    img.thumbnail((300,300), Image.Resampling.LANCZOS)\n    ax.imshow(img)\n    ax.set_title(\"LAA\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:09:33.038271Z","iopub.execute_input":"2022-08-07T07:09:33.039053Z","iopub.status.idle":"2022-08-07T07:10:50.410029Z","shell.execute_reply.started":"2022-08-07T07:09:33.039006Z","shell.execute_reply":"2022-08-07T07:10:50.409001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just like how we did for analysing the CE images, we plotted down 5 randomly-picked images that labeled as LAA (Large Artery Atherosclerosis). This happens when it leads to ischemic symptomatology with two mechanisms, which is embolic phenomena and regional brain hyperfusion.","metadata":{}},{"cell_type":"markdown","source":"Now, let's find out the images in the \"other\" data directory! We create another dataframe, other_image_df, to what we did for creating the image_df dataframe out of the train_df dataframe, but this time, we use the other file variable.","metadata":{}},{"cell_type":"code","source":"img_prop = defaultdict(list)\n\nfor i, path in enumerate(other):\n    img_path = other[i]\n    slide = OpenSlide(img_path)\n    img_prop[\"image_id\"].append(img_path[-12:-4])\n    img_prop[\"width\"].append(slide.dimensions[0])\n    img_prop[\"height\"].append(slide.dimensions[1])\n    img_prop[\"size\"].append(round(os.path.getsize(img_path) / 1e-6, 2))\n    img_prop[\"path\"].append(img_path)\n    \nother_image_df = pd.DataFrame(img_prop)\nother_image_df[\"img_aspect_ratio\"] = image_df[\"width\"] / image_df[\"height\"]\nother_image_df.sort_values(by=\"image_id\", inplace=True)\nother_image_df.reset_index(inplace=True, drop=True)\n\nother_image_df = other_image_df.merge(other_df, on=\"image_id\")\nother_image_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:10:50.411251Z","iopub.execute_input":"2022-08-07T07:10:50.411645Z","iopub.status.idle":"2022-08-07T07:11:00.8984Z","shell.execute_reply.started":"2022-08-07T07:10:50.411614Z","shell.execute_reply":"2022-08-07T07:11:00.897221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we've load up our images from the other directory of the competition data, let's define two variables, unknown and other, to the other_image_df dataframe that has the location with the loc attribute to the other_image_df dataframe again but with the label data index equal to Unknown or Other, and set it to the \"path\" index.","metadata":{}},{"cell_type":"code","source":"unknown = other_image_df.loc[other_image_df[\"label\"] == \"Unknown\", \"path\"]\nother = other_image_df.loc[other_image_df[\"label\"] == \"Other\", \"path\"]","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:11:00.904188Z","iopub.execute_input":"2022-08-07T07:11:00.905114Z","iopub.status.idle":"2022-08-07T07:11:00.913153Z","shell.execute_reply.started":"2022-08-07T07:11:00.905073Z","shell.execute_reply":"2022-08-07T07:11:00.911704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With all of what being said, let's visualize the other images in this competition data!","metadata":{}},{"cell_type":"markdown","source":"### **Unknown Data**\nTo get started on that, we define fig and axes to the plt module with the subplots function, setting the subplots to 1 row and 5 columns thus setting the figsize parameter to 16 by 16. Then, we use the for loop, looping the ax variable to the axes variable that is reshaped by -1 with the reshape function.\n\nInside of that, the img_path is defined to the np module with the random attribute along with the choice function, containing the unknown variable to generate 5 random file directories to any images. Then, the img variable is defined to opening the image provided by the img_path variable with the Image module that has the open function. After that, the img variable get their thumbnail resized by 300 by 300 with the thumbnail function, along setting up the Lanczos algorithm with the Image module that has the Resampling attribute along with the LANCZOS attribute. Thus, the ax variable will show the images provided by the img variable with the imshow function thus setting the title with the set_title function to \"Unknown\". Finally, outside of that for loop, the plt module with the show function will show the figure we made.","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 5, figsize=(16,16))\n\nfor ax in axes.reshape(-1):\n    img_path = np.random.choice(unknown)\n    img = Image.open(img_path)\n    img.thumbnail((300,300), Image.Resampling.LANCZOS)\n    ax.imshow(img), ax.set_title(\"Unknown\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:11:00.915269Z","iopub.execute_input":"2022-08-07T07:11:00.915773Z","iopub.status.idle":"2022-08-07T07:12:10.373041Z","shell.execute_reply.started":"2022-08-07T07:11:00.915726Z","shell.execute_reply":"2022-08-07T07:12:10.3718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Per the five, \"Unknown\" labeled images, we see that the images weren't classified as LAA or CE, since the scan to each of them were blurry, or the way of the clot is shaped as weird.","metadata":{}},{"cell_type":"markdown","source":"### **Other Data**\nNext, let's find out the images labeled as \"Other\"! It's vaguely the same as what we did to visualizing the images of LAA, CE, and Unknown labels, but this time, we are going to analyze the \"Other\" labeled data from the other variable thus labeling it with \"Other\"!","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 5, figsize=(16,16))\n\nfor ax in axes.reshape(-1):\n    img_path = np.random.choice(other)\n    img = Image.open(img_path)\n    img.thumbnail((300,300), Image.Resampling.LANCZOS)\n    ax.imshow(img), ax.set_title(\"Other\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-07T07:12:10.374689Z","iopub.execute_input":"2022-08-07T07:12:10.375099Z","iopub.status.idle":"2022-08-07T07:13:51.380624Z","shell.execute_reply.started":"2022-08-07T07:12:10.375062Z","shell.execute_reply":"2022-08-07T07:13:51.378972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see from this graph of these images above, we see that there are images labeled as other. The reason why is because they were dissected, embolized with tumors, or in a hypercoagulable state.","metadata":{}},{"cell_type":"markdown","source":"## Conclusion\nAnd with all of that being covered up with our data analysis on the STRIP AI Stroke Detection! We go through all of the data entities in train.csv, other.csv, and all of the images specified! Now that we've done all of that, we may classify images of asthma, heart disease, or other diseases like ebola!","metadata":{}}]}