{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"text-align: center;\"><b>EDA Simplified:<span style=\"color: #3268a8;\"> RSNA Screening Mammography Breast Cancer Detection</span></b></h1>","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #3268a8; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Introduction</h2>\n\nCancer, the disease in which some cells abnormally divide uncontrollably and then devaste the whole body tissue. It can have some normal types, like Melanoma, Leukemia, Lymphoma, Prostate cancer, Lung cancer, Colon cancer, and on the top of that, Breast cancer. In terms of that kind of cancer, it forms in the cells of the breasts and can commonly occur in women and rarely in men. And the treatment over this type of cancer depends on the stage of it, and consists of chemotherapy, surgery, and most importantly, radiation, in which RSNA hosted a competition from that.\n\n<center>\n    <img src=\"https://my.clevelandclinic.org/-/scassets/Images/org/health/articles/3986-breast-cancer\" width=500>\n    <figcaption style=\"color: gray;\">A diagram about breast cancer from Cleveland Clinic</figcaption>\n</center>","metadata":{}},{"cell_type":"markdown","source":"And speaking about RSNA, they pointed out that the early detection of breast cancer needs professionally-trained human observers, which costs an arm and a leg (lots of money) for screening mammography programs thus also leading to large results of false positives, as it trails to unnecessary anxiety, inconvenient follow-up care, extra imaging tests, and sometimes a need for tissue sampling. So, based on machine learning, RSNA wants us to identify breast cancer, so that their endgoals can save more patients with breast cancer. However, prior for just getting started, check out our previous notebook on EDA Simplified on Cervical Spine Fracture Detection here: https://www.kaggle.com/code/dinowun/eda-simplified-rsna-cfd\n\nWith no farther ado, let's splash into data analysis towards the breast cancer data!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #3268a8; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Imports + Setup</h2>\n\nLike the previous competition over the fracture detection in cervical spine, we import following important modules like the pandas module as pd for data science, the numpy module as np for linear algebra, and the cv2 module for computer vision. However, we import plotly with the express submodule as px and then installing and importing the pandas_bokeh module thus displaying our bokeh plots with the output_notebook function from the pandas_bokeh module. And last and very important, we import the nibabel module as nib for neuroimaging, followed by installing the python-gdcm, pylibjpeg, pylibjpeg-libjpeg, and the pydicom packages then importing the pydicom module and the apply_voi_lut function from the pydicom module with the pixel_data_handlers submodule and then the util submodule for visualizing the dicom images, since there are dicom files in another RSNA competition, just like the previous competition we analyzed from them.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport cv2\n\n! pip install pandas_bokeh\nimport plotly.express as px\nimport pandas_bokeh\npandas_bokeh.output_notebook()\n\n! pip install python-gdcm\n! pip install pylibjpeg pylibjpeg-libjpeg pydicom\nimport nibabel as nib\nimport pydicom\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\n\nimport os\nfrom glob import glob\nfrom tqdm import tqdm","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:33:44.145693Z","iopub.execute_input":"2023-01-11T13:33:44.146496Z","iopub.status.idle":"2023-01-11T13:34:25.685568Z","shell.execute_reply.started":"2023-01-11T13:33:44.146399Z","shell.execute_reply":"2023-01-11T13:34:25.684029Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we imported a lot of modules, we create our standard training dataframe, train_df, to read out the train.csv file from RSNA's competition data with the read_csv function from the pd module. Following from that, we exhibit the first five rows of our newly-created train_df dataframe with the head function.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(\"/kaggle/input/rsna-breast-cancer-detection/train.csv\")\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:25.688269Z","iopub.execute_input":"2023-01-11T13:34:25.689392Z","iopub.status.idle":"2023-01-11T13:34:25.836334Z","shell.execute_reply.started":"2023-01-11T13:34:25.689339Z","shell.execute_reply":"2023-01-11T13:34:25.83515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style=\"background-color: #3268a8; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">Basic Dataframe Analysis</h3>\n\nThenceforth from our train_df dataframe creation, let's dive deep into our basic, quick dataframe analysis of that dataframe we've generated! First, let's visualize the total number data logged into the train_df dataframe with the len function to find the number of overall data units calculated.","metadata":{}},{"cell_type":"code","source":"len(train_df)","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:25.83789Z","iopub.execute_input":"2023-01-11T13:34:25.838353Z","iopub.status.idle":"2023-01-11T13:34:25.846211Z","shell.execute_reply.started":"2023-01-11T13:34:25.838319Z","shell.execute_reply":"2023-01-11T13:34:25.845107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see, there are 54706 total units of data inside the train_df dataframe. The thing behind the huge number of data entities from the train_df dataframe is that there are a lot of patient stats needed for detecting the breast cancer efficiently and accurately for the machine learning models.","metadata":{}},{"cell_type":"markdown","source":"Let's proceed to visualize the calculation of the overall NaN data values! We use the train_df dataframe to the isna function to find the present NaN values followed by the sum function to sum all the counted nan values in this specific dataframe.","metadata":{}},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:25.849706Z","iopub.execute_input":"2023-01-11T13:34:25.85269Z","iopub.status.idle":"2023-01-11T13:34:25.871064Z","shell.execute_reply.started":"2023-01-11T13:34:25.852637Z","shell.execute_reply":"2023-01-11T13:34:25.869799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the overall NaN values calculated after running this code cell above, we found out that 37 of them were in the age column, 28420 in the BIRADS column, and 25236 in the density column, which sums them to total 53693 NaN values in the train_df dataframe. In other words, the NaN data being present in the train_df dataframe hinted us that some data in this dataframe contained anonymous entities and unclear data from each mammography radiation session.","metadata":{}},{"cell_type":"markdown","source":"Lastly, let's find out how many columns are there in the train_df dataframe! We use the shape attribute into the train_df dataframe to find the shape of this dataframe, along obtaining the second part of it with the slice index of 1.","metadata":{}},{"cell_type":"code","source":"train_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:25.872428Z","iopub.execute_input":"2023-01-11T13:34:25.873056Z","iopub.status.idle":"2023-01-11T13:34:25.880973Z","shell.execute_reply.started":"2023-01-11T13:34:25.873015Z","shell.execute_reply":"2023-01-11T13:34:25.879586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the output of the compiled code cell on the top, we've noticed that there are 14 columns counted in the train_df dataframe. Nevertheless, the fourteen columns of the train_df dataframe shows us that those columns were collected for the data based on the ids for site, patient, image, and machine, along with laterality, view, age, cancer, biopsy, invasive, BIRADS, implant, and difficult negative cases.\n\nNow that we've got through our straightforward basic analysis about the train_df dataframe, it's time to enter onwards the full data analysis on breast cancer data from the dataframe we've created, which it is the train_df dataframe!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #3268a8; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Data Analysis in the Training Dataframe</h2>\n\nOnce we got into the section where we start our data analysis in RSNA's Breast Cancer detection competition, let's visualize the site_id data, which is the ID for the hospital source into Bokeh's histogram. To get started on that, we use the train_df dataframe's site_id into the plot_bokeh function as we configure our chart, setting the kind parameter to hist as we set our chart's kind into a histogram.","metadata":{}},{"cell_type":"code","source":"train_df[\"site_id\"].plot_bokeh(\n    kind=\"hist\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:25.883057Z","iopub.execute_input":"2023-01-11T13:34:25.883588Z","iopub.status.idle":"2023-01-11T13:34:26.111725Z","shell.execute_reply.started":"2023-01-11T13:34:25.883542Z","shell.execute_reply":"2023-01-11T13:34:26.110566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From that histogram we created, it visualized the data bins from 1 to 1.1 and 1.9 to 2 under the site_id data column, meaning that there's ones and twos present in here. In other words, all the ones in the site_id data column is counted the most, with 29519 data entities listed, than the twos, which has 25187 data entities calculated. Moreover, the ones and twos mentioned in the site_id data column shows us that there are two hospitals listed anonymously in the RSNA competition.","metadata":{}},{"cell_type":"markdown","source":"Let's switch to Plotly over visualizing the patient_id data by distributing them into a histogram! We characterize our figure variable, fig, to the histogram function from the px module for configuring the histogram, setting the train_df dataframe as our data for the histogram and the x parameter to the patient_id column for the x-axis of our histogram. And after that, we simply exhibit the fig variable figure by using the show function to it.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"patient_id\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:26.11326Z","iopub.execute_input":"2023-01-11T13:34:26.113625Z","iopub.status.idle":"2023-01-11T13:34:27.218222Z","shell.execute_reply.started":"2023-01-11T13:34:26.113589Z","shell.execute_reply":"2023-01-11T13:34:27.217141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From every distribution of the data from the patient_id column in the train_df dataframe, we found out that there's a lot of counting from 0 to 65999, in which the highest count of the distribution of patient_id column is in between 17500 to 17999, with 521 units of data calculated, while the lowest is in between 65500 and 65999, with just 22 units of data calculated. In other words, the patient_id data distributed from the train_df dataframe shows us that it is used for organizing a data of train images for the machine learning model to identify breast cancers.","metadata":{}},{"cell_type":"markdown","source":"Let's return back to plotting with Bokeh over the image_id data with our histogram! We use the train_df dataframe's image_id data column to the plot_bokeh function for configuring our Bokeh chart, setting the kind parameter to hist as we set our graph to be a histogram type.","metadata":{}},{"cell_type":"code","source":"train_df[\"image_id\"].plot_bokeh(\n    kind=\"hist\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:27.219746Z","iopub.execute_input":"2023-01-11T13:34:27.220141Z","iopub.status.idle":"2023-01-11T13:34:27.430893Z","shell.execute_reply.started":"2023-01-11T13:34:27.220105Z","shell.execute_reply":"2023-01-11T13:34:27.429585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Holy smokes! From the histogram we've created from our code cell above after running it, we've visualized that there's a \"e+\" on every image_id x-axis, and not only that, we've espied huge columns of counted data, in which the highest is in between 1288510504 to 1503250839, with 5586 units of data counted while the lowest is in 2148088265 to 4295419620, with 5333 units of data calculated. Moreover, the various distribution of the image_id data column from the train_df dataframe gave us the point that this data column is attributed to the training images inside the folders from the patient_id column, by each 4 images for the machine learning model to train over with.","metadata":{}},{"cell_type":"markdown","source":"Let's move onwards to the laterality data visualization into the pie chart with help from Plotly! We define the laterality dataframe from the train_df dataframe's laterality column being counted with the value_counts function and then characterize the fig variable to the px module's pie function for configuring the pie chart, setting the laterality dataframe as our data for the pie chart, the names parameter to the indexes of the laterality dataframe with the index attribute for specifiying the labels of our pie chart, and the values parameter to the values of the laterality dataframe with the values attribute for specifying the values of the pie chart. Thenceforth, we display our fig figure variable by applying the show function to it!","metadata":{}},{"cell_type":"code","source":"laterality = train_df[\"laterality\"].value_counts()\n\nfig = px.pie(laterality, names=laterality.index, values=laterality.values)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:27.43243Z","iopub.execute_input":"2023-01-11T13:34:27.433469Z","iopub.status.idle":"2023-01-11T13:34:27.506521Z","shell.execute_reply.started":"2023-01-11T13:34:27.433427Z","shell.execute_reply":"2023-01-11T13:34:27.505227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once the code cell above compiled our pie chart, we noted that our pie chart contained two data of laterality: R and L, in which their data counts were nearly have the same values. In other words, R has the highest count of data, with 50.2% of data taken in the laterality column, than L, with 49.8% of data present in the laterality column. Furthermore, the Rs and Ls in the laterality data from the train_df dataframe represents the data showing the left and right breasts in images, which there are four of them each in the patient_id folder, with two L images and two R images inside of them.","metadata":{}},{"cell_type":"markdown","source":"Let's transfer to Bokeh over plotting the pie chart consisting over the view data column from the train_df dataframe! Prior to plotting, we inherit our dataframe, view, to the train_df dataframe's view column being counted with the value_counts function. Once finished, we use the view column to the plot_bokeh function to create our bokeh plot, setting the kind parameter to pie for configuring our graph to be a pie chart and the x parameter to the indexes of the view dataframe with the index attribute for configuring our labels of our diagram.","metadata":{}},{"cell_type":"code","source":"view = train_df[\"view\"].value_counts()\n\nview.plot_bokeh(\n    kind=\"pie\",\n    x=view.index\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:27.512046Z","iopub.execute_input":"2023-01-11T13:34:27.512442Z","iopub.status.idle":"2023-01-11T13:34:27.620639Z","shell.execute_reply.started":"2023-01-11T13:34:27.512407Z","shell.execute_reply":"2023-01-11T13:34:27.619484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How strange. The MLO (Medio Lateral Obique) and CC (Cranical Caudad) data from the train_df dataframe's view column only appeared in the pie chart, while the AT (Axillary Tail), LM (Latero-Medial), ML (Medial-Latero), and LMO (Latero-Medial Obique) doesn't appear in the pie chart, in which the MLO has more data than the CC as there are 27903 entities in MLO and 26765 in CC. Nevertheless, the absent data from the train_df dataframe's view column hinted us that the \"Other\" data mentioned in the view column from the train.csv file shows 0% and remained hidden once we're viewing it from the Data tab in the RSNA Competition page, for an unknown reason.","metadata":{}},{"cell_type":"markdown","source":"Let's splash into Plotly over distributing the age data column from the train_df dataframe and visualizing it with a histogram! We simply create our variable, fig, to the histogram function provided by the px module as we create our histogram, setting the train_df dataframe as our data input and the x parameter to the age column for the x-axis configuration of our histogram. Finally, we exhibit our figure variable, fig, with the show function.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"age\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:27.621867Z","iopub.execute_input":"2023-01-11T13:34:27.622241Z","iopub.status.idle":"2023-01-11T13:34:27.6986Z","shell.execute_reply.started":"2023-01-11T13:34:27.622212Z","shell.execute_reply":"2023-01-11T13:34:27.697324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As for the age distribution in a histogram, we've remarked that there are various data counts of patients aged 26 to 89, in which the highest of them is age 50, with 2248 data entities listed while age 31 is the lowest, with 4 entities calculated. In other words, most patients that are aged 50 were diagnosed to breast cancer as they are found in woman who are 50 years old or older, according to [CDC](https://www.cdc.gov/cancer/breast/basic_info/risk_factors.htm#:~:text=The%20main%20factors%20that%20influence,factors%20that%20they%20know%20of.).","metadata":{}},{"cell_type":"markdown","source":"Let's divert back to Bokeh for distributing the train_df dataframe's cancer column in a bar chart! We define the cancer dataframe from the train_df dataframe's cancer column counted by value with the value_counts function, and then we use the cancer dataframe to the plot_bokeh function to plot our bokeh graph, setting the kind parameter to the bar for configuring our graph into a bar chart and the x parameter to indexes of the cancer dataframe with the index attribute for the x-axis configuration of our bar chart.","metadata":{}},{"cell_type":"code","source":"cancer = train_df[\"cancer\"].value_counts()\n\ncancer.plot_bokeh(\n    kind=\"bar\",\n    x=cancer.index,\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:27.700292Z","iopub.execute_input":"2023-01-11T13:34:27.700712Z","iopub.status.idle":"2023-01-11T13:34:27.813492Z","shell.execute_reply.started":"2023-01-11T13:34:27.70067Z","shell.execute_reply":"2023-01-11T13:34:27.812283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From our bar chart we've created, we found out that most patients weren't diagnosed with breast cancer (the 0 data entity) than the patients that were diagnosed with it (the 1 data entity). In other words, there are 53548 data units of zeros listed in the train_df dataframe's cancer column, while there are 1158 ones counted in the same dataframe with the cancer data index. Moreover, most patients had some risk factors but however, most of them didn't get breast cancer by being physically active, not taking hormones, and not drinking alcohol (as mentioned by CDC's video below).","metadata":{}},{"cell_type":"code","source":"# Nothing to see here!\nfrom IPython.display import HTML\nimport warnings\nwarnings.filterwarnings('ignore')\n\nHTML(\"\"\"<iframe width=\"560\" height=\"315\" src=\"https://www.youtube.com/embed/JAVV-Hmp9o8\" title=\"YouTube video player\" frameborder=\"0\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture\" allowfullscreen></iframe>\"\"\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-01-11T13:34:27.814995Z","iopub.execute_input":"2023-01-11T13:34:27.815837Z","iopub.status.idle":"2023-01-11T13:34:27.823647Z","shell.execute_reply.started":"2023-01-11T13:34:27.815802Z","shell.execute_reply":"2023-01-11T13:34:27.822421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's swap back to Plotly for distributing the biopsy and invasive columns with a histogram, but this time, have fun experiementing subplots! Before that, we must import the make_subplots function from the plotly module with the subplots submodule and the plotly module's graph_objects submodule as go. After that, we characterize the fig variable to configuring our subplots with the make_subplots function, setting the rows parameter to 1 for making our subplot a single row, the cols parameter to 2 for making our subplot to have two columns, and the shared_yaxes parameter to True for making our subplots to share y-axis to any amount of graphs. And following from our creation of our subplots, we add a trace for our subplot with the add_trace function to the fig figure variable, setting the histogram configuration inside of it with the Histogram function from the go module that has the x configuration set to the train_df dataframe's biopsy column for the x-axis of the histogram and the name parameter to biopsy for making the histogram graph label clearly, the row parameter to 1 for placing our graph in the first row, and the col parameter to 1 for placing our graph to the left column. Meanwhile, we add another trace to our subplot with the add_trace function again, setting another histogram with the go module's Histogram function just like what we did for the first trace in our subplot, but the x parameter in it is set to the invasive column from the train_df dataframe for the x-axis of the histogram and the name parameter is set to invasive for specifying the label of the histogram, the row parameter to 1 for placing the histogram in the first row, and the col parameter to 2 for placing it on the second column or to the right of the first histogram. And just like that, we display our fig variable figure with the show function.","metadata":{}},{"cell_type":"code","source":"from plotly.subplots import make_subplots\nimport plotly.graph_objects as go\n\nfig = make_subplots(rows=1, cols=2, shared_yaxes=True)\nfig.add_trace(go.Histogram(x=train_df[\"biopsy\"], name=\"biopsy\"), row=1, col=1)\nfig.add_trace(go.Histogram(x=train_df[\"invasive\"], name=\"invasive\"), row=1, col=2)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:27.825525Z","iopub.execute_input":"2023-01-11T13:34:27.825876Z","iopub.status.idle":"2023-01-11T13:34:27.875523Z","shell.execute_reply.started":"2023-01-11T13:34:27.825846Z","shell.execute_reply":"2023-01-11T13:34:27.874318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the first histogram, with the blue bars shown on the left, we visualized that most patients doesn't have a biopsy, in which there are 51737 data entities of zeros, while there are 2969 of people does have a biopsy. In other words, some patients got biopsies for visualizing the sample cells from them for examining the disease around it, because of medical research or in most cases, high incidence of false-positive results from mammography screens.\n\nMeanwhile for the second histogram, shown in red, we noted that just like the first histogram over the data of the biopsy column, we visualized most data in the invasive data column were zeroes, meaning that there are 53888 patients doesn't have invasive cancer to their breasts, even if they have breast cancer, while 818 of the patients specified as ones does have invasive cancer when they have breast cancer. Moreover, most breast cancer patients that were labeled as zeroes over the invasive data shows us that they have [DCIS (ductal carcinoma in stiu)](https://my.clevelandclinic.org/health/diseases/3986-breast-cancer), which it is treatable although prompt care is necessary.","metadata":{}},{"cell_type":"markdown","source":"Let's proceed into plotting our BIRADS data column to a histogram with help from Bokeh! We simply use the train_df dataframe's BIRADS data column to the plot_bokeh function to configure our bokeh plot, setting the kind parameter to hist for configuring our plot to a histogram.","metadata":{}},{"cell_type":"code","source":"train_df[\"BIRADS\"].plot_bokeh(\n    kind=\"hist\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:27.877029Z","iopub.execute_input":"2023-01-11T13:34:27.877429Z","iopub.status.idle":"2023-01-11T13:34:28.099043Z","shell.execute_reply.started":"2023-01-11T13:34:27.877393Z","shell.execute_reply":"2023-01-11T13:34:28.098054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After our histogram was generated by our compiled cell, we found out that the range from 1 to 1.2 is counted the most in this diagram, with 15772 data units total and on the other hand, the range between 1.8 to 2 is counted the least, with 2265 data units summed together. Moreover, the zeroes, ones, and twos mentioned in the BIRADS data column hinted us that most patients in the RSNA competition data were tested negative for breast cancer, while others were required the follow-ups for their breasts or rated as normal to their breasts.","metadata":{}},{"cell_type":"markdown","source":"Let's use two different graphs for plotting the implant and density columns of data into our two subplots! Before we begin this, we define the implant dataframe to the train_df dataframe's implant column having its values counted with the value_counts function. After that, we create our variable, fig, to the function call of make_subplots function for creating our subplots, setting the rows parameter to 1 to make our subplots have one row, the cols parameter to 2 to make our subplots have 2 columns, and the shared_yaxes parameter to True for the graphs to share their y-axes, and then we add our trace to our first subplot with the add_trace function, containing the bar chart configuration by the Bar function from the go module, in which the x parameter is set to the indexes of the implant dataframe with the index attribute for the bar graph's x-axis, the y parameter set to the values of the implant dataframe with the values attribute for the y-axis specification of the bar graph, and the name parameter to \"implant\" since we need to specify our first graph, thus setting the row parameter to 1 for placing our graph to the first row and the col parameter to 1 for placing our graph to the first column. Thenceforth from our first graph configuration of our subplots, we configure another graph, which is the histogram configuration by the go module's Histogram function with the x parameter set to the density columns from the train_df dataframe for configuring the x-axis of a histogram and the name parameter to \"density\" for specifying our second histogram graph clearly, hence setting the row parameter to 1 for placing the histogram graph to the first row and the col parameter to 2 for placing the histogram into the second column, which is to the right of the implant graph. Eventually, we display the fig variable figure with the show function.","metadata":{}},{"cell_type":"code","source":"implant = train_df[\"implant\"].value_counts()\n\nfig = make_subplots(rows=1, cols=2, shared_yaxes=True)\nfig.add_trace(go.Bar(x=implant.index, y=implant.values, name=\"implant\"), row=1, col=1)\nfig.add_trace(go.Histogram(x=train_df[\"density\"], name=\"density\"), row=1, col=2)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:28.100666Z","iopub.execute_input":"2023-01-11T13:34:28.101498Z","iopub.status.idle":"2023-01-11T13:34:28.346114Z","shell.execute_reply.started":"2023-01-11T13:34:28.101454Z","shell.execute_reply":"2023-01-11T13:34:28.344912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For the implant data graph to the left of our subplot, we espied that most patients doesn't have an implant (with 53229 zeros counted), while a few people have their implant (1477 ones counted). In other words, a few patients got their breast implant since they wanted to have [larger breasts or rebuild them after cancer injury/surgery](https://www.mayoclinic.org/healthy-lifestyle/womens-health/in-depth/breast-implants/art-20045957#:~:text=Breast%20implants%20are%20used%20for,or%20injury%2C%20also%20called%20reconstruction.).\n\nMeanwhile, for the density distribution in a histogram to the right of the implant bar graph, we've noted that \"B\" has the highest data counted, with 12651 entities listed in the density column, while \"D\" has 1539 entities calculated, as it is the lowest counted data in the density column. Moreover, \"B\" and \"C\" in the density column are listed as medium dense, as it can sometimes make the diagnosis difficult.","metadata":{}},{"cell_type":"markdown","source":"Now we're switching back to Bokeh just for plotting the machine_id column into our histogram! We insert the train_df dataframe's machine_id column to the plot_bokeh function for plotting our graph with Bokeh, setting the kind parameter to hist for configuring our graph to a histogram.","metadata":{}},{"cell_type":"code","source":"train_df[\"machine_id\"].plot_bokeh(\n    kind=\"hist\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:28.347517Z","iopub.execute_input":"2023-01-11T13:34:28.347937Z","iopub.status.idle":"2023-01-11T13:34:28.5786Z","shell.execute_reply.started":"2023-01-11T13:34:28.347906Z","shell.execute_reply":"2023-01-11T13:34:28.577222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we saw this distribution of the data provided from the machine_id column, we found out that from 21 to 60, there's more counts of distributed data while the range from 79.5 to 99 and 157.5 to 216 were isolated to each other. In other words, the 40.5 to 60 range is distributed the most, with 32228 data entities listed, while 177 to 196.5 is counted the least, with 145 data entities calculated. Moreover, the data listed on the machine_id column shows us that the data in it comes in four values each and some data values can repeat inside the train_df dataframe's machine_id column for another scan in the imaging device.","metadata":{}},{"cell_type":"markdown","source":"Last but not least, let's visualize the data in the difficult_negative_case column from the train_df dataframe into our bar chart! Prior to setting up our bar diagram with Plotly, we characterize our difficult_neg_case dataframe to the train_df dataframe's difficult_negative_case column counted by value with the value_counts function. Thenceforth, we create our fig variable to the bar function from the px module, setting the difficult_neg_case dataframe as our data for the bar chart, the x parameter to the indexes of the difficult_neg_case dataframe with the index attribute as the x-axis of our bar chart, the y parameter to the values of the diffcult_neg_case dataframe with the values attribute as the y-axis of our bar chart, and the color parameter to the indexes of the difficult_neg_case dataframe with the index attribute again for specifying the color of the bars in our bar chart. Following from our bar chart configuration, we exhibit our fig variable with the show function.","metadata":{}},{"cell_type":"code","source":"difficult_neg_case = train_df[\"difficult_negative_case\"].value_counts()\n\nfig = px.bar(difficult_neg_case, x=difficult_neg_case.index, y=difficult_neg_case.values, color=difficult_neg_case.index)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:28.580065Z","iopub.execute_input":"2023-01-11T13:34:28.580428Z","iopub.status.idle":"2023-01-11T13:34:28.668663Z","shell.execute_reply.started":"2023-01-11T13:34:28.580396Z","shell.execute_reply":"2023-01-11T13:34:28.667438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After our code cell above compiled our bar graph, We noticed that most data inside the difficult_negative_case column were marked as False, with 41001 entities counted, than the ones that were marked as True, with 7705 entities counted. Furthermore, the data about the difficult_negative_case column shows us that some of the patients mentioned in the data had their breast cancer cases difficult even when they don't have cancer, while others doesn't have their breast cancer cases difficult. That's because there's some cases of unclear mammography screening.","metadata":{}},{"cell_type":"markdown","source":"And just like that, we've visualized all fourteen columns of data in the train_df dataframe! With that done, let's move onwards into visualizing the dicom images from RSNA's mammography competition data!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #3268a8; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Dicom Image Visualization</h2>\n\nJust like what we did previously on visualizing the cervical spine fractures in the last competition, we're blasting into fun and the last part of our data analysis in breast cancer detection by visualizing the dicom images! However, we found out that there are no .nii files around this competition we're in, so using the nibabel module as nib doesn't worth it for this section. But on the bright side, we're going to read out the dicom images since we have the pydicom modules equipped in our data visualization proceedures! \n\nFirst, let's generate a function called read_dicom, defining the path parameter variable inside of it, thus setting the voi_lut parameter to True for allowing the non-linear conversion for the output of the conceptual modality LUT values to the input of the presentation LUT, and the fix_monochrome parameter to True for fixing the monochrome images if they are broken. Inside the read_dicom file, the dicom_input variable is assigned to read out the dicom files with the read_file function from the pydicom module, with the path variable parameter specified. After that, an if statement is compiled, whether the voi_lut parameter is set to True, then the data variable is defined to the function call of the apply_voi_lut for applying the non-linear conversion for the output of the conceptual modality LUT values to the input of the presentation LUT, setting the dicom_input variable inside of it that is converted into a pixel array by the pixel_array attribute, followed by the dicom_input variable again, otherwise, the data variable is assigned to the pixel array of the dicom_input variable specified by the pixel_array attribute. After the first if-statement is created, another if-statement is also created, whether the fix_monochrome parameter is set to true and the dicom_input variable's photometric interpretation specified by the PhotometricInterpretation attribute is equal to \"MONOCHROME1\", then the data variable is redefined to the maximum array of itself specified by the np module's amax function along subtracting itself. Once finished, the data variable is redefined to itself minus the minimum values of the data variable itself, then redefined to itself divided by the maximum values of the data variable itself, and then redefined again to the specified type of itself times 255 by the astype function, containing the uint8 type specified by the np module's uint8 attribute. Finally, the dicom_input and the data variables is returned by the read_dicom function.","metadata":{}},{"cell_type":"code","source":"def read_dicom(path, voi_lut=True, fix_monochrome=True):\n    dicom_input = pydicom.read_file(path)\n    \n    if voi_lut:\n        data = apply_voi_lut(dicom_input.pixel_array, dicom_input)\n    else:\n        data = dicom_input.pixel_array\n        \n    if fix_monochrome and dicom_input.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n        \n    data = data - np.min(data)\n    data = data / np.max(data)\n    data = (data * 255).astype(np.uint8)\n    return data, dicom_input","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:28.670373Z","iopub.execute_input":"2023-01-11T13:34:28.670822Z","iopub.status.idle":"2023-01-11T13:34:28.679007Z","shell.execute_reply.started":"2023-01-11T13:34:28.670777Z","shell.execute_reply":"2023-01-11T13:34:28.677784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we setup the read_dicom function, we characterize the dicom_files variable to the glob function call to list out the dcm files from any patient_id folder. After that, we also characterize the temp variable to the array containing the read_dicom function call, containing the i variable that is looped in before the 4 files of the dicom_files variable specified by the slice index of a colon and 4, followed by the first index specified by 0, then specified again by the slice index of None, and the three dots thing. We then redefine the temp variable to concatenate itself with the concatenate function from the np module, in which the axis parameter is set to 0 for the axis configuration.","metadata":{}},{"cell_type":"code","source":"dicom_files = glob(\n    \"/kaggle/input/rsna-breast-cancer-detection/train_images/10006/*.dcm\"\n)\n\ntemp = [read_dicom(i)[0][None, ...] for i in dicom_files[:4]]\ntemp = np.concatenate(temp, axis=0)","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:28.680635Z","iopub.execute_input":"2023-01-11T13:34:28.680968Z","iopub.status.idle":"2023-01-11T13:34:38.272534Z","shell.execute_reply.started":"2023-01-11T13:34:28.680937Z","shell.execute_reply":"2023-01-11T13:34:38.271315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Following from that, let's create another function over creating our plotting batches with the plot_batch function, containing the samples variable parameter, and the cmap parameter set to None for the color map configuration. Inside this function, we import the matplotlib module with the pyplot submodule as plt and then we set our figure configuration with the figure function, setting the figsize parameter to 30 by 30 for configuring our size of the figure. After that, the xy variable is assigned to the int conversion (by the int function) of the square root (specified as the np module's sqrt function) of the first item from the shape of the samples variable specified by the shape attribute and the slice index of 0. Thenceforth, the for loop is created, looping the i and sample variables in enumering the samples variable with the enumerate function. From the for loop inside the plot_batch function, the try-except statement is generated, and inside the try statement, the ax variable is defined to the subplot creation with the subplot function from the plt module, setting the xy variables twice in rows and columns, thus increasing the i variable's index by 1 inside the index. Meanwhile in the except statement, the ax variable is defined to the same thing from the try statement, but inside the subplot function, the rows and columns were set to 8. Outside from the try-except statement, the axis of the matplotlib graph is disabled by the axis function from the plt module, which contains the \"open\" string inside of it thus showing the image with the plt module's imshow function, setting the sample variable inside of it, and the cmap parameter is set to cmap for configuring the colormap of the graph. Finally, on the outside from the for loop inside the function, the graph is displayed with the show function from the plt module.","metadata":{}},{"cell_type":"code","source":"def plot_batch(samples, cmap=None):\n    import matplotlib.pyplot as plt\n    \n    plt.figure(figsize=(30, 30))\n    xy = int(np.sqrt(samples.shape[0]))\n    \n    for i, sample in enumerate(samples):\n        try:\n            ax = plt.subplot(xy, xy, i+1)\n        except:\n            ax = plt.subplot(8, 8, i+1)\n        plt.axis(\"off\")\n        plt.imshow(sample, cmap=cmap)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:38.274899Z","iopub.execute_input":"2023-01-11T13:34:38.276589Z","iopub.status.idle":"2023-01-11T13:34:38.284671Z","shell.execute_reply.started":"2023-01-11T13:34:38.276528Z","shell.execute_reply":"2023-01-11T13:34:38.283243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we created the plot_batch function, it's time for us to plot! We use the plot_batch function to plot four batches of dicom images, setting the temp variable inside of it for configuring the images of it, and the cmap parameter to gray, jet, and bone separately for setting the colormap of our images.","metadata":{}},{"cell_type":"code","source":"plot_batch(temp, cmap=\"gray\")","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:38.286341Z","iopub.execute_input":"2023-01-11T13:34:38.287822Z","iopub.status.idle":"2023-01-11T13:34:49.968523Z","shell.execute_reply.started":"2023-01-11T13:34:38.287772Z","shell.execute_reply":"2023-01-11T13:34:49.967263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_batch(temp, cmap=\"jet\")","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:34:49.970311Z","iopub.execute_input":"2023-01-11T13:34:49.971089Z","iopub.status.idle":"2023-01-11T13:35:01.700399Z","shell.execute_reply.started":"2023-01-11T13:34:49.971046Z","shell.execute_reply":"2023-01-11T13:35:01.698779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_batch(temp, cmap=\"bone\")","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:35:01.702581Z","iopub.execute_input":"2023-01-11T13:35:01.703004Z","iopub.status.idle":"2023-01-11T13:35:13.515894Z","shell.execute_reply.started":"2023-01-11T13:35:01.702965Z","shell.execute_reply":"2023-01-11T13:35:13.514614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, we visualized four dicom images in each three diagram we plotted in different colors, as the breasts were pictured from the breast cancer patients from left to right. In other words, the images plotted into the imshow diagram explains us that each image_id folder specified come in four images, with views of left and right were counted by 2. However, we clearly see that there are no signs of breast cancer in this image we analyzed, as we find some dense areas around the patient's breasts. ","metadata":{}},{"cell_type":"markdown","source":"Nevertheless, let's analyze one of the positive breast cancer images! To begin with this, we characterize the positive_cancer variable to list the files associated with the dcm images from one of the patient_id folders that were marked as positive with the glob function.","metadata":{}},{"cell_type":"code","source":"positive_cancer = glob(\n    \"/kaggle/input/rsna-breast-cancer-detection/train_images/10130/*.dcm\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:35:13.517229Z","iopub.execute_input":"2023-01-11T13:35:13.517844Z","iopub.status.idle":"2023-01-11T13:35:13.534362Z","shell.execute_reply.started":"2023-01-11T13:35:13.51781Z","shell.execute_reply":"2023-01-11T13:35:13.533119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we characterize the positive_cancer variable, we characterize the temp_cancer variable to the same thing as what we did for the previous temp variable (including concatenating itself with the concatenate function from the np module) but for the looping part, we use the positive_cancer variable for the i variable to loop over.","metadata":{}},{"cell_type":"code","source":"temp_cancer = [read_dicom(i)[0][None, ...] for i in positive_cancer[:4]]\ntemp_cancer = np.concatenate(temp_cancer, axis=0)","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:35:13.536466Z","iopub.execute_input":"2023-01-11T13:35:13.536917Z","iopub.status.idle":"2023-01-11T13:35:17.997265Z","shell.execute_reply.started":"2023-01-11T13:35:13.536871Z","shell.execute_reply":"2023-01-11T13:35:17.995875Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once we finished characterizing the temp_cancer and positive_cancer variables, we use the plot_batch function we created previously, setting the temp_cancer inside for visualizing the images in that variable thus setting the cmap parameter to gray, jet, and bone separately for visualizing the images in different colormaps.","metadata":{}},{"cell_type":"code","source":"plot_batch(temp_cancer, cmap=\"gray\")","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:35:17.998834Z","iopub.execute_input":"2023-01-11T13:35:17.99979Z","iopub.status.idle":"2023-01-11T13:35:24.519107Z","shell.execute_reply.started":"2023-01-11T13:35:17.999749Z","shell.execute_reply":"2023-01-11T13:35:24.517924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_batch(temp_cancer, cmap=\"jet\")","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:35:24.523819Z","iopub.execute_input":"2023-01-11T13:35:24.524691Z","iopub.status.idle":"2023-01-11T13:35:31.183334Z","shell.execute_reply.started":"2023-01-11T13:35:24.524651Z","shell.execute_reply":"2023-01-11T13:35:31.181803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_batch(temp_cancer, cmap=\"bone\")","metadata":{"execution":{"iopub.status.busy":"2023-01-11T13:35:31.184763Z","iopub.execute_input":"2023-01-11T13:35:31.185532Z","iopub.status.idle":"2023-01-11T13:35:37.783933Z","shell.execute_reply.started":"2023-01-11T13:35:31.185488Z","shell.execute_reply":"2023-01-11T13:35:37.782617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From what we saw in these images in the four subplots, we noticed that there's some clusters of tumors spotted in the patient's breasts in four images. Specifically, some patients' cases of breast cancers vary in different situations, whether their size of a tumor is 2 cm or smaller or their lymph nodes feel lumpy or not. However, we found that the patient's breasts seems a bit dense, as it staged a problem for mammography screening leading to other problems like uneccessary anxiety, inconvenient follow-up care, or even more tests.","metadata":{}},{"cell_type":"markdown","source":"And with our image analysis completed, we're officially done with our visualization part of the dicom files specified in the RSNA Competition!","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #3268a8; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Conclusion</h2>\n\nBased out of our analysis from visualizing the data in the train_df dataframe by graphs and visualizing some dicom images from the competition data, we found out that most patients were marked as negative, as they were tested with mammography screens or biopsied. That's because most patients have a balanced diet, weighted healthily, and stayed active. However, there are some difficult negative cases because of various reasons like the density in some patient's breasts were more dense, making the diagnosis of breast cancer complicated for screening. And from all of our explaination and data analysis from breast cancer, avoid/limit alcohol, don't smoke, be active, and control your weight, so that you'll not be in risk of being diagnosed with breast cancer.","metadata":{}}]}