{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nRecently, artificial intelligence (AI) has also been widely employed in the field of healthcare. Particularly, AI has contributed to medical diagnosis in radiology by convolutional neural network (CNN). CNN has a great capacity to classify images such as X-ray images by way of supervised learning. It is expected that AI may facilitate the quality of healthcare services and reduce workloads for healthcare professionals.\n\nThis time, a png dataset is used instead of the original RSMA dataset, because Tensorflow cannot directly read dicom file. The dataset is created by and available at [RSNA Breast Cancer Detection - 512x512 pngs](https://www.kaggle.com/datasets/theoviel/rsna-breast-cancer-512-pngs).","metadata":{}},{"cell_type":"markdown","source":"# Import Libraries\n\nResNet50 is a greatly popular and frequently used CNN model for medical AI research. The results of image classification performance by an AI model is generally estimated by classification report and confusion matrix.","metadata":{}},{"cell_type":"code","source":"# TensorFlow libraries\nimport tensorflow as tf\nfrom tensorflow.keras.applications.resnet_v2 import ResNet50V2\nfrom tensorflow.keras.optimizers import RMSprop\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n\nfrom tensorflow.keras.layers import Dense, GlobalAveragePooling2D, Dropout\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.optimizers import Adam\n\n# basic libraries\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import classification_report, confusion_matrix\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport cv2\nimport os\n\nimport glob\nfrom glob import glob","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:09.344083Z","iopub.execute_input":"2023-06-06T12:56:09.344961Z","iopub.status.idle":"2023-06-06T12:56:18.959119Z","shell.execute_reply.started":"2023-06-06T12:56:09.344923Z","shell.execute_reply":"2023-06-06T12:56:18.95795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read the CSV File\n\nNext, we get the paths to access the image data and csv data. The csv data include various information about patients, such as biopsy, malignant cancer, and invasive cancer.","metadata":{}},{"cell_type":"code","source":"# the path to the image data\nRSNA_512_path = '/kaggle/input/rsna-breast-cancer-512-pngs'","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:18.961729Z","iopub.execute_input":"2023-06-06T12:56:18.962627Z","iopub.status.idle":"2023-06-06T12:56:18.970001Z","shell.execute_reply.started":"2023-06-06T12:56:18.962585Z","shell.execute_reply":"2023-06-06T12:56:18.968914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read the csv data.\ndf_train = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:18.971514Z","iopub.execute_input":"2023-06-06T12:56:18.971988Z","iopub.status.idle":"2023-06-06T12:56:19.110421Z","shell.execute_reply.started":"2023-06-06T12:56:18.971948Z","shell.execute_reply":"2023-06-06T12:56:19.109294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Attention\n\nGenerally, mammography is conducted with 2 views (from 2 angles), because **the position of breast cancer is various, although the upper outer quadrant of the breast is the most common site of breast cancer occurrence**. But the shape of the breast is almost the same regardless of different views. Moreover, the left and right breast generally have the same view. Therefore, we ignore the laterality and view for the machine learning purpose. The patient of the test data has no implant, so patients having implant can be excluded from the train data set. However, this patient has no risk of false positive for cancer because of implant. Although it is unknown whether the other patients in the hidden test data set have an implant, generally few patients have implant. Thus, implant appears to have little influence on this machine learning.","metadata":{}},{"cell_type":"code","source":"# the number of total patients\nlen(df_train)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:19.113461Z","iopub.execute_input":"2023-06-06T12:56:19.114116Z","iopub.status.idle":"2023-06-06T12:56:19.12121Z","shell.execute_reply.started":"2023-06-06T12:56:19.114074Z","shell.execute_reply":"2023-06-06T12:56:19.120155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# the number of patients having implant\nlen(df_train[df_train['implant'] == 1])","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:19.122746Z","iopub.execute_input":"2023-06-06T12:56:19.123187Z","iopub.status.idle":"2023-06-06T12:56:19.138043Z","shell.execute_reply.started":"2023-06-06T12:56:19.123148Z","shell.execute_reply":"2023-06-06T12:56:19.13688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# the number of patients without malignant cancer\nlen(df_train[df_train['cancer'] == 0])","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:19.139705Z","iopub.execute_input":"2023-06-06T12:56:19.140443Z","iopub.status.idle":"2023-06-06T12:56:19.154976Z","shell.execute_reply.started":"2023-06-06T12:56:19.140385Z","shell.execute_reply":"2023-06-06T12:56:19.15385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# the number of patient who took biopsy\nlen(df_train[df_train['biopsy'] == 1])","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:19.156725Z","iopub.execute_input":"2023-06-06T12:56:19.157108Z","iopub.status.idle":"2023-06-06T12:56:19.166381Z","shell.execute_reply.started":"2023-06-06T12:56:19.157071Z","shell.execute_reply":"2023-06-06T12:56:19.165118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# the number of patients having malignant cancer\nlen(df_train[df_train['cancer'] == 1])","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:19.168317Z","iopub.execute_input":"2023-06-06T12:56:19.168695Z","iopub.status.idle":"2023-06-06T12:56:19.177804Z","shell.execute_reply.started":"2023-06-06T12:56:19.168656Z","shell.execute_reply":"2023-06-06T12:56:19.176563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# the number of patients whose malignant cancer is invasive\nlen(df_train[df_train['invasive'] == 1])","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:19.179965Z","iopub.execute_input":"2023-06-06T12:56:19.180457Z","iopub.status.idle":"2023-06-06T12:56:19.189896Z","shell.execute_reply.started":"2023-06-06T12:56:19.180417Z","shell.execute_reply":"2023-06-06T12:56:19.188877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# the same as above\nlen(df_train[(df_train['cancer'] == 1) & (df_train['invasive'] == 1)])","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:19.195992Z","iopub.execute_input":"2023-06-06T12:56:19.196845Z","iopub.status.idle":"2023-06-06T12:56:19.205301Z","shell.execute_reply.started":"2023-06-06T12:56:19.196811Z","shell.execute_reply":"2023-06-06T12:56:19.204125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Most of the cases are normal or not-malignant cancer. Thus, physicians sometimes overlook cancer.\ndata = pd.DataFrame(np.concatenate([['Total'] * len(df_train) , ['Maglignant Cancer'] *  len(df_train[df_train['cancer'] == 1]), ['Invasive Cancer'] *  len(df_train[(df_train['cancer'] == 1) & (df_train['invasive'] == 1)])]), columns = [\"class\"])\n\nsns.countplot(x = 'class', data = data)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:19.206623Z","iopub.execute_input":"2023-06-06T12:56:19.207623Z","iopub.status.idle":"2023-06-06T12:56:19.51326Z","shell.execute_reply.started":"2023-06-06T12:56:19.207585Z","shell.execute_reply":"2023-06-06T12:56:19.512208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Around 3000 patient took biopsy and malignant cancer was found from some of them.\ndata = pd.DataFrame(np.concatenate([['Biopsy'] * len(df_train[df_train['biopsy'] == 1]) , ['Malignant Cancer'] *  len(df_train[df_train['cancer'] == 1]), ['Invasive Cancer'] *  len(df_train[(df_train['cancer'] == 1) & (df_train['invasive'] == 1)])]), columns = [\"class\"])\n\nsns.countplot(x = 'class', data = data)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:19.51656Z","iopub.execute_input":"2023-06-06T12:56:19.516849Z","iopub.status.idle":"2023-06-06T12:56:19.909238Z","shell.execute_reply.started":"2023-06-06T12:56:19.51682Z","shell.execute_reply":"2023-06-06T12:56:19.908163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# the number of not-malignant cancer cases from biopsy\nlen(df_train[(df_train['biopsy'] == 1) & (df_train['cancer'] == 0)])","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:19.911009Z","iopub.execute_input":"2023-06-06T12:56:19.91175Z","iopub.status.idle":"2023-06-06T12:56:19.921897Z","shell.execute_reply.started":"2023-06-06T12:56:19.911709Z","shell.execute_reply":"2023-06-06T12:56:19.920604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# the number of malignant cancer cases from biopsy\nlen(df_train[(df_train['biopsy'] == 1) & (df_train['cancer'] == 1)])","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:19.92418Z","iopub.execute_input":"2023-06-06T12:56:19.924676Z","iopub.status.idle":"2023-06-06T12:56:19.934757Z","shell.execute_reply.started":"2023-06-06T12:56:19.924638Z","shell.execute_reply":"2023-06-06T12:56:19.933546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 60% of biopsy resulted in not-malignanct cancer.\ndata = pd.DataFrame(np.concatenate([['Biopsy but Not Malignant'] * len(df_train[(df_train['biopsy'] == 1) & (df_train['cancer'] == 0)]) , ['Malignant Cancer'] *  len(df_train[df_train['cancer'] == 1]), ['Invasive Cancer'] *  len(df_train[(df_train['cancer'] == 1) & (df_train['invasive'] == 1)])]), columns = [\"class\"])\n\nsns.countplot(x = 'class', data = data)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:19.93667Z","iopub.execute_input":"2023-06-06T12:56:19.937032Z","iopub.status.idle":"2023-06-06T12:56:20.162125Z","shell.execute_reply.started":"2023-06-06T12:56:19.936995Z","shell.execute_reply":"2023-06-06T12:56:20.161162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Attention\n\nThis is a very difficult question whether **\"not-malignant cancer\" cases should be selected from (1)\"the whole patients without malignant cancer,\" (2)\"the whole patients without taking biopsy,\" or (3)\"patients with biopsy but without malignant cancer.\" The first choice is optimal for initial screening for general people**, most of whom are healthy.\n\nPeople generally take biopsy because not only of suspected mammography image of breast cancer, but also of their clinical findings or family history. Thus, images from **\"patients with biospy but without malignant cancer\" may be totally healthy or may include benign cancer or other diseases, such as inflammation**. Therefore, **the second choice may be optimal, if the initial screening is only conducted for detection of malignant breast cancer**, and the patients do not suffer from any other diseases.\n\n**The third choice is optimal** for special screening for suspected cases of malignanct cancer. This is particularly useful **when it is supected that a patient might suffer malignant cancer**. This AI would be used before biopsy is conducted.\n\nIn this competition, the purpose of screening is not suffiently clear, because the test data do not inlude information as to biopsy. This time we take the third choice, but another choice might be better for the purpose of the competition.","metadata":{}},{"cell_type":"code","source":"# The not-malignant cancer cases were limited into biopsy cases.\nDF_train = df_train[df_train['biopsy'] == 1].reset_index(drop = True)\nDF_train.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:20.163798Z","iopub.execute_input":"2023-06-06T12:56:20.164143Z","iopub.status.idle":"2023-06-06T12:56:20.18801Z","shell.execute_reply.started":"2023-06-06T12:56:20.164106Z","shell.execute_reply":"2023-06-06T12:56:20.187237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The number of positive (malignant) and negative (not-malignat) cases should be the same\n# to create a balanced dataset.\nDF_train = DF_train.groupby(['cancer']).apply(lambda x: x.sample(1158, replace = True)\n                                                      ).reset_index(drop = True)\nprint('New Data Size:', DF_train.shape[0])","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:20.189483Z","iopub.execute_input":"2023-06-06T12:56:20.190502Z","iopub.status.idle":"2023-06-06T12:56:20.210203Z","shell.execute_reply.started":"2023-06-06T12:56:20.190449Z","shell.execute_reply":"2023-06-06T12:56:20.209021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Generally, invasive cancer is confirmed by biopsy, not by mammography.\n# Maybe it is also extremely difficult for AI to detect invasive cancer from mammography.\ndata = pd.DataFrame(np.concatenate([['Biopsy but Not Malignant'] * len(DF_train[(DF_train['biopsy'] == 1) & (DF_train['cancer'] == 0)]) , ['Malignant Cancer'] *  len(DF_train[DF_train['cancer'] == 1]), ['Invasive Cancer'] *  len(DF_train[(DF_train['cancer'] == 1) & (DF_train['invasive'] == 1)])]), columns = [\"class\"])\n\nsns.countplot(x = 'class', data = data)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:20.212085Z","iopub.execute_input":"2023-06-06T12:56:20.21247Z","iopub.status.idle":"2023-06-06T12:56:20.435295Z","shell.execute_reply.started":"2023-06-06T12:56:20.212432Z","shell.execute_reply":"2023-06-06T12:56:20.434177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create the Path to Each Image","metadata":{}},{"cell_type":"code","source":"# Create the path to each image.\nfor i in range(len(DF_train)):\n    DF_train.loc[i, 'path'] = os.path.join(RSNA_512_path + '/' + str(DF_train.loc[i, 'patient_id']) + '_' + str(DF_train.loc[i, 'image_id']) + '.png')\nDF_train.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:20.43669Z","iopub.execute_input":"2023-06-06T12:56:20.437146Z","iopub.status.idle":"2023-06-06T12:56:21.487606Z","shell.execute_reply.started":"2023-06-06T12:56:20.437105Z","shell.execute_reply":"2023-06-06T12:56:21.486598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# a sample path\nDF_train.loc[0, 'path']","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:21.48928Z","iopub.execute_input":"2023-06-06T12:56:21.489666Z","iopub.status.idle":"2023-06-06T12:56:21.497234Z","shell.execute_reply.started":"2023-06-06T12:56:21.489627Z","shell.execute_reply":"2023-06-06T12:56:21.495902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# a sample image\nimg = cv2.imread(DF_train.loc[0, 'path'])\nplt.imshow(img, cmap = 'gray')","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:21.498753Z","iopub.execute_input":"2023-06-06T12:56:21.500109Z","iopub.status.idle":"2023-06-06T12:56:21.830263Z","shell.execute_reply.started":"2023-06-06T12:56:21.500063Z","shell.execute_reply":"2023-06-06T12:56:21.829181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:21.831711Z","iopub.execute_input":"2023-06-06T12:56:21.832179Z","iopub.status.idle":"2023-06-06T12:56:21.84061Z","shell.execute_reply.started":"2023-06-06T12:56:21.832137Z","shell.execute_reply":"2023-06-06T12:56:21.839362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img.shape","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:21.842212Z","iopub.execute_input":"2023-06-06T12:56:21.842722Z","iopub.status.idle":"2023-06-06T12:56:21.854302Z","shell.execute_reply.started":"2023-06-06T12:56:21.842684Z","shell.execute_reply":"2023-06-06T12:56:21.853089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Divide the Data into Training and Validation\n\nThis time we devide the data into training and validation data, because test data originally exist.","metadata":{}},{"cell_type":"code","source":"# Normal and cancer images must be equally distrubuted.\ntrain_df, val_df = train_test_split(DF_train, \n                                   test_size = 0.20, \n                                   random_state = 2018,\n                                   stratify = DF_train[['cancer']])\n\nprint('train', train_df.shape[0], 'validation', val_df.shape[0])\nprint('train', train_df['cancer'].value_counts())\nprint('validation', val_df['cancer'].value_counts())\ntrain_df.sample(1)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:21.856152Z","iopub.execute_input":"2023-06-06T12:56:21.856905Z","iopub.status.idle":"2023-06-06T12:56:21.895806Z","shell.execute_reply.started":"2023-06-06T12:56:21.856866Z","shell.execute_reply":"2023-06-06T12:56:21.894694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# training data\ntrain_df.head(5)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:21.897444Z","iopub.execute_input":"2023-06-06T12:56:21.897804Z","iopub.status.idle":"2023-06-06T12:56:21.915647Z","shell.execute_reply.started":"2023-06-06T12:56:21.897767Z","shell.execute_reply":"2023-06-06T12:56:21.914501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# validation data\nval_df.head(5)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:21.917174Z","iopub.execute_input":"2023-06-06T12:56:21.918499Z","iopub.status.idle":"2023-06-06T12:56:21.942261Z","shell.execute_reply.started":"2023-06-06T12:56:21.918459Z","shell.execute_reply":"2023-06-06T12:56:21.941276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Pick up normal images from the training data.\ntrain_df_normal = train_df[train_df['cancer'] == 0].reset_index(drop = True)\nprint(len(train_df_normal))\ntrain_df_normal.head(5)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:21.943891Z","iopub.execute_input":"2023-06-06T12:56:21.944214Z","iopub.status.idle":"2023-06-06T12:56:21.966243Z","shell.execute_reply.started":"2023-06-06T12:56:21.94418Z","shell.execute_reply":"2023-06-06T12:56:21.965233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Pick up cancer images from the training data.\ntrain_df_cancer = train_df[train_df['cancer'] == 1].reset_index(drop = True)\nprint(len(train_df_cancer))\ntrain_df_cancer.head(5)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:21.976894Z","iopub.execute_input":"2023-06-06T12:56:21.977478Z","iopub.status.idle":"2023-06-06T12:56:21.999948Z","shell.execute_reply.started":"2023-06-06T12:56:21.977438Z","shell.execute_reply":"2023-06-06T12:56:21.998813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Pick up normal images from the validation data.\nval_df_normal = val_df[val_df['cancer'] == 0].reset_index(drop = True)\nprint(len(val_df_normal))\nval_df_normal.head(5)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:22.001562Z","iopub.execute_input":"2023-06-06T12:56:22.001986Z","iopub.status.idle":"2023-06-06T12:56:22.023595Z","shell.execute_reply.started":"2023-06-06T12:56:22.001947Z","shell.execute_reply":"2023-06-06T12:56:22.022275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Pick up cancer images from the validation data.\nval_df_cancer = val_df[val_df['cancer'] == 1].reset_index(drop = True)\nprint(len(val_df_cancer))\nval_df_cancer.head(5)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:22.025605Z","iopub.execute_input":"2023-06-06T12:56:22.02644Z","iopub.status.idle":"2023-06-06T12:56:22.049512Z","shell.execute_reply.started":"2023-06-06T12:56:22.026401Z","shell.execute_reply":"2023-06-06T12:56:22.048468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**We have to store the 4 datasets in each folder to be used for machine learing of TensorFlow AI model.**","metadata":{}},{"cell_type":"code","source":"import shutil\n# Define the destination directory.\ndestination_dir = '/kaggle/working/train'\ndestination_dir_sub = '/kaggle/working/train/normal'\n\n# Create the destination directory if it doesn't exist.\nif not os.path.exists(destination_dir):\n    os.makedirs(destination_dir)\n\nif not os.path.exists(destination_dir_sub):\n    os.makedirs(destination_dir_sub)   \n    \n# Copy the images to the destination directory.\nfor path in train_df_normal['path']:\n    shutil.copy2(path, destination_dir_sub)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:22.050945Z","iopub.execute_input":"2023-06-06T12:56:22.051589Z","iopub.status.idle":"2023-06-06T12:56:25.612728Z","shell.execute_reply.started":"2023-06-06T12:56:22.051552Z","shell.execute_reply":"2023-06-06T12:56:25.611601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the destination directory.\ndestination_dir = '/kaggle/working/train'\ndestination_dir_sub = '/kaggle/working/train/cancer'\n\n# Create the destination directory if it doesn't exist.\nif not os.path.exists(destination_dir):\n    os.makedirs(destination_dir)\n\nif not os.path.exists(destination_dir_sub):\n    os.makedirs(destination_dir_sub)   \n    \n# Copy the images to the destination directory.\nfor path in train_df_cancer['path']:\n    shutil.copy2(path, destination_dir_sub)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:25.616325Z","iopub.execute_input":"2023-06-06T12:56:25.616685Z","iopub.status.idle":"2023-06-06T12:56:28.743047Z","shell.execute_reply.started":"2023-06-06T12:56:25.616652Z","shell.execute_reply":"2023-06-06T12:56:28.741879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the destination directory.\ndestination_dir = '/kaggle/working/val'\ndestination_dir_sub = '/kaggle/working/val/normal'\n\n# Create the destination directory if it doesn't exist.\nif not os.path.exists(destination_dir):\n    os.makedirs(destination_dir)\n\nif not os.path.exists(destination_dir_sub):\n    os.makedirs(destination_dir_sub)   \n    \n# Copy the images to the destination directory.\nfor path in val_df_normal['path']:\n    shutil.copy2(path, destination_dir_sub)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:28.747196Z","iopub.execute_input":"2023-06-06T12:56:28.747561Z","iopub.status.idle":"2023-06-06T12:56:29.377669Z","shell.execute_reply.started":"2023-06-06T12:56:28.747526Z","shell.execute_reply":"2023-06-06T12:56:29.376576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the destination directory.\ndestination_dir = '/kaggle/working/val'\ndestination_dir_sub = '/kaggle/working/val/cancer'\n\n# Create the destination directory if it doesn't exist.\nif not os.path.exists(destination_dir):\n    os.makedirs(destination_dir)\n\nif not os.path.exists(destination_dir_sub):\n    os.makedirs(destination_dir_sub)   \n    \n# Copy the images to the destination directory.\nfor path in val_df_cancer['path']:\n    shutil.copy2(path, destination_dir_sub)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:29.380104Z","iopub.execute_input":"2023-06-06T12:56:29.381088Z","iopub.status.idle":"2023-06-06T12:56:29.872006Z","shell.execute_reply.started":"2023-06-06T12:56:29.381045Z","shell.execute_reply":"2023-06-06T12:56:29.870852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Sample Images","metadata":{}},{"cell_type":"code","source":"import glob\nnormal_train_images = glob.glob('/kaggle/working/train/normal/*.png')\ncancer_train_images = glob.glob('/kaggle/working/train/cancer/*.png')","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:29.876205Z","iopub.execute_input":"2023-06-06T12:56:29.876561Z","iopub.status.idle":"2023-06-06T12:56:29.891134Z","shell.execute_reply.started":"2023-06-06T12:56:29.87653Z","shell.execute_reply":"2023-06-06T12:56:29.889895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# See normal images from the training dataset.\nfig, axes = plt.subplots(nrows = 2, ncols = 5, figsize = (15, 10), subplot_kw = {'xticks':[], 'yticks':[]})\nfor i, ax in enumerate(axes.flat):\n    img = cv2.imread(normal_train_images[i])\n    ax.imshow(img)\n    ax.set_title('Normal')\nfig.tight_layout()    \n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:29.894936Z","iopub.execute_input":"2023-06-06T12:56:29.895512Z","iopub.status.idle":"2023-06-06T12:56:31.043842Z","shell.execute_reply.started":"2023-06-06T12:56:29.895483Z","shell.execute_reply":"2023-06-06T12:56:31.041299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# See cancer images from the training dataset.\nfig, axes = plt.subplots(nrows = 2, ncols = 5, figsize = (15, 10), subplot_kw = {'xticks':[], 'yticks':[]})\nfor i, ax in enumerate(axes.flat):\n    img = cv2.imread(cancer_train_images[i])\n    ax.imshow(img)\n    ax.set_title('Cancer')\n    \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:31.045492Z","iopub.execute_input":"2023-06-06T12:56:31.04652Z","iopub.status.idle":"2023-06-06T12:56:31.924182Z","shell.execute_reply.started":"2023-06-06T12:56:31.046479Z","shell.execute_reply":"2023-06-06T12:56:31.923192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create Image data Generators\n\nThe dataset has already been divided into train and validation datasets, and each dataset includes normal and cancer image files. Thus, image data generators were easily created. ","metadata":{}},{"cell_type":"code","source":"train_datagen = ImageDataGenerator(rescale = 1./255.,\n                                   zoom_range = 0.2)\nval_datagen = ImageDataGenerator(rescale = 1./255.,)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:31.925737Z","iopub.execute_input":"2023-06-06T12:56:31.926433Z","iopub.status.idle":"2023-06-06T12:56:31.931974Z","shell.execute_reply.started":"2023-06-06T12:56:31.926394Z","shell.execute_reply":"2023-06-06T12:56:31.930803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_path = '/kaggle/working/train'\nval_path = '/kaggle/working/val'\n\ntrain_generator = train_datagen.flow_from_directory(\n    train_path,\n    target_size = (512, 512),\n    batch_size = 32,\n    class_mode = 'binary'\n)\nvalidation_generator = val_datagen.flow_from_directory(\n        val_path,\n        target_size = (512, 512),\n        batch_size = 16,\n        class_mode = 'binary'\n)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:31.933514Z","iopub.execute_input":"2023-06-06T12:56:31.934129Z","iopub.status.idle":"2023-06-06T12:56:32.154982Z","shell.execute_reply.started":"2023-06-06T12:56:31.934091Z","shell.execute_reply":"2023-06-06T12:56:32.153949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Define the Model (Transfer Learning)\n\nHere, ResNet50V2 was employed as the base model for transfer learning. It would be also possible to conduct fine tuning with the model. The choice of hyperparameters depends on the type of task by machine learning. Sigmoid, adam, and binary cross entropy were selected for the final activation function, optimizer, and loss function, respectively.","metadata":{}},{"cell_type":"code","source":"base_model = ResNet50V2(weights = 'imagenet', input_shape = (512, 512, 3), include_top = False)\n\nfor layer in base_model.layers:\n    layer.trainable = False\n    \nmodel = Sequential()\nmodel.add(base_model)\nmodel.add(GlobalAveragePooling2D())\nmodel.add(Dense(128, activation = 'relu'))\nmodel.add(Dropout(0.2))\nmodel.add(Dense(1, activation = 'sigmoid'))\n\nmodel.compile(optimizer = \"adam\", loss = 'binary_crossentropy', metrics = [\"accuracy\"])","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:32.156392Z","iopub.execute_input":"2023-06-06T12:56:32.156747Z","iopub.status.idle":"2023-06-06T12:56:37.883722Z","shell.execute_reply.started":"2023-06-06T12:56:32.156707Z","shell.execute_reply":"2023-06-06T12:56:37.882495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:37.888835Z","iopub.execute_input":"2023-06-06T12:56:37.891412Z","iopub.status.idle":"2023-06-06T12:56:37.959268Z","shell.execute_reply.started":"2023-06-06T12:56:37.891371Z","shell.execute_reply":"2023-06-06T12:56:37.958287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train the Model\n\nThe model was trained with the train and validation data. Early stopping was added to prevent overfitting. In fact, it might be preferable to set the larger number of epochs.","metadata":{}},{"cell_type":"code","source":"callback = tf.keras.callbacks.EarlyStopping(monitor = \"val_loss\", mode = \"min\", patience = 4)\n\nhistory = model.fit(train_generator, validation_data = validation_generator, steps_per_epoch = 20, epochs = 15, callbacks = callback)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T12:56:37.960841Z","iopub.execute_input":"2023-06-06T12:56:37.96183Z","iopub.status.idle":"2023-06-06T13:09:49.872656Z","shell.execute_reply.started":"2023-06-06T12:56:37.961779Z","shell.execute_reply":"2023-06-06T13:09:49.871542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Save the Model","metadata":{}},{"cell_type":"code","source":"model.save('mammography_pred_model.h5')","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:49.874078Z","iopub.execute_input":"2023-06-06T13:09:49.874467Z","iopub.status.idle":"2023-06-06T13:09:50.347594Z","shell.execute_reply.started":"2023-06-06T13:09:49.874425Z","shell.execute_reply":"2023-06-06T13:09:50.346469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Metrics\n\nThe accuracy and loss were calculated for both train and validation data in the model training process.","metadata":{}},{"cell_type":"code","source":"accuracy = history.history['accuracy']\nval_accuracy = history.history['val_accuracy']\n\nloss = history.history['loss']\nval_loss = history.history['val_loss']","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:50.349194Z","iopub.execute_input":"2023-06-06T13:09:50.349601Z","iopub.status.idle":"2023-06-06T13:09:50.356354Z","shell.execute_reply.started":"2023-06-06T13:09:50.349559Z","shell.execute_reply":"2023-06-06T13:09:50.355282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualizing Accuracy and Loss","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (15,10))\n\nplt.subplot(2, 2, 1)\nplt.plot(accuracy, label = \"Training Accuracy\")\nplt.plot(val_accuracy, label = \"Validation Accuracy\")\nplt.ylim(0.4, 1)\nplt.legend(['Train', 'Validation'], loc = 'upper left')\nplt.title(\"Training vs Validation Accuracy\")\nplt.xlabel('epoch')\nplt.ylabel('accuracy')\n\n\nplt.subplot(2, 2, 2)\nplt.plot(loss, label = \"Training Loss\")\nplt.plot(val_loss, label = \"Validation Loss\")\nplt.legend(['Train', 'Validation'], loc = 'upper left')\nplt.title(\"Training vs Validation Loss\")\nplt.xlabel('epoch')\nplt.ylabel('loss')","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:50.357941Z","iopub.execute_input":"2023-06-06T13:09:50.359794Z","iopub.status.idle":"2023-06-06T13:09:50.787195Z","shell.execute_reply.started":"2023-06-06T13:09:50.359752Z","shell.execute_reply":"2023-06-06T13:09:50.786167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Predictions\n\nIn order to evaluate the quality of the trained model, the outcome was predicted from the test (validation) data and compared with the observed value. **The threshold was set as 0.5 and the predicted value of more than 0.5 was treated as 1, which is positive.**","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.models import load_model\nmodel = load_model('/kaggle/working/mammography_pred_model.h5')","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:50.78888Z","iopub.execute_input":"2023-06-06T13:09:50.78963Z","iopub.status.idle":"2023-06-06T13:09:52.910076Z","shell.execute_reply.started":"2023-06-06T13:09:50.789589Z","shell.execute_reply":"2023-06-06T13:09:52.908946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred = model.predict(validation_generator)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:52.911704Z","iopub.execute_input":"2023-06-06T13:09:52.91208Z","iopub.status.idle":"2023-06-06T13:09:58.942268Z","shell.execute_reply.started":"2023-06-06T13:09:52.912042Z","shell.execute_reply":"2023-06-06T13:09:58.941087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# prediction by the AI\npred","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:58.943855Z","iopub.execute_input":"2023-06-06T13:09:58.944899Z","iopub.status.idle":"2023-06-06T13:09:58.96231Z","shell.execute_reply.started":"2023-06-06T13:09:58.944853Z","shell.execute_reply":"2023-06-06T13:09:58.961205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = []\nfor prob in pred:\n    if prob >= 0.5:\n        y_pred.append(1)\n    else:\n        y_pred.append(0)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:58.964099Z","iopub.execute_input":"2023-06-06T13:09:58.964551Z","iopub.status.idle":"2023-06-06T13:09:58.975799Z","shell.execute_reply.started":"2023-06-06T13:09:58.96451Z","shell.execute_reply":"2023-06-06T13:09:58.974712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(y_pred)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:58.977209Z","iopub.execute_input":"2023-06-06T13:09:58.977606Z","iopub.status.idle":"2023-06-06T13:09:58.986783Z","shell.execute_reply.started":"2023-06-06T13:09:58.977569Z","shell.execute_reply":"2023-06-06T13:09:58.985579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.Series(y_pred).value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:58.988585Z","iopub.execute_input":"2023-06-06T13:09:58.989071Z","iopub.status.idle":"2023-06-06T13:09:59.000213Z","shell.execute_reply.started":"2023-06-06T13:09:58.989034Z","shell.execute_reply":"2023-06-06T13:09:58.998891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_true = validation_generator.classes","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.002087Z","iopub.execute_input":"2023-06-06T13:09:59.002467Z","iopub.status.idle":"2023-06-06T13:09:59.007611Z","shell.execute_reply.started":"2023-06-06T13:09:59.002429Z","shell.execute_reply":"2023-06-06T13:09:59.006384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(y_true)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.009414Z","iopub.execute_input":"2023-06-06T13:09:59.00982Z","iopub.status.idle":"2023-06-06T13:09:59.020173Z","shell.execute_reply.started":"2023-06-06T13:09:59.009783Z","shell.execute_reply":"2023-06-06T13:09:59.018876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Confusion Matrix\n\nConfusion matrix was created with the predicted and observed values. The matrix indicated almost correct predictions by the trained model except that there were two cases observed as false positive. False positive means that a case is actually negative but predicted as positive.","metadata":{}},{"cell_type":"code","source":"cm = confusion_matrix(y_true, y_pred)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.022114Z","iopub.execute_input":"2023-06-06T13:09:59.022532Z","iopub.status.idle":"2023-06-06T13:09:59.029817Z","shell.execute_reply.started":"2023-06-06T13:09:59.022494Z","shell.execute_reply":"2023-06-06T13:09:59.028719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the class names.\nclass_names = ['Normal', 'Cancer']\n\n# Create the heatmap with class names as tick labels.\nax = sns.heatmap(cm, annot = True, fmt = '.0f', cmap = \"Blues\", annot_kws = {\"size\": 16},\\\n           xticklabels = class_names, yticklabels = class_names)\n\n# Set the axis labels.\nax.set_xlabel(\"Prediction\")\nax.set_ylabel(\"Truth\")","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.031101Z","iopub.execute_input":"2023-06-06T13:09:59.031843Z","iopub.status.idle":"2023-06-06T13:09:59.315558Z","shell.execute_reply.started":"2023-06-06T13:09:59.031805Z","shell.execute_reply":"2023-06-06T13:09:59.314581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Classification Report\n\nClassification report describes precision, recall, and f1-score as to each value. Many people regard accuracy and f1-score as the most important indicator to evaluate an AI model. This may be correct, but is not necessarily correct in the clinical field. **It must be considered why AI can be useful for and accepted by healthcare professionals. They expect that AI may be able to reduce their workload.** What does it mean to reduce their workload by AI? One idea is to **exclude lots of negative cases by AI that healthcare professionals would not have to see in order that they would be able to concentrate on the remaining positive cases to be treated**. Thus, it is required that the AI should be able to **exclude negative cases without false negatives**, which are actually positive but predicted as negative. **Otherwise, they would have to re-check the negative cases in order not to miss actually positive cases.** They must absolutely avoid clinical negligence! Therefore, **if the AI model does not give rise to false negative cases, the model will be considerably acceptable in the healthcare field** regardless of the accuracy or f1-score.","metadata":{}},{"cell_type":"code","source":"print(classification_report(y_true, y_pred))","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.316852Z","iopub.execute_input":"2023-06-06T13:09:59.317321Z","iopub.status.idle":"2023-06-06T13:09:59.332058Z","shell.execute_reply.started":"2023-06-06T13:09:59.317282Z","shell.execute_reply":"2023-06-06T13:09:59.330918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Analysing the Results\n\nIt is crucial in the medical field to analyze what kinds of cases were misclassified by the AI model, because medical misdiagnosis must be avoided as much as possible. Thus, it is necessary to identify false positive and false negative cases. Making a data frame and confusion table can visualize the results. As discussed above, false negative cases must be particularly avoided.","metadata":{}},{"cell_type":"code","source":"confusion = []\n\nfor i, j in zip(y_true, y_pred):\n  if i == 0 and j == 0:\n    confusion.append('TN')\n  elif i == 1 and j == 1:\n    confusion.append('TP')\n  elif i == 0 and j == 1:\n    confusion.append('FP')\n  else:\n    confusion.append('FN')","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.333759Z","iopub.execute_input":"2023-06-06T13:09:59.334134Z","iopub.status.idle":"2023-06-06T13:09:59.344115Z","shell.execute_reply.started":"2023-06-06T13:09:59.334096Z","shell.execute_reply":"2023-06-06T13:09:59.342651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(confusion)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.345645Z","iopub.execute_input":"2023-06-06T13:09:59.346504Z","iopub.status.idle":"2023-06-06T13:09:59.354889Z","shell.execute_reply.started":"2023-06-06T13:09:59.346465Z","shell.execute_reply":"2023-06-06T13:09:59.353474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"confusion_table = pd.DataFrame(data = confusion, columns = [\"Results\"])\nconfusion_table","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.356593Z","iopub.execute_input":"2023-06-06T13:09:59.357821Z","iopub.status.idle":"2023-06-06T13:09:59.374019Z","shell.execute_reply.started":"2023-06-06T13:09:59.357779Z","shell.execute_reply":"2023-06-06T13:09:59.373035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"confusion_table = pd.DataFrame({'Predicton':y_pred,\n                                'Truth': y_true,\n                                'Results': confusion})\nconfusion_table","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.375758Z","iopub.execute_input":"2023-06-06T13:09:59.376041Z","iopub.status.idle":"2023-06-06T13:09:59.392355Z","shell.execute_reply.started":"2023-06-06T13:09:59.376015Z","shell.execute_reply":"2023-06-06T13:09:59.390903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"confusion_table.Results == 'FP'","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.3939Z","iopub.execute_input":"2023-06-06T13:09:59.3946Z","iopub.status.idle":"2023-06-06T13:09:59.40532Z","shell.execute_reply.started":"2023-06-06T13:09:59.394562Z","shell.execute_reply":"2023-06-06T13:09:59.404412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# list of false positive images\nFPs = confusion_table[confusion_table['Results'] == 'FP']\nFPs","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.406629Z","iopub.execute_input":"2023-06-06T13:09:59.407559Z","iopub.status.idle":"2023-06-06T13:09:59.423662Z","shell.execute_reply.started":"2023-06-06T13:09:59.407489Z","shell.execute_reply":"2023-06-06T13:09:59.42226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"FPs.index","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.426016Z","iopub.execute_input":"2023-06-06T13:09:59.426887Z","iopub.status.idle":"2023-06-06T13:09:59.435443Z","shell.execute_reply.started":"2023-06-06T13:09:59.426848Z","shell.execute_reply":"2023-06-06T13:09:59.434173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# list of false negative images\nFNs = confusion_table[confusion_table['Results'] == 'FN']\nFNs","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.437599Z","iopub.execute_input":"2023-06-06T13:09:59.438112Z","iopub.status.idle":"2023-06-06T13:09:59.454682Z","shell.execute_reply.started":"2023-06-06T13:09:59.438075Z","shell.execute_reply":"2023-06-06T13:09:59.453547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"FNs.index","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.456382Z","iopub.execute_input":"2023-06-06T13:09:59.456748Z","iopub.status.idle":"2023-06-06T13:09:59.464575Z","shell.execute_reply.started":"2023-06-06T13:09:59.45671Z","shell.execute_reply":"2023-06-06T13:09:59.463358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Misclassficiation Cases\n\nIt is impoertant to pick up wrong cases judged by the AI and to analyze why the AI made wrong judgements for these images.","metadata":{}},{"cell_type":"code","source":"import glob\nval_images = glob.glob('/kaggle/working/val/*/*.png')","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.466039Z","iopub.execute_input":"2023-06-06T13:09:59.46737Z","iopub.status.idle":"2023-06-06T13:09:59.47548Z","shell.execute_reply.started":"2023-06-06T13:09:59.467328Z","shell.execute_reply":"2023-06-06T13:09:59.474303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# False positive imgages\nfig, axes = plt.subplots(nrows = 2, ncols = 5, figsize = (15, 10), subplot_kw = {'xticks':[], 'yticks':[]})\nfor i, ax in zip(FPs.index, axes.flat):\n    img = cv2.imread(val_images[i])\n    ax.imshow(img)\n    ax.set_title(\"False Positive Case\")\nfig.tight_layout()    \n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:09:59.476896Z","iopub.execute_input":"2023-06-06T13:09:59.480175Z","iopub.status.idle":"2023-06-06T13:10:00.589832Z","shell.execute_reply.started":"2023-06-06T13:09:59.480132Z","shell.execute_reply":"2023-06-06T13:10:00.58529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# False negative imgages\nfig, axes = plt.subplots(nrows = 2, ncols = 5, figsize = (15, 10), subplot_kw = {'xticks':[], 'yticks':[]})\nfor i, ax in zip(FNs.index, axes.flat):\n    img = cv2.imread(val_images[i])\n    ax.imshow(img)\n    ax.set_title(\"False Negative Case\")\nfig.tight_layout()    \n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:10:00.591248Z","iopub.execute_input":"2023-06-06T13:10:00.59231Z","iopub.status.idle":"2023-06-06T13:10:01.68011Z","shell.execute_reply.started":"2023-06-06T13:10:00.592254Z","shell.execute_reply":"2023-06-06T13:10:01.679205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Finetuning the Model (Unfreeying the Layers of the Model)","metadata":{}},{"cell_type":"code","source":"base_model = ResNet50V2(weights = 'imagenet', input_shape = (512, 512, 3), include_top = False)\n\nfor layer in base_model.layers:\n    layer.trainable = True # Change from False to True.\n    \nmodel = Sequential()\nmodel.add(base_model)\nmodel.add(GlobalAveragePooling2D())\nmodel.add(Dense(128, activation = 'relu'))\nmodel.add(Dropout(0.2))\nmodel.add(Dense(1, activation = 'sigmoid'))\n\nmodel.compile(optimizer = \"adam\", loss = 'binary_crossentropy', metrics = [\"accuracy\"])","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:10:01.68161Z","iopub.execute_input":"2023-06-06T13:10:01.682649Z","iopub.status.idle":"2023-06-06T13:10:04.059845Z","shell.execute_reply.started":"2023-06-06T13:10:01.682609Z","shell.execute_reply":"2023-06-06T13:10:04.058756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:10:04.061392Z","iopub.execute_input":"2023-06-06T13:10:04.061775Z","iopub.status.idle":"2023-06-06T13:10:04.111163Z","shell.execute_reply.started":"2023-06-06T13:10:04.061725Z","shell.execute_reply":"2023-06-06T13:10:04.110109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"callback = tf.keras.callbacks.EarlyStopping(monitor = \"val_loss\", mode = \"min\", patience = 4)\n\nhistory = model.fit(train_generator, validation_data = validation_generator, steps_per_epoch = 20, epochs = 50, callbacks = callback)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:10:04.11294Z","iopub.execute_input":"2023-06-06T13:10:04.113726Z","iopub.status.idle":"2023-06-06T13:17:12.582695Z","shell.execute_reply.started":"2023-06-06T13:10:04.113688Z","shell.execute_reply":"2023-06-06T13:17:12.581586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This time we use validation data to calculate the final accuracy.\nfinal_accuracy = model.evaluate_generator(validation_generator)[1]","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:17:12.584565Z","iopub.execute_input":"2023-06-06T13:17:12.584952Z","iopub.status.idle":"2023-06-06T13:17:17.562818Z","shell.execute_reply.started":"2023-06-06T13:17:12.584909Z","shell.execute_reply":"2023-06-06T13:17:17.561667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_accuracy","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:17:17.564336Z","iopub.execute_input":"2023-06-06T13:17:17.564723Z","iopub.status.idle":"2023-06-06T13:17:17.571379Z","shell.execute_reply.started":"2023-06-06T13:17:17.564681Z","shell.execute_reply":"2023-06-06T13:17:17.570295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Metrics","metadata":{}},{"cell_type":"code","source":"accuracy = history.history['accuracy']\nval_accuracy  = history.history['val_accuracy']\n\nloss = history.history['loss']\nval_loss = history.history['val_loss']","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:17:17.572734Z","iopub.execute_input":"2023-06-06T13:17:17.57355Z","iopub.status.idle":"2023-06-06T13:17:17.583748Z","shell.execute_reply.started":"2023-06-06T13:17:17.57351Z","shell.execute_reply":"2023-06-06T13:17:17.581775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualizing Accuracy and Loss","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize = (15,10))\n\nplt.subplot(2, 2, 1)\nplt.plot(accuracy, label = \"Training Accuracy\")\nplt.plot(val_accuracy, label = \"Validation Accuracy\")\nplt.ylim(0.4, 1)\nplt.legend(['Train', 'Validation'], loc = 'upper left')\nplt.title(\"Training vs Validation Accuracy\")\nplt.xlabel('epoch')\nplt.ylabel('accuracy')\n\n\nplt.subplot(2, 2, 2)\nplt.plot(loss, label = \"Training Loss\")\nplt.plot(val_loss, label = \"Validation Loss\")\nplt.legend(['Train', 'Validation'], loc = 'upper left')\nplt.title(\"Training vs Validation Loss\")\nplt.xlabel('epoch')\nplt.ylabel('loss')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:17:17.585481Z","iopub.execute_input":"2023-06-06T13:17:17.585844Z","iopub.status.idle":"2023-06-06T13:17:17.95724Z","shell.execute_reply.started":"2023-06-06T13:17:17.585806Z","shell.execute_reply":"2023-06-06T13:17:17.956226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Actually, Finetuning is worse than Transfer Learning, because the number of data images is small!","metadata":{}},{"cell_type":"markdown","source":"# Save the Model","metadata":{}},{"cell_type":"code","source":"model.save('mammography_pred_model_finetuning.h5')","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:17:17.958993Z","iopub.execute_input":"2023-06-06T13:17:17.95978Z","iopub.status.idle":"2023-06-06T13:17:19.37688Z","shell.execute_reply.started":"2023-06-06T13:17:17.95973Z","shell.execute_reply":"2023-06-06T13:17:19.375759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#VGG16\n\nfrom keras.applications.vgg16 import VGG16\n\n\n #FineTuning\nbase_model = VGG16(weights = 'imagenet', input_shape = (512, 512, 3), include_top = False)\n\nfor layer in base_model.layers:\n    layer.trainable = True # Change from False to True.\n    \nmodel = Sequential()\nmodel.add(base_model)\nmodel.add(GlobalAveragePooling2D())\nmodel.add(Dense(128, activation = 'relu'))\nmodel.add(Dropout(0.2))\nmodel.add(Dense(1, activation = 'sigmoid'))\n\nmodel.compile(optimizer = \"adam\", loss = 'binary_crossentropy', metrics = [\"accuracy\"])","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:17:19.378635Z","iopub.execute_input":"2023-06-06T13:17:19.379002Z","iopub.status.idle":"2023-06-06T13:17:20.196086Z","shell.execute_reply.started":"2023-06-06T13:17:19.378962Z","shell.execute_reply":"2023-06-06T13:17:20.195046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:17:20.197786Z","iopub.execute_input":"2023-06-06T13:17:20.198154Z","iopub.status.idle":"2023-06-06T13:17:20.204278Z","shell.execute_reply.started":"2023-06-06T13:17:20.198114Z","shell.execute_reply":"2023-06-06T13:17:20.203266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"callback = tf.keras.callbacks.EarlyStopping(monitor = \"val_loss\", mode = \"min\", patience = 4)\n\nhistory = model.fit(train_generator, validation_data = validation_generator, steps_per_epoch = 10, epochs = 30, callbacks = callback)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:17:20.205753Z","iopub.execute_input":"2023-06-06T13:17:20.206412Z","iopub.status.idle":"2023-06-06T13:22:32.579707Z","shell.execute_reply.started":"2023-06-06T13:17:20.206368Z","shell.execute_reply":"2023-06-06T13:22:32.578588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This time we use validation data to calculate the final accuracy.\nfinal_accuracy = model.evaluate_generator(validation_generator)[1]","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:22:32.581377Z","iopub.execute_input":"2023-06-06T13:22:32.581745Z","iopub.status.idle":"2023-06-06T13:22:37.786361Z","shell.execute_reply.started":"2023-06-06T13:22:32.581705Z","shell.execute_reply":"2023-06-06T13:22:37.785091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_accuracy","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:22:37.787798Z","iopub.execute_input":"2023-06-06T13:22:37.790425Z","iopub.status.idle":"2023-06-06T13:22:37.800092Z","shell.execute_reply.started":"2023-06-06T13:22:37.790375Z","shell.execute_reply":"2023-06-06T13:22:37.798801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#VGG19\nfrom tensorflow.keras.applications.vgg19 import VGG19\nbase_model = VGG19(weights = 'imagenet', input_shape = (512, 512, 3), include_top = False)\n\nfor layer in base_model.layers:\n    layer.trainable = True # Change from False to True.\n    \nmodel = Sequential()\nmodel.add(base_model)\nmodel.add(GlobalAveragePooling2D())\nmodel.add(Dense(128, activation = 'relu'))\nmodel.add(Dropout(0.2))\nmodel.add(Dense(1, activation = 'sigmoid'))\n\nmodel.compile(optimizer = \"adam\", loss = 'binary_crossentropy', metrics = [\"accuracy\"])","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:39:38.629381Z","iopub.execute_input":"2023-06-06T13:39:38.629774Z","iopub.status.idle":"2023-06-06T13:39:39.152192Z","shell.execute_reply.started":"2023-06-06T13:39:38.629738Z","shell.execute_reply":"2023-06-06T13:39:39.151051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:39:46.375167Z","iopub.execute_input":"2023-06-06T13:39:46.376158Z","iopub.status.idle":"2023-06-06T13:39:46.405536Z","shell.execute_reply.started":"2023-06-06T13:39:46.376102Z","shell.execute_reply":"2023-06-06T13:39:46.404687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"callback = tf.keras.callbacks.EarlyStopping(monitor = \"val_loss\", mode = \"min\", patience = 4)\n\nhistory = model.fit(train_generator, validation_data = validation_generator, steps_per_epoch = 10, epochs = 30, callbacks = callback)","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:39:56.169717Z","iopub.execute_input":"2023-06-06T13:39:56.170417Z","iopub.status.idle":"2023-06-06T13:46:33.350978Z","shell.execute_reply.started":"2023-06-06T13:39:56.170377Z","shell.execute_reply":"2023-06-06T13:46:33.349868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This time we use validation data to calculate the final accuracy.\nfinal_accuracy = model.evaluate_generator(validation_generator)[1]\n","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:55:06.532358Z","iopub.execute_input":"2023-06-06T13:55:06.533197Z","iopub.status.idle":"2023-06-06T13:55:12.143189Z","shell.execute_reply.started":"2023-06-06T13:55:06.533155Z","shell.execute_reply":"2023-06-06T13:55:12.142054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_accuracy","metadata":{"execution":{"iopub.status.busy":"2023-06-06T13:55:34.764348Z","iopub.execute_input":"2023-06-06T13:55:34.764736Z","iopub.status.idle":"2023-06-06T13:55:34.771815Z","shell.execute_reply.started":"2023-06-06T13:55:34.7647Z","shell.execute_reply":"2023-06-06T13:55:34.770722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conclusion\n\nThis AI model may be useful for general physicians who occasionally see female patients for the screening purpose. Further improvement is required to prevent misdiagnosis and unnecessary treatment.","metadata":{}}]}