{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n","metadata":{}},{"cell_type":"markdown","source":"\n# <span style=\"color:teal\"> RSNA Screening Mammography Breast Cancer Detection<a class=\"anchor\"  id=\"projectTopic\"></a></span>\n### <span style=\"color:teal\"> Detect breast cancers in screening mammograms <a class=\"anchor\"  id=\"detect\"></a></span>","metadata":{}},{"cell_type":"markdown","source":"# <span style=\"color:teal\">Developed by : <a class=\"anchor\"  id=\"detect\"></a></span>\n\n* [Gebreyowhans Hailekiros](https://www.kaggle.com/gebreyowhansbahre/)\n* [Mahbub Hasan](https://www.kaggle.com/mahbubhasanuunical)\n* [Muhammad Danish Sadiq](https://www.kaggle.com/muhammaddanishsadiq/)","metadata":{}},{"cell_type":"markdown","source":"* This notebook is developed to make inference on the trained model [RSNA_BCD_Train[TPU_VM]_EfficientNet](https://www.kaggle.com/code/gebreyowhansbahre/rsna-bcd-train-tpu-vm-bc86d1) using unseen test data and submit it to competition. Since we trained to models during training; we have to use average of the predictions of the two models to get the final prediction.","metadata":{}},{"cell_type":"markdown","source":"# <span style=\"color:teal\"> Notebooks <a class=\"anchor\"  id=\"notebooks\"></a></span>\n* Image preprocessing Notebook: [RSNA_BCD_DICOM_PNG_ROI](https://www.kaggle.com/code/gebreyowhansh/rsna-bcd-dicom-png-roi)\n\n* Training Notebook: [RSNA_BCD_Train[TPU_VM]_EfficientNet](https://www.kaggle.com/code/gebreyowhansbahre/rsna-bcd-train-tpu-vm-bc86d1)\n\n* Test Notebook: [RSNA-BCD-GPU-TEST_EfficientNet](https://www.kaggle.com/code/gebreyowhansbahre/rsna-bcd-gpu-test/edit/run/128324368)\n\n ","metadata":{}},{"cell_type":"markdown","source":"# <span style=\"color:teal\">1. Imporing and installing libraries <a class=\"anchor\"  id=\"libraries\"></a></span>","metadata":{}},{"cell_type":"code","source":"from IPython.display import clear_output\n!pip install -qU --upgrade pip\nclear_output()\n!pip install -qU /kaggle/input/whl-files/dicomsdl-0.109.1-cp37-cp37m-manylinux_2_12_x86_64.manylinux2010_x86_64.whl/dicomsdl-0.109.1-cp37-cp37m-manylinux_2_12_x86_64.manylinux2010_x86_64.whl\n!pip install -qU /kaggle/input/whl-files/pylibjpeg_libjpeg-1.3.2-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl/pylibjpeg_libjpeg-1.3.2-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl\n!pip install -qU /kaggle/input/whl-files/python_gdcm-3.0.20-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl/python_gdcm-3.0.20-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:54:47.53092Z","iopub.execute_input":"2023-06-18T22:54:47.531468Z","iopub.status.idle":"2023-06-18T22:59:03.204238Z","shell.execute_reply.started":"2023-06-18T22:54:47.531416Z","shell.execute_reply":"2023-06-18T22:59:03.202936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os, random, cv2, dicomsdl\nimport numpy as np\nimport pandas as pd\nfrom IPython import display as ipd\n\nfrom tqdm import tqdm\nfrom joblib import Parallel, delayed\nfrom matplotlib import pyplot as plt\nfrom mpl_toolkits.axes_grid1 import ImageGrid\n\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.python.client import device_lib\nfrom kaggle_datasets import KaggleDatasets\nfrom kaggle_secrets import UserSecretsClient\n\n\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3'\ntf.get_logger().setLevel('ERROR')\n","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:03.209125Z","iopub.execute_input":"2023-06-18T22:59:03.209476Z","iopub.status.idle":"2023-06-18T22:59:03.221352Z","shell.execute_reply.started":"2023-06-18T22:59:03.209439Z","shell.execute_reply":"2023-06-18T22:59:03.220333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('np:', np.__version__)\nprint('pd:', pd.__version__)\nprint('tf:',tf.__version__)","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:03.223119Z","iopub.execute_input":"2023-06-18T22:59:03.223556Z","iopub.status.idle":"2023-06-18T22:59:03.234867Z","shell.execute_reply.started":"2023-06-18T22:59:03.223517Z","shell.execute_reply":"2023-06-18T22:59:03.233447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:teal\"> 2. Create basic configuration class <a class=\"anchor\"  id=\"configuration\"></a></span>\n * Configuration class consisting of information about project","metadata":{}},{"cell_type":"code","source":"class Config:\n    \n    def __init__(self):\n        \n        self.debug = False\n        \n        self.device = 'GPU'\n        self.num_devices = 1\n        \n        self.seed = 150\n        \n        self.models_path = '/kaggle/input/rsna-bcd-train-tpu-vm/'\n        self.weights = \"/kaggle/input/whl-files/efficientnetb3_notop.h5/efficientnetb3_notop.h5\"\n        \n        self.input_data_path = '/kaggle/input/rsna-breast-cancer-detection/'\n        \n        self.output_path = '/kaggle/working/'\n        self.test_images_path=self.output_path+'test_images/'\n        \n        self.test_path = self.input_data_path + 'test_images/'\n        self.test_csv = self.input_data_path + 'test.csv'\n        \n        self.sub_csv = '/kaggle/working/submission.csv'  \n        self.sample_sub_csv = self.input_data_path + 'sample_submission.csv'\n        self.models = ['model_3.h5','model_4.h5']\n        \n        self.threshold = 0.6\n        self.batch_size=32\n        self.epochs=10\n        self.dropout=0.4\n        self.optimizer='adam'\n        self.loss='binary_crossentropy'\n        \n        self.img_size =(512,256)\n        self.resize_dim = 512\n        self.img_ext = 'png'\n\nconfig = Config()","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:03.237164Z","iopub.execute_input":"2023-06-18T22:59:03.237604Z","iopub.status.idle":"2023-06-18T22:59:03.249304Z","shell.execute_reply.started":"2023-06-18T22:59:03.237567Z","shell.execute_reply":"2023-06-18T22:59:03.247177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:teal\"> 3. Device Configurations <a class=\"anchor\"  id=\"configuration\"></a></span>","metadata":{}},{"cell_type":"code","source":"num_devices = len(tf.config.list_physical_devices('GPU'))\n\nif num_devices > 1:\n    config.num_devices = num_devices\n    strategy = tf.distribute.MirroredStrategy()\n    print(f'Running on {num_devices} GPU devices')\nelif num_devices == 1:\n    strategy = tf.distribute.get_strategy()\n    print(f'Running on {num_devices} GPU device')\nelse:\n    strategy = tf.distribute.get_strategy()\n    config.device = 'CPU'\n    print(f'Running on CPU')\n\ntf.config.optimizer.set_jit(True)\ntf.keras.mixed_precision.set_global_policy(policy=\"float32\")\nconfig.batch_size = config.batch_size * config.num_devices","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:03.25361Z","iopub.execute_input":"2023-06-18T22:59:03.254228Z","iopub.status.idle":"2023-06-18T22:59:03.270173Z","shell.execute_reply.started":"2023-06-18T22:59:03.254199Z","shell.execute_reply":"2023-06-18T22:59:03.269128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:teal\">4.Functions to convert ,extract ROI and save dicom images <a class=\"anchor\"  id=\"utilityfunctions\"></a></span>\n ","metadata":{}},{"cell_type":"markdown","source":"### <span style=\"color:teal\">4.1 Dicom to png <a class=\"anchor\"  id=\"dicomtopng\"></a></span>","metadata":{}},{"cell_type":"code","source":"def dicom_to_png(dicom_path):\n    dicom = dicomsdl.open(dicom_path)\n    image = dicom.pixelData(storedvalue=False)\n    image = image - np.min(image)\n    image = image / np.max(image)\n\n    if dicom.PhotometricInterpretation == 'MONOCHROME1':\n        image = 1.0 - image\n        \n    image = cv2.resize(image, (config.resize_dim, config.resize_dim), interpolation=cv2.INTER_LINEAR)\n    image = (image * 255).astype(np.uint8)\n    return image\n","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:03.271498Z","iopub.execute_input":"2023-06-18T22:59:03.272548Z","iopub.status.idle":"2023-06-18T22:59:03.280924Z","shell.execute_reply.started":"2023-06-18T22:59:03.272501Z","shell.execute_reply":"2023-06-18T22:59:03.279651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:teal\">4.2 Extract region of interest <a class=\"anchor\"  id=\"regionofInterest\"></a></span>","metadata":{}},{"cell_type":"code","source":"def png_to_roi(image, image_path):\n    bin_image = cv2.threshold(image, 20, 255, cv2.THRESH_BINARY)[1]\n    contours, _ = cv2.findContours(bin_image, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_NONE)\n    contour = max(contours, key=cv2.contourArea)\n    ys = contour.squeeze()[:, 0]\n    xs = contour.squeeze()[:, 1]\n    roi = image[np.min(xs):np.max(xs), np.min(ys):np.max(ys)]\n    return cv2.resize(roi, config.img_size[::-1], interpolation=cv2.INTER_LINEAR)","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:03.282294Z","iopub.execute_input":"2023-06-18T22:59:03.283263Z","iopub.status.idle":"2023-06-18T22:59:03.29283Z","shell.execute_reply.started":"2023-06-18T22:59:03.283203Z","shell.execute_reply":"2023-06-18T22:59:03.291696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def process(dicom_path, image_path):\n    image = dicom_to_png(dicom_path)\n    os.makedirs(os.path.dirname(image_path), exist_ok=True)\n    image = png_to_roi(image, image_path)\n    cv2.imwrite(image_path, image)","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:03.294358Z","iopub.execute_input":"2023-06-18T22:59:03.294915Z","iopub.status.idle":"2023-06-18T22:59:03.303055Z","shell.execute_reply.started":"2023-06-18T22:59:03.294869Z","shell.execute_reply":"2023-06-18T22:59:03.302264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:teal\"> 5. Dataframe informations <a class=\"anchor\"  id=\"dataframeinfo\"></a></span>","metadata":{}},{"cell_type":"code","source":"print('\\n  Train:')\ntrain_df = pd.read_csv(config.input_data_path + 'train.csv')\ndisplay(train_df.head())\n\nprint('\\n  Test :')\ntest_df = pd.read_csv(config.test_csv)\ntest_df['dicom_path'] = config.test_path + test_df['patient_id'].astype(str) + '/' + test_df['image_id'].astype(str) + '.dcm'\ntest_df['image_path'] = config.test_images_path + test_df['patient_id'].astype(str) + '/' + test_df['image_id'].astype(str) + '.png'\ndisplay(test_df.head())","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:03.304408Z","iopub.execute_input":"2023-06-18T22:59:03.305205Z","iopub.status.idle":"2023-06-18T22:59:03.422388Z","shell.execute_reply.started":"2023-06-18T22:59:03.305167Z","shell.execute_reply":"2023-06-18T22:59:03.421121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:03.424097Z","iopub.execute_input":"2023-06-18T22:59:03.4252Z","iopub.status.idle":"2023-06-18T22:59:03.44042Z","shell.execute_reply.started":"2023-06-18T22:59:03.425157Z","shell.execute_reply":"2023-06-18T22:59:03.439152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Parallel(n_jobs=4, backend='threading')(delayed(process)(dicom_path, image_path) \n                   for dicom_path, image_path in tqdm(zip(test_df['dicom_path'], \n                                                          test_df['image_path'])))\nclear_output()","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:03.442317Z","iopub.execute_input":"2023-06-18T22:59:03.442699Z","iopub.status.idle":"2023-06-18T22:59:05.959338Z","shell.execute_reply.started":"2023-06-18T22:59:03.44266Z","shell.execute_reply":"2023-06-18T22:59:05.958306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n#### <span style=\"color:teal\"> Check If Data Exists ? <a class=\"anchor\"  id=\"existence\"></a></span>","metadata":{}},{"cell_type":"code","source":"tf.io.gfile.exists(test_df.dicom_path.iloc[0])","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:05.963034Z","iopub.execute_input":"2023-06-18T22:59:05.963659Z","iopub.status.idle":"2023-06-18T22:59:05.971997Z","shell.execute_reply.started":"2023-06-18T22:59:05.963626Z","shell.execute_reply":"2023-06-18T22:59:05.970809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:teal\">6. Defining input features <a class=\"anchor\"  id=\"inputefeatures\"></a></span>","metadata":{}},{"cell_type":"code","source":"train_df_processed = pd.read_csv(config.models_path+'train_df_processed.csv')\n    \nprocessed_train_df_columns = train_df_processed.columns\nprocessed_train_df_columns = np.append(processed_train_df_columns, 'prediction_id')\nprint(processed_train_df_columns)\n\ntest_df = pd.DataFrame(test_df, columns=processed_train_df_columns).fillna(0.0)\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:05.973783Z","iopub.execute_input":"2023-06-18T22:59:05.974564Z","iopub.status.idle":"2023-06-18T22:59:06.290705Z","shell.execute_reply.started":"2023-06-18T22:59:05.974525Z","shell.execute_reply":"2023-06-18T22:59:06.289434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:06.297417Z","iopub.execute_input":"2023-06-18T22:59:06.297761Z","iopub.status.idle":"2023-06-18T22:59:06.316317Z","shell.execute_reply.started":"2023-06-18T22:59:06.297729Z","shell.execute_reply":"2023-06-18T22:59:06.31488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"exclude_cols = ['patient_id', 'image_id', 'site_id','machine_id', 'cancer', 'age', 'stratify', \n                'image_path', 'fold', 'prediction_id']\ninput_features = test_df.columns.difference(exclude_cols)\n\ninput_features","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:06.318386Z","iopub.execute_input":"2023-06-18T22:59:06.319123Z","iopub.status.idle":"2023-06-18T22:59:06.330878Z","shell.execute_reply.started":"2023-06-18T22:59:06.31908Z","shell.execute_reply":"2023-06-18T22:59:06.329527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:teal\">7. Normalize the inpute features <a class=\"anchor\"  id=\"nomalization\"></a></span>","metadata":{}},{"cell_type":"code","source":"test_df[input_features] = (test_df[input_features] - train_df_processed[input_features].mean()) / train_df_processed[input_features].std()\ntest_df[input_features] = test_df[input_features].astype('float32')\n\ntest_df[input_features]","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:06.332867Z","iopub.execute_input":"2023-06-18T22:59:06.333994Z","iopub.status.idle":"2023-06-18T22:59:06.394721Z","shell.execute_reply.started":"2023-06-18T22:59:06.333959Z","shell.execute_reply":"2023-06-18T22:59:06.393446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:teal\">8. Data Pipeline <a class=\"anchor\"  id=\"pipline\"></a></span>\n\n### <span style=\"color:teal\">8.1 Decode images <a class=\"anchor\"  id=\"decode\"></a></span>\n * **tf.image.decode_png()** and **tf.image.decode_jpeg** are TensorFlow functions that decodes a PNG-encoded   image into a tensor of type uint8.\n \n * These functions takes the following arguments: \n     * **image**: A string tensor containing a PNG or jpeg -encoded image.\n     * **channels**: An optional integer specifying the number of color channels in the decoded image. By  default, this is set to 3,\n     \n\n* The function returns a **uint8** tensor representing the decoded image with the shape of (height, width, channels) \n ","metadata":{}},{"cell_type":"code","source":"def decode_image(label=True, img_size=config.img_size, ext=config.img_ext):\n    \n    def _decode_image(Input_Image, label=None):\n        image = tf.io.read_file(Input_Image['input_image'])\n        \n        if ext == 'png':\n            ## PNG-encoded image into a tensor of type uint8.\n            image = tf.image.decode_png(image, channels=3)\n        elif ext in ['jpg', 'jpeg']:\n            ## jpeg-encoded image into a tensor of type uint8.\n            image = tf.image.decode_jpeg(image, channels=3)\n        else:\n            raise ValueError(\"Image extension not supported\")\n        \n        ## explicit size needed for TPU\n        image = tf.image.resize(image, img_size)\n        ## convert image to floats in [0, 1] range\n        image = tf.cast(image, tf.float32) / 255.0\n        \n        Input_Image['input_image'] = image\n        \n        if label is None:\n            return Input_Image\n        else:\n            return Input_Image, label\n    \n    if label:\n        return _decode_image\n    else:\n        return lambda x: _decode_image(x, None)","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:06.396673Z","iopub.execute_input":"2023-06-18T22:59:06.39713Z","iopub.status.idle":"2023-06-18T22:59:06.407894Z","shell.execute_reply.started":"2023-06-18T22:59:06.397091Z","shell.execute_reply":"2023-06-18T22:59:06.406387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:teal\">8.2 Data augumentation <a class=\"anchor\"  id=\"augumentation\"></a></span>\n* Using Augmentations to reduce overfitting and make model more robust by :\n * 1. random_flip_left_right for applying position transforamtion\n * 2. perofrming some random_hue,random_saturation,random_contrast,random_brightness for pixel transforamtion","metadata":{}},{"cell_type":"code","source":"def data_augment(label=True):\n    def _augment(Input_Image, label=None):\n        image = Input_Image['input_image']\n        #position transforamtion\n        image = tf.image.random_flip_left_right(image)\n        # pixel-augment\n        image = tf.image.random_hue(image, config.hue)\n        image = tf.image.random_saturation(image,config.sat[0], config.sat[1])\n        image = tf.image.random_contrast(image,config.cont[0], config.cont[1])\n        image = tf.image.random_brightness(image,config.bri)\n        Input_Image['input_image'] = image\n        if label is not None:\n            return Input_Image, label\n        else:\n            return Input_Image\n\n    if label:\n        return _augment\n    else:\n        return lambda x: _augment(x, None)","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:06.409917Z","iopub.execute_input":"2023-06-18T22:59:06.410709Z","iopub.status.idle":"2023-06-18T22:59:06.42307Z","shell.execute_reply.started":"2023-06-18T22:59:06.410651Z","shell.execute_reply":"2023-06-18T22:59:06.421286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:teal\">8.3 Build tf.data.dataset <a class=\"anchor\"  id=\"dataset\"></a></span>","metadata":{}},{"cell_type":"code","source":"def build_dataset(df,input_features,image_size=config.img_size,batch_size=config.batch_size, \n                  label=True,cache=False,ext=config.img_ext):\n    \n    decode = decode_image(label, img_size=image_size, ext=ext)\n    input_data = {'input_image': df['image_path'].values, 'input_features': df[input_features].values}\n    \n    if label:\n        label_data = df['cancer'].apply(lambda x: int(x)).values\n        dataset = tf.data.Dataset.from_tensor_slices((input_data, label_data))\n    else:\n        dataset = tf.data.Dataset.from_tensor_slices(input_data)\n        \n    dataset = dataset.map(decode, num_parallel_calls=tf.data.AUTOTUNE)\n    \n    if cache:\n        dataset = dataset.cache()\n        \n    dataset = dataset.batch(batch_size)\n    dataset = dataset.prefetch(tf.data.AUTOTUNE)\n    return dataset\n","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:06.424704Z","iopub.execute_input":"2023-06-18T22:59:06.425202Z","iopub.status.idle":"2023-06-18T22:59:06.437008Z","shell.execute_reply.started":"2023-06-18T22:59:06.425159Z","shell.execute_reply":"2023-06-18T22:59:06.435844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_dataset = build_dataset(test_df, input_features, \n                             batch_size=config.batch_size, \n                             label=False, \n                             cache=False)","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:06.438656Z","iopub.execute_input":"2023-06-18T22:59:06.439064Z","iopub.status.idle":"2023-06-18T22:59:06.471979Z","shell.execute_reply.started":"2023-06-18T22:59:06.439018Z","shell.execute_reply":"2023-06-18T22:59:06.471003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# for item in test_dataset.take(1):\n#     print(item)","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:06.47362Z","iopub.execute_input":"2023-06-18T22:59:06.473973Z","iopub.status.idle":"2023-06-18T22:59:06.478763Z","shell.execute_reply.started":"2023-06-18T22:59:06.473926Z","shell.execute_reply":"2023-06-18T22:59:06.477652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:teal\">9. Evaluation metric(f1 score)<a class=\"anchor\"  id=\"dataset\"></a></span>","metadata":{}},{"cell_type":"code","source":"def p_f1(y_true, y_pred):\n    tp = tf.reduce_sum(tf.cast(y_true * y_pred, tf.float32))\n    tn = tf.reduce_sum(tf.cast((1 - y_true) * (1 - y_pred), tf.float32))\n    fp = tf.reduce_sum(tf.cast((1 - y_true) * y_pred, tf.float32))\n    fn = tf.reduce_sum(tf.cast(y_true * (1 - y_pred), tf.float32))\n\n    p = tp / (tp + fp + tf.keras.backend.epsilon())\n    r = tp / (tp + fn + tf.keras.backend.epsilon())\n    \n    numerator=2*p*r\n    denominator=p+r\n    \n    f1 = numerator / denominator\n    f1 = tf.where(tf.math.is_nan(f1), tf.zeros_like(f1), f1)\n    \n    return tf.reduce_mean(f1)","metadata":{"execution":{"iopub.status.busy":"2023-06-18T22:59:06.480421Z","iopub.execute_input":"2023-06-18T22:59:06.481127Z","iopub.status.idle":"2023-06-18T22:59:06.491378Z","shell.execute_reply.started":"2023-06-18T22:59:06.481084Z","shell.execute_reply":"2023-06-18T22:59:06.490217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:teal\">10. Defining the model <a class=\"anchor\"  id=\"model\"></a></span>","metadata":{}},{"cell_type":"code","source":"def build_model(input_features, \n                loss=config.loss, \n                dropout=config.dropout, \n                optimizer=config.optimizer, \n                img_size=config.img_size):\n    with strategy.scope():\n        input_image = tf.keras.layers.Input(shape=(*img_size,3), \n                                            name='input_image')\n        input_features = tf.keras.layers.Input(shape=[len(input_features)],\n                                               name='input_features')\n        \n        efficientNetB3Base_model = tf.keras.applications.EfficientNetB3(input_shape=(*img_size,3),\n                                                 include_top=False, \n                                                 drop_connect_rate=0.3,\n                                                 weights=config.weights)(input_image)\n        #reduce the spatial dimensions of a feature map produced by a pre-trained EfficientNetB3\n        x = tf.keras.layers.GlobalAveragePooling2D()(efficientNetB3Base_model)\n        \n        x = tf.keras.layers.Dense(128,activation=\"relu\")(x)\n        _output = tf.keras.layers.Dropout(dropout)(x)\n        #normalizes its input by subtracting the batch mean and dividing by the batch standard deviation\n        _output = tf.keras.layers.BatchNormalization()(_output)\n        _output = tf.keras.layers.Dense(64,activation=\"relu\")(_output)\n        _output = tf.keras.layers.Dropout(dropout)(_output)\n        _output = tf.keras.layers.BatchNormalization()(_output)\n        _output = tf.keras.layers.Dense(32,activation=\"relu\")(_output)\n        _output = tf.keras.layers.BatchNormalization()(_output)\n        _output = tf.keras.layers.Dense(16,activation=\"relu\")(_output)\n        _output = tf.keras.layers.BatchNormalization()(_output)  \n        _output = tf.keras.layers.Concatenate()([_output, input_features])\n        _output = tf.keras.layers.Dense(1, activation='sigmoid')(_output)\n        \n        model = tf.keras.Model(inputs=[input_image, input_features], outputs=_output)\n        \n        model.compile(optimizer=optimizer,\n                      loss=loss,\n                      metrics=['accuracy', \n                               p_f1])\n\n        return model","metadata":{"execution":{"iopub.status.busy":"2023-06-18T23:04:04.210414Z","iopub.execute_input":"2023-06-18T23:04:04.210802Z","iopub.status.idle":"2023-06-18T23:04:04.227378Z","shell.execute_reply.started":"2023-06-18T23:04:04.210765Z","shell.execute_reply":"2023-06-18T23:04:04.226087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = build_model(input_features,config.loss,config.dropout,config.optimizer,config.img_size)\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2023-06-18T23:04:08.082883Z","iopub.execute_input":"2023-06-18T23:04:08.083364Z","iopub.status.idle":"2023-06-18T23:04:18.449097Z","shell.execute_reply.started":"2023-06-18T23:04:08.083321Z","shell.execute_reply":"2023-06-18T23:04:18.447979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.keras.utils.plot_model(model, show_shapes=True, dpi=64)","metadata":{"execution":{"iopub.status.busy":"2023-06-18T23:05:42.137882Z","iopub.execute_input":"2023-06-18T23:05:42.138891Z","iopub.status.idle":"2023-06-18T23:05:42.50204Z","shell.execute_reply.started":"2023-06-18T23:05:42.138847Z","shell.execute_reply":"2023-06-18T23:05:42.500803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:teal\">11. Predictions <a class=\"anchor\"  id=\"prediction\"></a></span>","metadata":{}},{"cell_type":"code","source":"test_dataset = build_dataset(test_df, input_features, batch_size=config.batch_size,\n                             label=False, cache=False)\n\npredictions = []\n\nfor _mdl in config.models:\n    print(f'Predicting with model {_mdl}...')\n    model.load_weights(f'{config.models_path}/models/'+_mdl)\n    pred = model.predict(test_dataset)\n    predictions.append(pred)\n    \npredictions = np.mean(predictions, axis=0)\npredictions","metadata":{"execution":{"iopub.status.busy":"2023-06-18T23:05:51.658396Z","iopub.execute_input":"2023-06-18T23:05:51.65949Z","iopub.status.idle":"2023-06-18T23:06:20.867292Z","shell.execute_reply.started":"2023-06-18T23:05:51.659443Z","shell.execute_reply":"2023-06-18T23:06:20.866264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:teal\">12. Prepare submission data <a class=\"anchor\"  id=\"submission\"></a></span>","metadata":{}},{"cell_type":"code","source":"prediction_reshaped=predictions.reshape(-1)\nprint(\"Predictions after reshaped :\",prediction_reshaped)\n\nprediction_ids=test_df['prediction_id']\nprint(\"Prediction Id's : \",prediction_ids)","metadata":{"execution":{"iopub.status.busy":"2023-06-18T23:06:45.008474Z","iopub.execute_input":"2023-06-18T23:06:45.008867Z","iopub.status.idle":"2023-06-18T23:06:45.018081Z","shell.execute_reply.started":"2023-06-18T23:06:45.008832Z","shell.execute_reply":"2023-06-18T23:06:45.016907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predicted_df = pd.DataFrame({'prediction_id':prediction_ids, \n                        'cancer':prediction_reshaped})\npredicted_df","metadata":{"execution":{"iopub.status.busy":"2023-06-18T23:06:50.893164Z","iopub.execute_input":"2023-06-18T23:06:50.89423Z","iopub.status.idle":"2023-06-18T23:06:50.913668Z","shell.execute_reply.started":"2023-06-18T23:06:50.894191Z","shell.execute_reply":"2023-06-18T23:06:50.912214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df = pd.read_csv(config.sample_sub_csv)\ndel submission_df['cancer']\n\nsubmission_df = submission_df.merge(predicted_df, on='prediction_id', how='left')\nsubmission_df = submission_df.groupby('prediction_id')['cancer'].max().reset_index()\n\nsubmission_df.to_csv(config.sub_csv, index=False)","metadata":{"execution":{"iopub.status.busy":"2023-06-18T23:06:54.326315Z","iopub.execute_input":"2023-06-18T23:06:54.327494Z","iopub.status.idle":"2023-06-18T23:06:54.360795Z","shell.execute_reply.started":"2023-06-18T23:06:54.327443Z","shell.execute_reply":"2023-06-18T23:06:54.359483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-06-18T23:06:57.022856Z","iopub.execute_input":"2023-06-18T23:06:57.02361Z","iopub.status.idle":"2023-06-18T23:06:57.039025Z","shell.execute_reply.started":"2023-06-18T23:06:57.023545Z","shell.execute_reply":"2023-06-18T23:06:57.037631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-18T23:07:02.097952Z","iopub.execute_input":"2023-06-18T23:07:02.098397Z","iopub.status.idle":"2023-06-18T23:07:02.111775Z","shell.execute_reply.started":"2023-06-18T23:07:02.098358Z","shell.execute_reply":"2023-06-18T23:07:02.110424Z"},"trusted":true},"execution_count":null,"outputs":[]}]}