{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<center>\n    <img align=\"center\" src=\"https://www.linkpicture.com/q/unical-logo-640x640_1.png\"> \n    <img align=\"center\" src=\"https://www.linkpicture.com/q/logo-rsna.png\"> \n<center>","metadata":{}},{"cell_type":"markdown","source":"# RSNA Screening Mammography Breast Cancer Detection\n## Find breast cancers in screening mammograms\n\n### Context\n* The RSNA International Conference on Artificial Intelligence in Radiology (RSNA-AI) is a new conference that will be held in conjunction with the 2019 RSNA Annual Meeting in Chicago, Illinois, USA.\n* The goal of the conference is to bring together radiologists, radiology trainees, and radiology researchers to discuss the latest advances in artificial intelligence (AI) and machine learning (ML) in radiology.\n* The RSNA and the American College of Radiology (ACR) provide the RSNA-AI Challenge to promote the development of AI and ML algorithms for the detection of breast cancer in mammography.\n* The goal of the challenge is to develop an algorithm that can automatically detect breast cancer in screening mammograms.\n* The work of improving the automation of detection in screening mammography may enable radiologists to be more accurate and efficient, improving the quality and safety of patient care. It could also help reduce costs and unnecessary medical procedures.\n  \nLink to the competition: [Here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/overview) <br>\nLink to the dataset: [Here](https://www.kaggle.com/c/rsna-breast-cancer-detection/data)","metadata":{}},{"cell_type":"markdown","source":"### Notebooks\n* [RSNA BCD | DICOM ➜ ROI-PNG](https://www.kaggle.com/code/matteoperfidio/rsna-bcd-dicom-roi-png): This notebook converts the DICOM images to PNG images with the ROI (Region of Interest).\n* [RSNA BCD | EfficientNetB3 [Train]](https://www.kaggle.com/code/matteoperfidio/rsna-bcd-efficientnetb3-train): This notebook uses the EfficientNetB3 model to train the provided dataset.\n* [RSNA BCD | EfficientNetB3 [Test]](https://www.kaggle.com/code/matteoperfidio/rsna-bcd-efficientnetb3-test): This notebook uses the EfficientNetB3 model to test the provided dataset and submit the results to the competition.\n  \n**Note**: *The test notebook is required because the model is trained on the TPU but the submission must be done via the GPU.*","metadata":{}},{"cell_type":"markdown","source":"### Experiments\nIn the following link you can find the results of the experiments carried out during the training phase of the model:\n* [Comet ML | Training Model 0](https://www.comet.com/nottyche/rsna-bcd-train-model-0)\n* [Comet ML | Training Model 1](https://www.comet.com/nottyche/rsna-bcd-train-model-1)","metadata":{}},{"cell_type":"markdown","source":"### Overview\n\n#### Inferencing\n* In this notebook, we are going to use the ensamble of models that we trained in the [RSNA BCD | EfficientNetB3 [Train]](https://www.kaggle.com/code/matteoperfidio/rsna-bcd-efficientnetb3-train) notebook to test the provided dataset and submit the results to the competition.\n* Since we have 2 models, we are will use the average of the predictions of the two models to get the final prediction.\n  \n#### Medical Analysis\n* Due to the fact that the test set is not labeled, we consulted the opinion of a doctor to understand if in the test set there are cases of breast cancer.","metadata":{}},{"cell_type":"markdown","source":"### Outline\n* [1. Install libraries](#1)\n* [2. Import Libraries](#2)\n* [3. Configuration](#3)\n* [4. Device Configuration](#4)\n* [5. DICOM to PNG](#5)\n    * [5.1 Utility functions](#6)\n    * [5.2 Create ROIs](#7)\n* [6. Add Missing Features](#8)\n* [7. Fix Data Types](#9)\n* [8. Numerical Categorization](#10)\n* [9. Dummy Variables](#11)\n* [10. Define Input Features](#12)\n* [11. Normalization](#13)\n* [12. Build Dataset](#14)\n* [13. Custom Metrics](#15)\n* [14. Build Model](#16)\n* [15. Predictions](#17)\n* [16. Submission](#18)\n* [17. Medical Analysis](#19)","metadata":{}},{"cell_type":"markdown","source":"### 1. Install libraries <a id=\"1\"></a>","metadata":{}},{"cell_type":"code","source":"from IPython.display import clear_output","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:19:23.92097Z","iopub.execute_input":"2023-01-19T23:19:23.921359Z","iopub.status.idle":"2023-01-19T23:19:23.926482Z","shell.execute_reply.started":"2023-01-19T23:19:23.921327Z","shell.execute_reply":"2023-01-19T23:19:23.925375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install -q /kaggle/input/rsna-bcd-whl-ds/dicomsdl-0.109.1-cp37-cp37m-manylinux_2_12_x86_64.manylinux2010_x86_64.whl\nclear_output()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:19:23.935581Z","iopub.execute_input":"2023-01-19T23:19:23.935938Z","iopub.status.idle":"2023-01-19T23:19:54.662487Z","shell.execute_reply.started":"2023-01-19T23:19:23.935902Z","shell.execute_reply":"2023-01-19T23:19:54.661158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2. Import Libraries <a id=\"2\"></a>","metadata":{}},{"cell_type":"code","source":"import os, re, math, random, shutil, warnings, cv2, dicomsdl\nimport numpy as np\nimport pandas as pd\nimport tensorflow as tf\n\nfrom tqdm import tqdm\nfrom joblib import Parallel, delayed\nfrom tensorflow import keras\nfrom tensorflow.python.client import device_lib\nfrom kaggle_datasets import KaggleDatasets\nfrom kaggle_secrets import UserSecretsClient\n\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3'\ntf.get_logger().setLevel('ERROR')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-01-19T23:19:54.665125Z","iopub.execute_input":"2023-01-19T23:19:54.665523Z","iopub.status.idle":"2023-01-19T23:19:59.307157Z","shell.execute_reply.started":"2023-01-19T23:19:54.665479Z","shell.execute_reply":"2023-01-19T23:19:59.306219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3. Configuration <a id=\"3\"></a>","metadata":{}},{"cell_type":"code","source":"img_size = [(1024,512), (512,256)]\nresize_dim = [1024, 512]\nimg_ext = ['png', 'jpg', 'jpeg']\n\nclass Config:\n    \n    def __init__(self):\n        \n        self.debug = False\n        \n        self.device = 'GPU'\n        self.num_devices = 1\n        self.model_name = 'EfficientNetB3'\n        self.seed = 55\n        \n        self.models_path = '/kaggle/input/rsnabcd-train-results-dataset/'\n        self.weights = \"/kaggle/input/efficientnetb3-notop/efficientnetb3_notop.h5\"\n\n        self.data_path = '/kaggle/input/rsna-breast-cancer-detection/'\n        \n        self.test_output_path = '/tmp/dataset/rsna-bcd/test_images/'\n        self.test_path = self.data_path + 'test_images/'\n        self.test_csv = self.data_path + 'test.csv'\n        \n        self.train_csv = self.data_path + 'train.csv'\n        self.train_df = pd.read_csv(self.train_csv)\n        self.train_df_mod = pd.read_csv(self.models_path + 'train_df.csv')\n    \n        self.sub_csv = '/kaggle/working/submission.csv'  \n        self.sample_sub_csv = self.data_path + 'sample_submission.csv'\n        self.num_models = 2\n        \n        self.threshold = 0.2\n        \n        self.parameters = {\n            'batch_size': 32,\n            'epochs': 12,\n            'dropout': 0.3,\n            'optimizer': 'adam',\n            'loss': 'binary_crossentropy'\n        }\n        \n        self.img_size = img_size[1]\n        self.resize_dim = resize_dim[1]\n        self.img_ext = img_ext[0]\n\nconfig = Config()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:19:59.30851Z","iopub.execute_input":"2023-01-19T23:19:59.309164Z","iopub.status.idle":"2023-01-19T23:19:59.667667Z","shell.execute_reply.started":"2023-01-19T23:19:59.309126Z","shell.execute_reply":"2023-01-19T23:19:59.666474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4. Device Configuration <a id=\"4\"></a>","metadata":{}},{"cell_type":"code","source":"num_devices = len(tf.config.list_physical_devices('GPU'))\n\nif num_devices > 1:\n    config.num_devices = num_devices\n    strategy = tf.distribute.MirroredStrategy()\n    print(f'Running on {num_devices} GPU devices')\nelif num_devices == 1:\n    strategy = tf.distribute.get_strategy()\n    print(f'Running on {num_devices} GPU device')\nelse:\n    strategy = tf.distribute.get_strategy()\n    config.device = 'CPU'\n    print(f'Running on CPU')\n\ntf.config.optimizer.set_jit(True)\ntf.keras.mixed_precision.set_global_policy(\"mixed_float16\")\nconfig.parameters['batch_size'] = config.parameters['batch_size'] * config.num_devices","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:19:59.669973Z","iopub.execute_input":"2023-01-19T23:19:59.670274Z","iopub.status.idle":"2023-01-19T23:20:00.607919Z","shell.execute_reply.started":"2023-01-19T23:19:59.670245Z","shell.execute_reply":"2023-01-19T23:20:00.606971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = pd.read_csv(config.test_csv)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:00.609228Z","iopub.execute_input":"2023-01-19T23:20:00.609586Z","iopub.status.idle":"2023-01-19T23:20:00.621975Z","shell.execute_reply.started":"2023-01-19T23:20:00.609543Z","shell.execute_reply":"2023-01-19T23:20:00.621083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:00.623672Z","iopub.execute_input":"2023-01-19T23:20:00.624031Z","iopub.status.idle":"2023-01-19T23:20:00.647918Z","shell.execute_reply.started":"2023-01-19T23:20:00.623996Z","shell.execute_reply":"2023-01-19T23:20:00.646915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:00.649335Z","iopub.execute_input":"2023-01-19T23:20:00.649691Z","iopub.status.idle":"2023-01-19T23:20:00.668458Z","shell.execute_reply.started":"2023-01-19T23:20:00.649657Z","shell.execute_reply":"2023-01-19T23:20:00.667589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df['dicom_path'] = config.test_path + test_df['patient_id'].astype(str) + '/' + test_df['image_id'].astype(str) + '.dcm'\ntest_df['image_path'] = config.test_output_path + test_df['patient_id'].astype(str) + '/' + test_df['image_id'].astype(str) + '.png'","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:00.669771Z","iopub.execute_input":"2023-01-19T23:20:00.670228Z","iopub.status.idle":"2023-01-19T23:20:00.678078Z","shell.execute_reply.started":"2023-01-19T23:20:00.67018Z","shell.execute_reply":"2023-01-19T23:20:00.677121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:00.679601Z","iopub.execute_input":"2023-01-19T23:20:00.680433Z","iopub.status.idle":"2023-01-19T23:20:00.696158Z","shell.execute_reply.started":"2023-01-19T23:20:00.680396Z","shell.execute_reply":"2023-01-19T23:20:00.695004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.io.gfile.exists(test_df.dicom_path.iloc[0])","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:00.701206Z","iopub.execute_input":"2023-01-19T23:20:00.701495Z","iopub.status.idle":"2023-01-19T23:20:00.710797Z","shell.execute_reply.started":"2023-01-19T23:20:00.701456Z","shell.execute_reply":"2023-01-19T23:20:00.709556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 5. DICOM to PNG <a id=\"5\"></a>","metadata":{}},{"cell_type":"markdown","source":"#### 5.1 Utility functions<a id=\"6\"></a>","metadata":{}},{"cell_type":"code","source":"def dicom_to_png(dicom_path):\n    dicom = dicomsdl.open(dicom_path)\n    image = dicom.pixelData(storedvalue=False)\n    image = image - np.min(image)\n    image = image / np.max(image)\n\n    if dicom.PhotometricInterpretation == 'MONOCHROME1':\n        image = 1.0 - image\n    \n    image = cv2.resize(image, (config.resize_dim, config.resize_dim), interpolation=cv2.INTER_LINEAR)\n    image = (image * 255).astype(np.uint8)\n    return image","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:00.712282Z","iopub.execute_input":"2023-01-19T23:20:00.712637Z","iopub.status.idle":"2023-01-19T23:20:00.718459Z","shell.execute_reply.started":"2023-01-19T23:20:00.712586Z","shell.execute_reply":"2023-01-19T23:20:00.71749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def png_to_roi(image, image_path):\n    bin_image = cv2.threshold(image, 20, 255, cv2.THRESH_BINARY)[1]\n    contours, _ = cv2.findContours(bin_image, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_NONE)\n    contour = max(contours, key=cv2.contourArea)\n    ys = contour.squeeze()[:, 0]\n    xs = contour.squeeze()[:, 1]\n    roi = image[np.min(xs):np.max(xs), np.min(ys):np.max(ys)]\n    return cv2.resize(roi, config.img_size[::-1], interpolation=cv2.INTER_LINEAR)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:00.719962Z","iopub.execute_input":"2023-01-19T23:20:00.720678Z","iopub.status.idle":"2023-01-19T23:20:00.728744Z","shell.execute_reply.started":"2023-01-19T23:20:00.720642Z","shell.execute_reply":"2023-01-19T23:20:00.727659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 5.2 Create ROIs <a id=\"7\"></a>","metadata":{}},{"cell_type":"code","source":"def process(dicom_path, image_path):\n    image = dicom_to_png(dicom_path)\n    os.makedirs(os.path.dirname(image_path), exist_ok=True)\n    image = png_to_roi(image, image_path)\n    cv2.imwrite(image_path, image)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:00.730202Z","iopub.execute_input":"2023-01-19T23:20:00.731006Z","iopub.status.idle":"2023-01-19T23:20:00.737267Z","shell.execute_reply.started":"2023-01-19T23:20:00.730977Z","shell.execute_reply":"2023-01-19T23:20:00.736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Parallel(n_jobs=4, backend='threading')(delayed(process)(dicom_path, image_path) \n                   for dicom_path, image_path in tqdm(zip(test_df['dicom_path'], \n                                                          test_df['image_path'])))\nclear_output()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:00.739236Z","iopub.execute_input":"2023-01-19T23:20:00.739655Z","iopub.status.idle":"2023-01-19T23:20:03.433643Z","shell.execute_reply.started":"2023-01-19T23:20:00.739585Z","shell.execute_reply":"2023-01-19T23:20:03.432679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 6. Add Missing Features <a id=\"8\"></a>","metadata":{}},{"cell_type":"code","source":"config.train_df = config.train_df.drop(['BIRADS'], axis=1)\nconfig.train_df = config.train_df.drop(['cancer'], axis=1)\n\nconfig.train_df['age'] = config.train_df['age'].fillna(config.train_df['age'].mean())\nconfig.train_df['density'] = config.train_df['density'].fillna('E')\n\nconfig.train_df['laterality'] = config.train_df['laterality'].astype('category')\nconfig.train_df['difficult_negative_case'] = config.train_df['difficult_negative_case'].astype('category')\nconfig.train_df['density'] = config.train_df['density'].astype('category')\nconfig.train_df['view'] = config.train_df['view'].astype('category')","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:03.435232Z","iopub.execute_input":"2023-01-19T23:20:03.435593Z","iopub.status.idle":"2023-01-19T23:20:03.488836Z","shell.execute_reply.started":"2023-01-19T23:20:03.435559Z","shell.execute_reply":"2023-01-19T23:20:03.48784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for column in config.train_df.columns:\n    if column not in test_df:\n        if config.train_df[column].dtype == 'category':\n            test_df[column] = config.train_df[column].sample(1).values[0]\n        else:\n            test_df[column] = config.train_df[column].mean()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:03.493787Z","iopub.execute_input":"2023-01-19T23:20:03.496226Z","iopub.status.idle":"2023-01-19T23:20:03.511656Z","shell.execute_reply.started":"2023-01-19T23:20:03.496183Z","shell.execute_reply":"2023-01-19T23:20:03.510885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 7. Fix Data Types <a id=\"9\"></a>","metadata":{}},{"cell_type":"code","source":"test_df['laterality'] = test_df['laterality'].astype('category')\ntest_df['difficult_negative_case'] = test_df['difficult_negative_case'].astype('category')\ntest_df['density'] = test_df['density'].astype('category')\ntest_df['view'] = test_df['view'].astype('category')\ntest_df['age'] = test_df['age'].astype('int64')\ntest_df['image_path'] = test_df['image_path'].astype('string')","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:03.515469Z","iopub.execute_input":"2023-01-19T23:20:03.51754Z","iopub.status.idle":"2023-01-19T23:20:03.530462Z","shell.execute_reply.started":"2023-01-19T23:20:03.517506Z","shell.execute_reply":"2023-01-19T23:20:03.529657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 8. Numerical Categorization <a id=\"10\"></a>","metadata":{}},{"cell_type":"code","source":"test_df['age_bin'] = pd.cut(test_df['age'], bins=[0, 30, 50, 70, 90, 110], labels=False)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:03.534377Z","iopub.execute_input":"2023-01-19T23:20:03.536547Z","iopub.status.idle":"2023-01-19T23:20:03.550987Z","shell.execute_reply.started":"2023-01-19T23:20:03.536513Z","shell.execute_reply":"2023-01-19T23:20:03.550024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 9. Dummy Variables <a id=\"11\"></a>","metadata":{}},{"cell_type":"code","source":"cat_cols = ['laterality', 'view', 'density', 'difficult_negative_case']\ncat = []\n\nfor col in cat_cols:\n    if len(test_df[col].unique()) != 1:\n        cat.append(col)\n    else:\n        test_df.drop(col, axis=1, inplace=True)\n\ntest_df = pd.get_dummies(test_df, columns=cat)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:03.555264Z","iopub.execute_input":"2023-01-19T23:20:03.557668Z","iopub.status.idle":"2023-01-19T23:20:03.574985Z","shell.execute_reply.started":"2023-01-19T23:20:03.557614Z","shell.execute_reply":"2023-01-19T23:20:03.573965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 10. Define Input Features <a id=\"12\"></a>","metadata":{}},{"cell_type":"code","source":"for column in config.train_df_mod.columns:\n    if column not in test_df.columns:\n        test_df[column] = config.train_df_mod[column].unique()[0]","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:03.580699Z","iopub.execute_input":"2023-01-19T23:20:03.583405Z","iopub.status.idle":"2023-01-19T23:20:03.608539Z","shell.execute_reply.started":"2023-01-19T23:20:03.583369Z","shell.execute_reply":"2023-01-19T23:20:03.607593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"input_features = test_df.columns[~test_df.columns.isin(['patient_id', 'image_id', 'site_id', 'machine_id',\n                                                        'width', 'height', 'cancer', 'age', 'BIRADS', 'dicom_path',\n                                                        'stratify', 'image_path', 'fold', 'prediction_id'])]","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:03.612809Z","iopub.execute_input":"2023-01-19T23:20:03.615057Z","iopub.status.idle":"2023-01-19T23:20:03.622785Z","shell.execute_reply.started":"2023-01-19T23:20:03.61502Z","shell.execute_reply":"2023-01-19T23:20:03.621709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 11. Normalization <a id=\"13\"></a>","metadata":{}},{"cell_type":"code","source":"test_df[input_features] = (test_df[input_features] - config.train_df_mod[input_features].mean()) / config.train_df_mod[input_features].std()\ntest_df[input_features] = test_df[input_features].astype('float32')","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:03.643477Z","iopub.execute_input":"2023-01-19T23:20:03.646552Z","iopub.status.idle":"2023-01-19T23:20:03.692067Z","shell.execute_reply.started":"2023-01-19T23:20:03.646504Z","shell.execute_reply":"2023-01-19T23:20:03.690907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 12. Build Dataset <a id=\"14\"></a>","metadata":{}},{"cell_type":"code","source":"def decode_image(label=True, img_size=config.img_size, ext=config.img_ext):\n\n    def _decode_image(input):\n        image = tf.io.read_file(input['input_image'])\n        if ext == 'png':\n            image = tf.image.decode_png(image, channels=3)\n        elif ext in ['jpg', 'jpeg']:\n            image = tf.image.decode_jpeg(image, channels=3)\n        else:\n            raise ValueError(\"Image extension not supported\")\n        image = tf.image.resize(image, img_size)\n        image = tf.reshape(image, [*img_size, 3])\n        image = tf.cast(image, tf.float32) / 255.0\n        input['input_image'] = image\n        return input\n    \n    def _decode_labeled_image(input, label):\n        input = _decode_image(input)\n        return input, label\n\n    return _decode_labeled_image if label else _decode_image\n\n\ndef build_dataset(df, input_features, image_size=config.img_size, batch_size=config.parameters['batch_size'], \n                  label=True, cache=False, ext=config.img_ext):\n    \n    decode = decode_image(label, img_size=image_size, ext=ext)\n\n    if label:\n        dataset = tf.data.Dataset.from_tensor_slices(({'input_image': df['image_path'].values, \n                                                       'input_features' : df[input_features].values}, df['cancer'].values))\n    else:\n        dataset = tf.data.Dataset.from_tensor_slices({'input_image': df['image_path'].values, \n                                                      'input_features' : df[input_features].values})\n\n    dataset = dataset.map(decode, num_parallel_calls=tf.data.AUTOTUNE)\n\n    if cache:\n        dataset = dataset.cache()\n        \n    dataset = dataset.batch(batch_size)\n    dataset = dataset.prefetch(tf.data.AUTOTUNE)\n    return dataset","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:03.697065Z","iopub.execute_input":"2023-01-19T23:20:03.699526Z","iopub.status.idle":"2023-01-19T23:20:03.716114Z","shell.execute_reply.started":"2023-01-19T23:20:03.699473Z","shell.execute_reply":"2023-01-19T23:20:03.715012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 13. Custom Metrics <a id=\"15\"></a>","metadata":{}},{"cell_type":"code","source":"def p_f1(y_true, y_pred):\n    y_true = tf.cast(y_true, tf.float32)\n    y_pred = tf.cast(y_pred, tf.float32)\n    \n    tp = tf.reduce_sum(y_true * y_pred)\n    tn = tf.reduce_sum((1 - y_true) * (1 - y_pred))\n    fp = tf.reduce_sum((1 - y_true) * y_pred)\n    fn = tf.reduce_sum(y_true * (1 - y_pred))\n    \n    p = tp / (tp + fp + tf.keras.backend.epsilon())\n    r = tp / (tp + fn + tf.keras.backend.epsilon())\n    \n    f1 = 2 * p * r / (p + r + tf.keras.backend.epsilon())\n    f1 = tf.where(tf.math.is_nan(f1), tf.zeros_like(f1), f1)\n\n    return tf.reduce_mean(f1)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:03.721673Z","iopub.execute_input":"2023-01-19T23:20:03.724591Z","iopub.status.idle":"2023-01-19T23:20:03.735111Z","shell.execute_reply.started":"2023-01-19T23:20:03.724552Z","shell.execute_reply":"2023-01-19T23:20:03.734132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 14. Build Model <a id=\"16\"></a>","metadata":{}},{"cell_type":"code","source":"def build_model(input_features, loss=config.parameters['loss'], dropout=config.parameters['dropout'], \n                optimizer=config.parameters['optimizer'], img_size=config.img_size):\n    with strategy.scope():\n        inputs = tf.keras.layers.Input(shape=img_size+(3,), name='input_image')\n        features = tf.keras.layers.Input(shape=[len(input_features)], name='input_features')\n        x = tf.keras.applications.EfficientNetB3(input_shape=img_size+(3,), include_top=False, \n                                                 drop_connect_rate=0.4, weights=config.weights)(inputs)\n        x = tf.keras.layers.GlobalAveragePooling2D()(x)\n        x = tf.keras.layers.Dropout(dropout)(x)\n        x = tf.keras.layers.Dense(32, activation=\"relu\")(x)\n        x = tf.keras.layers.BatchNormalization()(x)\n        x = tf.keras.layers.Dropout(dropout)(x)\n        x = tf.keras.layers.Concatenate()([x, features])\n        x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n        model = tf.keras.Model(inputs=[inputs, features], outputs=x)\n        model.compile(optimizer=optimizer, loss=loss, metrics=['accuracy', p_f1, tf.keras.metrics.AUC(name='auc')])\n        return model\n\nmodel = build_model(input_features)\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:03.740313Z","iopub.execute_input":"2023-01-19T23:20:03.742118Z","iopub.status.idle":"2023-01-19T23:20:11.121091Z","shell.execute_reply.started":"2023-01-19T23:20:03.742069Z","shell.execute_reply":"2023-01-19T23:20:11.120049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 15. Predictions <a id=\"17\"></a>","metadata":{}},{"cell_type":"code","source":"test_dataset = build_dataset(test_df, input_features, batch_size=config.parameters['batch_size'], label=False, cache=False)\npredictions = []\n\nfor fold in range(config.num_models):\n    print(f'Predicting with model {fold}...')\n    model.load_weights(f'{config.models_path}/models/model_{fold}.h5')\n    pred = model.predict(test_dataset, verbose=1)\n    predictions.append(pred)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:11.122585Z","iopub.execute_input":"2023-01-19T23:20:11.123211Z","iopub.status.idle":"2023-01-19T23:20:36.24503Z","shell.execute_reply.started":"2023-01-19T23:20:11.123172Z","shell.execute_reply":"2023-01-19T23:20:36.24265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = np.mean(predictions, axis=0)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:36.252406Z","iopub.execute_input":"2023-01-19T23:20:36.25281Z","iopub.status.idle":"2023-01-19T23:20:36.261228Z","shell.execute_reply.started":"2023-01-19T23:20:36.252772Z","shell.execute_reply":"2023-01-19T23:20:36.260192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 16. Submission <a id=\"18\"></a>","metadata":{}},{"cell_type":"code","source":"pred_df = pd.DataFrame({'prediction_id': test_df['prediction_id'], 'cancer': predictions.reshape(-1)})\nsub_df = pd.read_csv(config.sample_sub_csv)\ndel sub_df['cancer']\n\nsub_df = sub_df.merge(pred_df, on='prediction_id', how='left')\nsub_df = sub_df.groupby('prediction_id')['cancer'].max().reset_index()\nsub_df['cancer'] = (sub_df.cancer > config.threshold).astype('float32')\n\nsub_df.to_csv(config.sub_csv, index=False)","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:36.262642Z","iopub.execute_input":"2023-01-19T23:20:36.265495Z","iopub.status.idle":"2023-01-19T23:20:36.308569Z","shell.execute_reply.started":"2023-01-19T23:20:36.265457Z","shell.execute_reply":"2023-01-19T23:20:36.307554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:36.312725Z","iopub.execute_input":"2023-01-19T23:20:36.315405Z","iopub.status.idle":"2023-01-19T23:20:36.328802Z","shell.execute_reply.started":"2023-01-19T23:20:36.31537Z","shell.execute_reply":"2023-01-19T23:20:36.327716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-19T23:20:36.333009Z","iopub.execute_input":"2023-01-19T23:20:36.335598Z","iopub.status.idle":"2023-01-19T23:20:36.348642Z","shell.execute_reply.started":"2023-01-19T23:20:36.335562Z","shell.execute_reply":"2023-01-19T23:20:36.347511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 17. Medical Analysis <a id=\"19\"></a>","metadata":{}},{"cell_type":"markdown","source":"The Doctor's opinion is that there are a case of breast cancer in the test set. In particular, the image with the ID `68070693` has a case of breast cancer. This image belongs to the right hand side of the screening mammogram. The subregions of the image in which the breast cancer are highlighted in the image below.","metadata":{}},{"cell_type":"markdown","source":"<center>\n    <img align=\"center\" src=\"https://www.linkpicture.com/q/med-analysis.png\" width=\"960\" height=\"500\">\n<center>","metadata":{}},{"cell_type":"markdown","source":"Unfortunately, the model is not able to detect the breast cancer in the image with the ID `68070693`. This is due to the fact that our model is not only based on the image but also on the informations releated to the patient. This behaviour is seems to be correct because also for an human being, the informations related to the patient are crucial to determine if there is a case of breast cancer or not.","metadata":{}}]}