{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# [RSNA Screening Mammography Breast Cancer Detection](https://www.kaggle.com/c/petfinder-pawpularity-score)\n> Find breast cancers in screening mammograms\n\n![](https://storage.googleapis.com/kaggle-competitions/kaggle/39272/logos/header.png?t=2022-11-28-17-29-35)","metadata":{}},{"cell_type":"markdown","source":"# Idea:\n* In this notebook will classify `breast cancer` from Mammography images. Intially, we'll use CNN models then we'll move towards Transformer models.\n* There is a huge **class-imbalance** between cancer and non-cancer class which makes it difficult for the model to learn properly. Moreover, cancer size could be from very small to very large, so there is also **pixel-imbalance** which makes the task even more difficult. There are couple of ways to handle imbalance ex: `WeightedLoss`, `FocalLoss`, `Upsample`. In our notebook we'll use both `FocalLoss` and `Upsample`.\n* We'll use whole image (**NoROI**) instead of ROI: Regoin Of Interest as its faster. To avoid distortion we'll also use **Rectangule** image for training instead of Squares\n* As dataset is quite large we'll utilize `TPU-1VM` aka `Local TPU` to speed up the training.\n* We'll use `cancer` column as target.\n* We'll use **CNN** model: EfficientNet for our training.\n* We won't use **Tabular** feature in this notebook.\n* **Wandb** is integrated hence we can use this notebook to track which experiemnt is peforming better.\n* We'll maximize the `pF1` score using **Threshold**\n","metadata":{}},{"cell_type":"markdown","source":"# Notebooks\n* Only Image:\n    * ROI:\n        * train: [RSNA-BCD: EfficientNet [TF][TPU-1VM][Train]](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train)\n        * infer: [RSNA-BCD: EfficientNet [TF][TPU-1VM][Infer]](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-infer)\n    * NoROI + KerasCV: \n        * train: [RSNA-BCD: NoROI KerasCV [TF][Train]](https://www.kaggle.com/awsaf49/rsna-bcd-noroi-kerascv-tf-train/)\n        * infer: [RSNA-BCD: NoROI KerasCV [TF][Infer]](https://www.kaggle.com/awsaf49/rsna-bcd-noroi-kerascv-tf-infer/)\n* Dataset:\n    * ROI:\n        * [RSNA-BCD: ROI 1024x PNG Dataset](https://www.kaggle.com/datasets/awsaf49/rsna-bcd-roi-1024x-png-dataset)\n    * NoROI:\n        * [RSNA-BCD: 512 PNG v2 PNG Dataset](https://www.kaggle.com/datasets/awsaf49/rsnabcd-512-png-v2-dataset)","metadata":{}},{"cell_type":"markdown","source":"# Update\n* `v6` - `14/12/2022`:\n    * Balanced Sampling with `tf.data.Dataset`","metadata":{}},{"cell_type":"markdown","source":"# Overview\n\n## NoROI + Square Image\n* This notebook will the whole image (**NoROI**) instead of ROI (Region Of Interest).\n* This notebook will use **square** image (width!=height) training to avoid resizing hence speed up.\n\n## Upsample `Cancer`\n* Upsample the cancer data `10x` to reduce class_imbalance effect on loss\n\n## `TPU-1VM` aka `Local TPU`\n* It doesn't require `GCS_PATH` like normaly `TPU`. So, for this, we don't have to make our dataset **public** anymore. \n* As datasets will be read from local path, it should speed up the first epoch (for normal `tpu` first epoch is quite slow).\n\n## Augmentations:\n* We'll use KerasCV as it comes with varies augmentations\n\n## WandB Integration:\n* You can track your training using **wandb**\n* It's very easy to compare model's performance using **wandb**.\n\n## `pF1` Maximization:\n* Use simple empirical method to attain best `threshold` for max `pF1`","metadata":{}},{"cell_type":"markdown","source":"# Install Libraries","metadata":{}},{"cell_type":"markdown","source":"## Install for `TPU 1VM v3-8`\nUsually these libraries come pre-installed for other accelerators but fo `tpu-1vm` we need to install them manually. If you want to use rather `remote-TPU` or `GPU` or `CPU`, then comment out the following cell.","metadata":{}},{"cell_type":"code","source":"!pip install -q /lib/wheels/tensorflow-2.9.1-cp38-cp38-linux_x86_64.whl\n!pip install -q tensorflow-addons==0.18.0\n!pip install -q tensorflow-probability==0.17.0\n!pip install -q keras-cv==0.3.4\n!pip install -q opencv-python-headless\n!pip install -q seaborn","metadata":{"_kg_hide-input":false,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-12-13T20:17:58.183799Z","iopub.execute_input":"2022-12-13T20:17:58.184543Z","iopub.status.idle":"2022-12-13T20:19:20.419924Z","shell.execute_reply.started":"2022-12-13T20:17:58.184453Z","shell.execute_reply":"2022-12-13T20:19:20.418976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Install Custom Libraries","metadata":{}},{"cell_type":"code","source":"!pip install -q efficientnet >> /dev/null\n!pip install -qU wandb\n!pip install -qU scikit-learn","metadata":{"_kg_hide-output":true,"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-12-13T20:19:20.421698Z","iopub.execute_input":"2022-12-13T20:19:20.422311Z","iopub.status.idle":"2022-12-13T20:19:43.752333Z","shell.execute_reply.started":"2022-12-13T20:19:20.422277Z","shell.execute_reply":"2022-12-13T20:19:43.751378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import Libraries","metadata":{}},{"cell_type":"code","source":"import os\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3'  # to avoid too many logging messages\nimport pandas as pd, numpy as np, random, shutil\nimport tensorflow as tf, re, math\ntf.get_logger().setLevel('ERROR')\nimport tensorflow.keras.backend as K\nimport sklearn\nimport matplotlib.pyplot as plt\nimport tensorflow_addons as tfa\nimport tensorflow_probability as tfp\nimport wandb\nimport yaml\n\nfrom IPython import display as ipd\nfrom glob import glob\nfrom tqdm import tqdm\nfrom sklearn.model_selection import KFold, StratifiedKFold, GroupKFold, StratifiedGroupKFold\nfrom sklearn.metrics import roc_auc_score","metadata":{"execution":{"iopub.status.busy":"2022-12-13T20:19:43.753593Z","iopub.execute_input":"2022-12-13T20:19:43.753875Z","iopub.status.idle":"2022-12-13T20:20:01.341788Z","shell.execute_reply.started":"2022-12-13T20:19:43.753848Z","shell.execute_reply":"2022-12-13T20:20:01.340686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Version Check","metadata":{}},{"cell_type":"code","source":"print('np:', np.__version__)\nprint('pd:', pd.__version__)\nprint('sklearn:', sklearn.__version__)\nprint('tf:',tf.__version__)\nprint('tfp:', tfp.__version__)\nprint('tfa:', tfa.__version__)\nprint('w&b:', wandb.__version__)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T20:20:01.344082Z","iopub.execute_input":"2022-12-13T20:20:01.344804Z","iopub.status.idle":"2022-12-13T20:20:01.351961Z","shell.execute_reply.started":"2022-12-13T20:20:01.344773Z","shell.execute_reply":"2022-12-13T20:20:01.351237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Wandb","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://camo.githubusercontent.com/dd842f7b0be57140e68b2ab9cb007992acd131c48284eaf6b1aca758bfea358b/68747470733a2f2f692e696d6775722e636f6d2f52557469567a482e706e67\" width=\"400\" alt=\"Weights & Biases\" />\n\nWeights & Biases (W&B) is MLOps platform for tracking our experiemnts. We can use it to Build better models faster with experiment tracking, dataset versioning, and model management\n\nSome of the cool features of **W&B**:\n\n* Track, compare, and visualize ML experiments\n* Get live metrics, terminal logs, and system stats streamed to the centralized dashboard.\n* Explain how your model works, show graphs of how model versions improved, discuss bugs, and demonstrate progress towards milestones.\n\n> **Note:** `kaggle_secrets` has not been added to `TPU-1VM` yet, so run anonymously then go to the dashbaord and simply `claim` the run.","metadata":{}},{"cell_type":"code","source":"import wandb\n\ntry:\n    from kaggle_secrets import UserSecretsClient\n    user_secrets = UserSecretsClient()\n    api_key = user_secrets.get_secret(\"WANDB\")\n\n    wandb.login(key=api_key)\n    anonymous = None\nexcept:\n    anonymous = \"must\"\n    print('To use your W&B account,\\nGo to Add-ons -> Secrets and provide your W&B access token. Use the Label name as WANDB. \\nGet your W&B access token from here: https://wandb.ai/authorize')","metadata":{"execution":{"iopub.status.busy":"2022-12-13T20:20:01.352959Z","iopub.execute_input":"2022-12-13T20:20:01.353238Z","iopub.status.idle":"2022-12-13T20:20:01.366088Z","shell.execute_reply.started":"2022-12-13T20:20:01.353212Z","shell.execute_reply":"2022-12-13T20:20:01.365272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Configuration","metadata":{}},{"cell_type":"code","source":"class CFG:\n    wandb         = True\n    competition   = 'rsna-bcd' \n    _wandb_kernel = 'awsaf49'\n    debug         = False\n    comment       = 'EfficientNetB4-1024x1024-up=10-lr6-kcv-aug=0.90'\n    exp_name      = 'whole-image' # name of the experiment, folds will be grouped using 'exp_name'\n    \n    # use verbose=0 for silent, vebose=1 for interactive,\n    verbose      = 1\n    display_plot = True\n\n    # device\n    device = \"TPU-1VM\" #or \"GPU\"\n\n    model_name = 'EfficientNetB4'\n\n    # seed for data-split, layer init, augs\n    seed = 42\n\n    # number of folds for data-split\n    folds = 5\n    \n    # which folds to train\n    selected_folds = [0, 1,]\n\n    # size of the image\n    img_size = [1024, 1024]\n\n    # batch_size and epochs\n    batch_size = 8\n    epochs = 10\n    \n    # upsample\n    upsample = 1\n    pos_prob = 0.125  # 1/8\n\n    # loss and optimizer\n    loss      = 'BCE'  # BCE, Focal\n    optimizer = 'Adam'\n\n    # augmentation\n    keras_cv  = True\n    augment   = 0.90\n    flip_mode = \"horizontal\"\n    rot_range = 1\n\n    # clip\n    clip = False\n\n    # lr-scheduler\n    scheduler   = 'exp' # cosine\n\n    # test-time augs\n    tta = 1\n    \n    # target column\n    target_col  = ['cancer']","metadata":{"execution":{"iopub.status.busy":"2022-12-13T20:29:05.275143Z","iopub.execute_input":"2022-12-13T20:29:05.276018Z","iopub.status.idle":"2022-12-13T20:29:05.290642Z","shell.execute_reply.started":"2022-12-13T20:29:05.275887Z","shell.execute_reply":"2022-12-13T20:29:05.289925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Reproducibility","metadata":{}},{"cell_type":"code","source":"def seeding(SEED):\n    np.random.seed(SEED)\n    random.seed(SEED)\n    os.environ['PYTHONHASHSEED'] = str(SEED)\n#     os.environ['TF_CUDNN_DETERMINISTIC'] = str(SEED)\n    tf.random.set_seed(SEED)\n    print('seeding done!!!')\nseeding(CFG.seed)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T20:20:01.376835Z","iopub.execute_input":"2022-12-13T20:20:01.377159Z","iopub.status.idle":"2022-12-13T20:20:01.387923Z","shell.execute_reply.started":"2022-12-13T20:20:01.377131Z","shell.execute_reply":"2022-12-13T20:20:01.387116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Device Configs\nThis notebook is compatible for **remote-tpu**, **local-tpu**, **multi-gpu** and **single-gpu**. Simple change to `device=\"TPU\"` for **remote-tpu** and `device=\"TPU-1VM\"` for **local-tpu** and finally, `device=\"GPU\"` for single or multi-gpu.","metadata":{}},{"cell_type":"code","source":"if \"TPU\" in CFG.device:\n    tpu = 'local' if CFG.device=='TPU-1VM' else None\n    print(\"connecting to TPU...\")\n    try:\n        tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=tpu)\n        strategy = tf.distribute.TPUStrategy(tpu)\n    except:\n        CFG.device = \"GPU\"\n        \nif CFG.device == \"GPU\"  or CFG.device==\"CPU\":\n    ngpu = len(tf.config.experimental.list_physical_devices('GPU'))\n    if ngpu>1:\n        print(\"Using multi GPU\")\n        strategy = tf.distribute.MirroredStrategy()\n    elif ngpu==1:\n        print(\"Using single GPU\")\n        strategy = tf.distribute.get_strategy()\n    else:\n        print(\"Using CPU\")\n        strategy = tf.distribute.get_strategy()\n        CFG.device = \"CPU\"\n\nif CFG.device == \"GPU\":\n    print(\"Num GPUs Available: \", ngpu)\n    \n\nAUTO     = tf.data.experimental.AUTOTUNE\nREPLICAS = strategy.num_replicas_in_sync\nprint(f'REPLICAS: {REPLICAS}')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-12-13T20:20:01.38889Z","iopub.execute_input":"2022-12-13T20:20:01.38916Z","iopub.status.idle":"2022-12-13T20:20:14.087479Z","shell.execute_reply.started":"2022-12-13T20:20:01.389137Z","shell.execute_reply":"2022-12-13T20:20:14.086458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# GCS Path for TPU\n* Remote-TPU requires **GCS** path. Luckily Kaggle Provides that for us :)\n* Local-TPU aka TPU-1VM doesn't need GCS path =)","metadata":{}},{"cell_type":"code","source":"BASE_PATH = '/kaggle/input/rsna-bcd-1024x1024-png-v5-dataset'\n\nif CFG.device==\"TPU\":\n    from kaggle_datasets import KaggleDatasets\n    GCS_PATH = KaggleDatasets().get_gcs_path(BASE_PATH.split('/')[-1])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-13T20:20:14.088776Z","iopub.execute_input":"2022-12-13T20:20:14.089205Z","iopub.status.idle":"2022-12-13T20:20:14.093622Z","shell.execute_reply.started":"2022-12-13T20:20:14.089178Z","shell.execute_reply":"2022-12-13T20:20:14.092851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Meta Data","metadata":{}},{"cell_type":"code","source":"# use gcs_path for remote-tpu\nif CFG.device==\"TPU\":\n    BASE_PATH = GCS_PATH\n\n# train\ndf = pd.read_csv(f'{BASE_PATH}/train.csv')\ndf['image_path'] = f'{BASE_PATH}/train_images'\\\n                    + '/' + df.patient_id.astype(str)\\\n                    + '/' + df.image_id.astype(str)\\\n                    + '.png'\nprint('Train:')\ndisplay(df.head(2))\n\n# test\ntest_df = pd.read_csv(f'{BASE_PATH}/test.csv')\ntest_df['image_path'] = f'{BASE_PATH}/test_images'\\\n                    + '/' + test_df.patient_id.astype(str)\\\n                    + '/' + test_df.image_id.astype(str)\\\n                    + '.png'\nprint('\\nTest:')\ndisplay(test_df.head(2))","metadata":{"execution":{"iopub.status.busy":"2022-12-13T20:20:14.096804Z","iopub.execute_input":"2022-12-13T20:20:14.097526Z","iopub.status.idle":"2022-12-13T20:20:14.425313Z","shell.execute_reply.started":"2022-12-13T20:20:14.097497Z","shell.execute_reply":"2022-12-13T20:20:14.42425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Check If Data Exist?","metadata":{}},{"cell_type":"code","source":"tf.io.gfile.exists(df.image_path.iloc[0]), tf.io.gfile.exists(test_df.image_path.iloc[0])","metadata":{"execution":{"iopub.status.busy":"2022-12-13T20:20:14.426612Z","iopub.execute_input":"2022-12-13T20:20:14.426937Z","iopub.status.idle":"2022-12-13T20:20:14.446466Z","shell.execute_reply.started":"2022-12-13T20:20:14.426889Z","shell.execute_reply":"2022-12-13T20:20:14.44559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train-Test Ditribution","metadata":{}},{"cell_type":"code","source":"print('train_files:',df.shape[0])\nprint('test_files:',test_df.shape[0])","metadata":{"execution":{"iopub.status.busy":"2022-12-13T20:20:14.447716Z","iopub.execute_input":"2022-12-13T20:20:14.448051Z","iopub.status.idle":"2022-12-13T20:20:14.452888Z","shell.execute_reply.started":"2022-12-13T20:20:14.448022Z","shell.execute_reply":"2022-12-13T20:20:14.452054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Split\n* Data is splited while stratifying, 'laterality', 'view', 'age','cancer', 'biopsy', 'invasive', 'BIRADS', 'implant', 'density','machine_id', 'difficult_negative_case','cancer'.\n* To avoid leakage, data is also split keeping images from same patient in either train or valid not in both.\n* `StratifiedGroupKFold` does the both job **stratifying** and **group** split.\n\n","metadata":{}},{"cell_type":"code","source":"num_bins = 5\ndf[\"age_bin\"] = pd.cut(df['age'].values.reshape(-1), bins=num_bins, labels=False)\n\nstrat_cols = [\n    'laterality', 'view', 'biopsy','invasive', 'BIRADS', 'age_bin',\n    'implant', 'density','machine_id', 'difficult_negative_case',\n    'cancer',\n]\n\ndf['stratify'] = ''\nfor col in strat_cols:\n    df['stratify'] += df[col].astype(str)\n\nskf = StratifiedGroupKFold(n_splits=CFG.folds, shuffle=True, random_state=CFG.seed)\nfor fold, (train_idx, val_idx) in enumerate(skf.split(df, df['stratify'], df[\"patient_id\"])):\n    df.loc[val_idx, 'fold'] = fold\ndisplay(df.groupby(['fold', \"cancer\"]).size())","metadata":{"execution":{"iopub.status.busy":"2022-12-13T20:20:14.454034Z","iopub.execute_input":"2022-12-13T20:20:14.454315Z","iopub.status.idle":"2022-12-13T20:20:23.653515Z","shell.execute_reply.started":"2022-12-13T20:20:14.454291Z","shell.execute_reply":"2022-12-13T20:20:23.652326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# KerasCV - Augmentation\n<div align=\"center\"><img src=\"https://keras.io/img/logo.png\" width=\"400\"></div>\n\n[**KerasCV**](https://github.com/keras-team/keras-cv/) is powerful tool which is made of modular building blocks (ops, functions, layers, metrics, losses, callbacks) that standardizes APIs for computer vision concepts such as **data-augmentation pipeline** and bounding boxes. Below I have demonstrated how to incorporate KerasCV in `tf.data` pipeline easily. \n    \n> **Note**: Some functions specifically `zoom`, `translation`, `rotation` throws error, most likely due to dynamic dimension. Need to carry out more experiments with them to debug it.","metadata":{}},{"cell_type":"markdown","source":"## KerasCV - RandomRot90\nRotate image 90 degree randomly","metadata":{}},{"cell_type":"code","source":"from keras_cv.layers.preprocessing.base_image_augmentation_layer import (\n    BaseImageAugmentationLayer,\n)\n\nclass RandomRot90(BaseImageAugmentationLayer):\n    def __init__(self, rot_range=3, seed=None, **kwargs):\n        super().__init__(seed=seed, force_generator=True, **kwargs)\n        self.rot_range = rot_range\n        self.seed = seed\n        self.auto_vectorize = True\n\n    def augment_label(self, label, transformation, **kwargs):\n        return label\n\n    def augment_image(self, image, transformation, **kwargs):\n        return RandomRot90._rot_image(image, transformation)\n\n    def get_random_transformation(self, **kwargs):\n        num_rotate = self._random_generator.random_uniform(shape=[],\n                                                           minval=-self.rot_range,\n                                                           maxval=self.rot_range+1)\n        return {\n            \"num_rotate\": tf.cast(num_rotate, dtype=tf.int32),\n        }\n\n    def _rot_image(image, transformation):\n        rotted_output = tf.image.rot90(image, k=transformation[\"num_rotate\"])\n        rotted_output.set_shape(image.shape)\n        return rotted_output","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-12-13T20:20:23.654945Z","iopub.execute_input":"2022-12-13T20:20:23.655311Z","iopub.status.idle":"2022-12-13T20:20:24.542359Z","shell.execute_reply.started":"2022-12-13T20:20:23.655281Z","shell.execute_reply":"2022-12-13T20:20:24.541071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## KerasCV - Pipeline\nWe have created a simple pipline using KerasCV, it uses `kcv.layers.Augmenter` to create pipeline. You can checkout this [link](https://github.com/keras-team/keras-cv/blob/master/keras_cv/layers/preprocessing/) for more transformation, there are so many don't get intimidated.","metadata":{}},{"cell_type":"code","source":"import keras_cv as kcv\n\naugmenter = kcv.layers.Augmenter(\n  layers=[\n    kcv.layers.RandomFlip(mode=CFG.flip_mode),\n    kcv.layers.RandAugment(value_range=(0, 1),\n                           augmentations_per_image=3,\n                           magnitude=0.05,\n                          geometric=True,\n                          rate=0.5),\n#     kcv.layers.Posterization(value_range=(0, 1), bits=4),\n#     kcv.layers.Solarization(value_range=(0, 1)),\n    kcv.layers.RandomBrightness(factor=0.1, value_range=(0, 1), rate=0.5,),\n    kcv.layers.RandomContrast(factor=0.1),\n    kcv.layers.RandomSaturation(factor=0.1),\n    kcv.layers.GridMask(ratio_factor=(0.10, 0.15)),\n#     kcv.layers.RandomGaussianBlur(kernel_size=(3,3), factor=0.05),\n    kcv.layers.RandomHue(0.05, value_range=(0, 1)),\n    kcv.layers.RandomShear(0.1, 0.1, fill_mode='constant'),\n#     RandomRot90(rot_range=CFG.rot_range)\n    ]\n)\n\n# apply function\ndef apply_kcv(image, label=None):\n    \"\"\"Accepts inputs in both image and (image, label) format\"\"\"\n    if tf.random.uniform([]) < CFG.augment:\n        image = augmenter(image)\n    if label is None:\n        return image\n    else:\n        return image, label","metadata":{"execution":{"iopub.status.busy":"2022-12-13T20:29:11.381997Z","iopub.execute_input":"2022-12-13T20:29:11.382653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Pipeline\n* Reads the raw file and then decodes it to tf.Tensor\n* Resizes the image in desired size\n* Chages the datatype to **float32**\n* Caches the Data for boosting up the speed.\n* Uses Augmentations to reduce overfitting and make model more robust.\n* Finally, splits the data into batches.\n","metadata":{}},{"cell_type":"code","source":"def build_decoder(with_labels=True, target_size=CFG.img_size, ext='png'):\n    def decode(path):\n        file_bytes = tf.io.read_file(path)\n        if ext == 'png':\n            img = tf.image.decode_png(file_bytes, channels=3, dtype=tf.uint16)\n        elif ext in ['jpg', 'jpeg']:\n            img = tf.image.decode_jpeg(file_bytes, channels=3)\n        else:\n            raise ValueError(\"Image extension not supported\")\n\n        img = tf.cast(img, tf.float32) / 65536.0\n        img = tf.image.resize(img, target_size, method='bilinear')\n        img = tf.reshape(img, [*target_size, 3])\n        return img\n    \n    def decode_with_labels(path, label):\n        return decode(path), tf.cast(label, tf.float32)\n    \n    return decode_with_labels if with_labels else decode\n\n\ndef build_dataset(paths, labels=None, batchify=False,\n                  batch_size=32, cache=True,\n                  decode_fn=None, augment_fn=None,\n                  augment=True, repeat=True, shuffle=1024, \n                  cache_dir=\"\", drop_remainder=False):\n    if cache_dir != \"\" and cache is True:\n        os.makedirs(cache_dir, exist_ok=True)\n    \n    if decode_fn is None:\n        decode_fn = build_decoder(labels is not None)\n        \n    if augment_fn is None:\n        augment_fn = apply_kcv\n    \n    AUTO = tf.data.experimental.AUTOTUNE\n    slices = paths if labels is None else (paths, labels)\n    \n    ds = tf.data.Dataset.from_tensor_slices(slices)\n    ds = ds.map(decode_fn, num_parallel_calls=AUTO)\n    ds = ds.cache(cache_dir) if cache else ds\n    ds = ds.repeat() if repeat else ds\n    if shuffle: \n        ds = ds.shuffle(shuffle, seed=CFG.seed)\n        opt = tf.data.Options()\n        opt.experimental_deterministic = False\n        ds = ds.with_options(opt)\n    ds = ds.map(augment_fn, num_parallel_calls=AUTO) if augment else ds\n    if batchify:\n        ds = ds.batch(batch_size, drop_remainder=drop_remainder)\n        ds = ds.prefetch(AUTO)\n    return ds\n\ndef make_balanced_ds(paths, \n                     labels, \n                     pos_prob=0.25,\n                     batch_size=32,\n                     cache=True,\n                     augment=True,\n                     repeat=True,\n                     drop_remainder=False,\n                    shuffle=1024):\n    \"\"\"create balanced dataset using weights\"\"\"\n    pos_idx = labels==1; neg_idx = labels==0;\n    pos_paths = paths[pos_idx]; pos_labels = labels[pos_idx]\n    neg_paths = paths[neg_idx]; neg_labels = labels[neg_idx]\n    pos_ds = build_dataset(pos_paths, \n                           pos_labels, \n                           batch_size=batch_size, \n                           cache=cache,\n                           augment=augment,\n                           shuffle=shuffle,\n                          repeat=repeat)\n    neg_ds = build_dataset(neg_paths, \n                           neg_labels, \n                           batch_size=batch_size, \n                           cache=cache,\n                           augment=augment,\n                           shuffle=shuffle,\n                          repeat=repeat,)\n    ds = tf.data.Dataset.sample_from_datasets([pos_ds, neg_ds], weights=[pos_prob, 1. -  pos_prob])\n    ds = ds.batch(batch_size, drop_remainder=drop_remainder)\n    ds = ds.prefetch(AUTO)\n    return ds","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualization\n* Check if augmentation is working properly or not.","metadata":{}},{"cell_type":"code","source":"def display_batch(batch, size=2):\n    if isinstance(batch, tuple):\n        imgs, tars = batch\n        tars = tars.numpy().squeeze()\n    else:\n        imgs = batch\n        tars = None\n    plt.figure(figsize=(size*5, 5))\n    for img_idx in range(size):\n        plt.subplot(1, size, img_idx+1)\n        if tars is not None:\n            plt.title(f'{CFG.target_col[0]}: {tars[img_idx]:0.3f}', fontsize=15)\n        plt.imshow(imgs[img_idx,:, :, :])\n        plt.xticks([]); plt.yticks([])\n    plt.tight_layout()\n    plt.show() ","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Generate Image","metadata":{}},{"cell_type":"code","source":"fold = 0\nfold_df = df.groupby('cancer').sample(frac=0.1)\npaths  = fold_df.image_path.values\nlabels = fold_df[CFG.target_col].values.ravel()\nds = make_balanced_ds(paths, labels, pos_prob=0.25, cache=False, batch_size=32, augment=False, repeat=True)\nbatch = next(iter(ds))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-13T20:20:24.656625Z","iopub.execute_input":"2022-12-13T20:20:24.65689Z","iopub.status.idle":"2022-12-13T20:20:29.65808Z","shell.execute_reply.started":"2022-12-13T20:20:24.656865Z","shell.execute_reply":"2022-12-13T20:20:29.656725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## No Augmentation","metadata":{}},{"cell_type":"code","source":"display_batch(batch, 5);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-13T20:20:29.659673Z","iopub.execute_input":"2022-12-13T20:20:29.660055Z","iopub.status.idle":"2022-12-13T20:20:31.12678Z","shell.execute_reply.started":"2022-12-13T20:20:29.660022Z","shell.execute_reply":"2022-12-13T20:20:31.125578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## After Augmentation","metadata":{}},{"cell_type":"code","source":"display_batch((apply_kcv(batch[0]), batch[1]), 5)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-13T20:20:31.128216Z","iopub.execute_input":"2022-12-13T20:20:31.128571Z","iopub.status.idle":"2022-12-13T20:20:40.339659Z","shell.execute_reply.started":"2022-12-13T20:20:31.12854Z","shell.execute_reply":"2022-12-13T20:20:40.338599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loss Function\n## BCE\nLoss Function for this notebook is **BCE: Binary Crossentropy** or **Focal** as the task is **binary classification**\n\n$$\\textrm{BCE}  = -\\frac{1}{N} \\sum_{i=1}^{N} y_i \\cdot log(\\hat{y}_i) + (1 - y_i) \\cdot log(1 - \\hat{y}_i)$$\n\nwhere $\\hat{y}_i$ is the **predicted** value and $y_i$ is the **original** value for each instance $i$.\n\n## Focal\n$$\\textrm{FL} = -\\alpha_{t}(1 - p_{t})^{\\gamma}\\log{p_{t}}$$\n\n$\\gamma$ controls the shape of the curve. The higher the value of $\\gamma$, the lower the loss for well-classified examples, so we could turn the attention of the model more towards ‘hard-to-classify examples. FL gives high weights to the rare class and small weights to the dominating or common class. These weights are referred to as $\\alpha$.\n\n\n## Code\n* `tf.keras.losses.BinaryCrossentropy`\n* `tfa.losses.SigmoidFocalCrossEntropy`","metadata":{}},{"cell_type":"markdown","source":"# Metric\nMetric for this competition is **probabilistic F1 score (pF1)**. This extension of the traditional F score accepts probabilities instead of binary classifications,\n\n$$\npF1=2\\frac{pPrecision⋅pRecall}{pPrecision+pRecall}\n$$\n\nWhere,\n$$\npPrecision=\\frac{pTP}{pTP+pFP}\n$$\n$$\npRecall=\\frac{pTP}{TP+FN}\n$$","metadata":{}},{"cell_type":"code","source":"# tensorflow\ndef pfbeta_tf(labels, preds, beta=1):\n    eps = 1e-5\n    preds = tf.clip_by_value(preds, 0, 1)\n    y_true_count = tf.reduce_sum(labels)\n    ctp = tf.reduce_sum(preds[labels==1])\n    cfp = tf.reduce_sum(preds[labels==0])\n    beta_squared = beta * beta\n    c_precision = ctp / (ctp + cfp + eps)\n    c_recall = ctp / (y_true_count + eps)\n    if (c_precision > 0 and c_recall > 0):\n        result = (1 + beta_squared) * (c_precision * c_recall) / (beta_squared * c_precision + c_recall + eps)\n        return result\n    else:\n        return tf.constant(0, dtype=tf.float32)\npfbeta_tf.__name__='pF1'\n\n\n# finds best pf1 using thresholds\ndef pfbeta_thr(labels, preds):\n    thrs = tf.range(0, 1, 0.05)\n    best_score = tf.constant(0, dtype=tf.float32)\n    for thr in thrs:\n        score = pfbeta_tf(labels, tf.cast(preds>thr, tf.float32))\n        best_score = tf.cond(score > best_score, lambda: score, lambda: best_score)\n    return best_score\n\npfbeta_thr.__name__='pF1_thr'\n\n# numpy\ndef pfbeta(labels, preds, beta=1):\n    eps = 1e-5\n    preds = preds.clip(0, 1)\n    y_true_count = labels.sum()\n    ctp = preds[labels==1].sum()\n    cfp = preds[labels==0].sum()\n    beta_squared = beta * beta\n    c_precision = ctp / (ctp + cfp + eps)\n    c_recall = ctp / (y_true_count + eps)\n    if (c_precision > 0 and c_recall > 0):\n        result = (1 + beta_squared) * (c_precision * c_recall) / (beta_squared * c_precision + c_recall + eps)\n        return result\n    else:\n        return 0.0","metadata":{"execution":{"iopub.status.busy":"2022-12-13T20:20:40.340996Z","iopub.execute_input":"2022-12-13T20:20:40.341409Z","iopub.status.idle":"2022-12-13T20:20:40.377722Z","shell.execute_reply.started":"2022-12-13T20:20:40.341378Z","shell.execute_reply":"2022-12-13T20:20:40.37682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Build Model","metadata":{}},{"cell_type":"code","source":"import efficientnet.tfkeras as efn\n\ndef build_model(model_name=CFG.model_name,\n                loss_name=CFG.loss,\n                dim=CFG.img_size,\n                compile_model=True,\n                include_top=False):         \n    base = getattr(efn, model_name)(input_shape=(*dim,3),\n                                    weights='imagenet',\n                                    include_top=False) # get base model (efficientnet), use imgnet weights\n    inp = base.inputs\n    x = base.output\n    x = tf.keras.layers.GlobalAveragePooling2D()(x) # use GAP to get pooling result form conv outputs\n#     x = tf.keras.layers.Dense(32, activation='silu')(x) # use activation to apply non-linearity\n    x = tf.keras.layers.Dense(1,activation='sigmoid')(x) # use sigmoid to convert predictions to [0-1]\n    model = tf.keras.Model(inputs=inp,outputs=x)\n    if compile_model:\n        # optimizer\n        opt = tf.keras.optimizers.Adam(learning_rate=0.0001)\n        # loss\n        if loss_name == 'BCE':\n            loss = tf.keras.losses.BinaryCrossentropy(label_smoothing=0.0)\n        elif loss_name == 'Focal':\n            loss = tfa.losses.SigmoidFocalCrossEntropy(alpha=0.80, gamma=2.0)\n        # metric\n        auc = tf.keras.metrics.AUC(name='auc')\n        pf1 = pfbeta_tf\n        pf1_thr = pfbeta_thr\n        metrics = [pf1, pf1_thr, auc]\n        # compile\n        model.compile(optimizer=opt,\n                      loss=loss,\n                      metrics=metrics)\n    return model","metadata":{"execution":{"iopub.status.busy":"2022-12-13T20:20:40.378901Z","iopub.execute_input":"2022-12-13T20:20:40.379211Z","iopub.status.idle":"2022-12-13T20:20:40.453796Z","shell.execute_reply.started":"2022-12-13T20:20:40.379184Z","shell.execute_reply":"2022-12-13T20:20:40.452968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model Check","metadata":{}},{"cell_type":"code","source":"tmp = build_model(CFG.model_name, dim=CFG.img_size, compile_model=True)\n# tmp.summary()  # too long","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-12-13T20:20:40.454941Z","iopub.execute_input":"2022-12-13T20:20:40.455199Z","iopub.status.idle":"2022-12-13T20:20:45.318527Z","shell.execute_reply.started":"2022-12-13T20:20:40.455176Z","shell.execute_reply":"2022-12-13T20:20:45.317389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Learning-Rate Scheduler","metadata":{}},{"cell_type":"code","source":"def get_lr_callback(batch_size=8, plot=False): \n    lr_start   = 0.00001\n    lr_max     = 0.00006\n    lr_min     = 0.00001\n    lr_ramp_ep = 4\n    lr_sus_ep  = 0\n    lr_decay   = 0.8\n   \n    def lrfn(epoch):\n        if epoch < lr_ramp_ep:\n            lr = (lr_max - lr_start) / lr_ramp_ep * epoch + lr_start\n            \n        elif epoch < lr_ramp_ep + lr_sus_ep:\n            lr = lr_max\n            \n        elif CFG.scheduler=='exp':\n            lr = (lr_max - lr_min) * lr_decay**(epoch - lr_ramp_ep - lr_sus_ep) + lr_min\n            \n        elif CFG.scheduler=='cosine':\n            decay_total_epochs = CFG.epochs - lr_ramp_ep - lr_sus_ep + 3\n            decay_epoch_index = epoch - lr_ramp_ep - lr_sus_ep\n            phase = math.pi * decay_epoch_index / decay_total_epochs\n            cosine_decay = 0.4 * (1 + math.cos(phase))\n            lr = (lr_max - lr_min) * cosine_decay + lr_min\n        return lr\n    if plot:\n        plt.figure(figsize=(10,5))\n        plt.plot(np.arange(1, CFG.epochs+1), [lrfn(epoch) for epoch in np.arange(CFG.epochs)], marker='o')\n        plt.xlabel('epoch'); plt.ylabel('learnig rate')\n        plt.title('Learning Rate Scheduler')\n        plt.show()\n\n    lr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose=False)\n    return lr_callback\n\n_=get_lr_callback(CFG.batch_size, plot=True )","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-13T20:20:45.319661Z","iopub.execute_input":"2022-12-13T20:20:45.32023Z","iopub.status.idle":"2022-12-13T20:20:45.508226Z","shell.execute_reply.started":"2022-12-13T20:20:45.320197Z","shell.execute_reply":"2022-12-13T20:20:45.507321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Grad-CAM\nGradient-weighted Class Activation Mapping (Grad-CAM), uses the class-specific gradient information flowing into the final convolutional layer of a CNN to produce a coarse localization map of the important regions in the image. Grad-CAM is a strict generalization of the Class Activation Mapping. Grad-CAM provides visual explanations to better understand image classification problems.\n\n<img src=\"http://gradcam.cloudcv.org/static/images/network.png\" width=\"800\">","metadata":{}},{"cell_type":"code","source":"import matplotlib.cm as cm, cv2\nfrom tensorflow import keras\n\ndef gen_gradcam_heatmap(img, model, last_conv='top_conv', pred_index=0):\n    # First, we create a model that maps the input image to the activations\n    # of the last conv layer as well as the output predictions\n    grad_model = tf.keras.models.Model(\n        [model.inputs], [model.get_layer(last_conv).output, model.output]\n    )\n\n    # Then, we compute the gradient of the top predicted class for our input image\n    # with respect to the activations of the last conv layer\n    with tf.GradientTape() as tape:\n        features, preds = grad_model(img)\n        if pred_index is None:\n            pred_index = tf.argmax(preds[0])\n        class_channel = preds[:, pred_index]\n\n    # This is the gradient of the output neuron (top predicted or chosen)\n    # with regard to the output feature map of the last conv layer\n    grads = tape.gradient(class_channel, features)\n\n    # This is a vector where each entry is the mean intensity of the gradient\n    # over a specific feature map channel\n    pooled_grads = tf.reduce_mean(grads, axis=(0, 1, 2))\n\n    # We multiply each channel in the feature map array\n    # by \"how important this channel is\" with regard to the top predicted class\n    # then sum all the channels to obtain the heatmap class activation\n    features = features[0]\n    heatmap = features @ pooled_grads[..., tf.newaxis]\n    heatmap = tf.squeeze(heatmap)\n\n    # For visualization purpose, we will also normalize the heatmap between 0 & 1\n    heatmap = tf.maximum(heatmap, 0) / tf.math.reduce_max(heatmap)\n    return heatmap.numpy()\n\ndef get_gradcam(img, model,  alpha=0.4, show=False):\n\n    heatmap = gen_gradcam_heatmap(img, model, last_conv='top_conv', pred_index=0)\n    img     = img[0]\n    # Rescale heatmap to a range 0-255\n    heatmap = np.uint8(255 * heatmap)\n\n    # Use jet colormap to colorize heatmap\n    jet = cm.get_cmap(\"jet\")\n\n    # Use RGB values of the colormap\n    jet_colors  = jet(np.arange(256))[:, :3]\n    jet_heatmap = jet_colors[heatmap]\n\n    # Create an image with RGB colorized heatmap\n    jet_heatmap = cv2.resize(jet_heatmap, dsize=(img.shape[1], img.shape[0]))\n\n    # Superimpose the heatmap on original image\n    superimposed_img = jet_heatmap * alpha + img\n    superimposed_img = keras.preprocessing.image.array_to_img(superimposed_img)\n#     superimposed_img = np.uint8(superimposed_img*255.0)\n\n    # Display Grad CAM\n    if show:\n        plt.imshow(superimposed_img)\n        \n    return superimposed_img","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-13T20:20:45.509444Z","iopub.execute_input":"2022-12-13T20:20:45.509814Z","iopub.status.idle":"2022-12-13T20:20:46.344126Z","shell.execute_reply.started":"2022-12-13T20:20:45.509786Z","shell.execute_reply":"2022-12-13T20:20:46.343133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Wandb** Logger\nLog:\n* Best Score\n* Attention MAP","metadata":{}},{"cell_type":"code","source":"# create directory to save gradcam imgs\n!mkdir -p gradcam","metadata":{"execution":{"iopub.status.busy":"2022-12-13T20:20:46.345248Z","iopub.execute_input":"2022-12-13T20:20:46.345627Z","iopub.status.idle":"2022-12-13T20:20:47.341585Z","shell.execute_reply.started":"2022-12-13T20:20:46.345599Z","shell.execute_reply":"2022-12-13T20:20:47.340231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# if CFG.wandb:\ndef wandb_init(fold):\n    config = {k:v for k,v in dict(vars(CFG)).items() if '__' not in k}\n    config.update({\"fold\":int(fold)})\n    yaml.dump(config, open(f'/kaggle/working/config fold-{fold}.yaml', 'w'),)\n    config = yaml.load(open(f'/kaggle/working/config fold-{fold}.yaml', 'r'), Loader=yaml.FullLoader)\n    run    = wandb.init(project=\"rsna-bcd-public\",\n               name=f\"fold-{fold}|dim-{CFG.img_size[0]}x{CFG.img_size[1]}|model-{CFG.model_name}\",\n               config=config,\n               anonymous=anonymous,\n               group=CFG.comment\n                    )\n    return run\n\ndef log_wandb(fold):\n    \"log best result for error analysis\"\n    valid_df = df.query(\"fold==@fold\").copy().reset_index(drop=True)\n    if CFG.debug:\n        valid_df = valid_df[:min_samples]\n    valid_df['pred'] = oof_pred[fold].reshape(-1)\n    valid_df['diff'] = abs(valid_df.cancer - valid_df.pred)\n    valid_df = valid_df.sort_values(by='diff', ascending=False)\n    \n    noimg_cols  = ['site_id', 'patient_id', 'image_id', 'laterality', 'view', 'age',\n                   'cancer', 'biopsy', 'invasive', 'BIRADS', 'implant', 'density',\n                   'machine_id', 'difficult_negative_case','width','height', 'fold'] + ['pred','diff']\n    \n    # select top and worst 10 cases for each class\n    gradcam_df  = pd.concat((valid_df.groupby('cancer').tail(10), \n                             valid_df.groupby('cancer').head(10)), axis=0, ignore_index=True)\n    gradcam_ds  = build_dataset(gradcam_df.image_path, labels=None, batchify=True, cache=False, \n                                batch_size=1, repeat=False, shuffle=False, augment=False)\n    \n    # create wandb table for upload\n    data = []\n    for idx, img in enumerate(tqdm(gradcam_ds, total=40, desc='gradcam ', position=0, leave=True)):\n        gradcam = get_gradcam(img, model)\n        row = gradcam_df[noimg_cols].iloc[idx].tolist()\n        img = img.numpy()[0]\n        img = (img*255.0).astype('uint8')\n        data+=[[*row, wandb.Image(img), wandb.Image(gradcam)]]\n        if idx<10: # save best ones\n            cv2.imwrite(f'gradcam/fold{fold}_{idx:02d}.png', np.concatenate([img, gradcam], axis=1))\n    wandb_table = wandb.Table(data=data, columns=[*noimg_cols, 'image', 'gradcam'])\n    \n    # log values to wandb\n    best_epoch = np.argmax(history.history['val_pF1_thr'])\n    wandb.log({\n               'best_pF1_batch':oof_val[-1], \n               'best_pF1':pF1,\n               'best_pF1_thr':pF1_thr,\n               'best_auc':auc,\n               'best_epoch':best_epoch,\n               'viz_table':wandb_table,\n              })","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-13T20:20:47.343156Z","iopub.execute_input":"2022-12-13T20:20:47.343552Z","iopub.status.idle":"2022-12-13T20:20:47.363259Z","shell.execute_reply.started":"2022-12-13T20:20:47.343523Z","shell.execute_reply":"2022-12-13T20:20:47.362458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train Model\n* Cross-Validation: 5 fold\n* **WandB** dashboard is shown end of the each fold. So we don't need to plot anything. We can select best model from here.","metadata":{}},{"cell_type":"code","source":"oof_pred = []; oof_tar = []; oof_val = []; oof_ids = []; oof_folds = []\npreds = np.zeros((test_df.shape[0],1))\n\nfor fold in np.arange(CFG.folds):\n    \n    # ignore not selected folds\n    if fold not in CFG.selected_folds:\n        continue\n        \n    # init wandb\n    if CFG.wandb:\n        run = wandb_init(fold)\n        WandbCallback = wandb.keras.WandbCallback(save_model=False)\n            \n    # train and valid dataframe\n    train_df = df.query(\"fold!=@fold\")\n    valid_df = df.query(\"fold==@fold\")\n    \n    # upsample cancer data\n    if CFG.upsample>1:\n        pos_df = train_df.query(\"cancer==1\").sample(frac=CFG.upsample, replace=True)\n        neg_df = train_df.query(\"cancer==0\")\n        train_df = pd.concat([pos_df, neg_df], axis=0, ignore_index=True)\n    \n    # get image_paths and labels\n    train_paths = train_df.image_path.values; train_labels = train_df[CFG.target_col].values.astype(np.float32)\n    valid_paths = valid_df.image_path.values; valid_labels = valid_df[CFG.target_col].values.astype(np.float32)\n    test_paths  = test_df.image_path.values\n    \n    # shuffle train data\n    index = np.arange(len(train_df))\n    np.random.shuffle(index)\n    train_paths  = train_paths[index]\n    train_labels = train_labels[index]\n    \n    # min samples in debug mode\n    min_samples = CFG.batch_size*REPLICAS*2\n    \n    # for debug model run on small portion\n    if CFG.debug:\n        train_paths = train_paths[:min_samples]; train_labels = train_labels[:min_samples]\n        valid_paths = valid_paths[:min_samples]; valid_labels = valid_labels[:min_samples]\n    \n    # show message\n    print('#'*40); print('#### FOLD: ',fold)\n    print('#### IMAGE_SIZE: (%i, %i) | MODEL_NAME: %s | BATCH_SIZE: %i'%\n          (CFG.img_size[0],CFG.img_size[1],CFG.model_name,CFG.batch_size*REPLICAS))\n    num_train = len(train_paths)\n    num_valid = len(valid_paths)\n    if CFG.wandb:\n        wandb.log({'num_train':num_train,\n                   'num_valid':num_valid})\n    print('#### NUM_TRAIN: {:,} | NUM_VALID: {:,}'.format(num_train, num_valid))\n    \n    # build model\n    K.clear_session()\n    with strategy.scope():\n        model = build_model(CFG.model_name, dim=CFG.img_size, compile_model=True)\n\n    # build dataset\n    cache = False # 'TPU' not in CFG.device\n    train_ds = make_balanced_ds(train_paths, train_labels.ravel(), pos_prob=CFG.pos_prob,\n                                cache=cache, batch_size=CFG.batch_size*REPLICAS, repeat=True, \n                                shuffle=True, augment=CFG.augment,)\n    valid_ds = build_dataset(valid_paths, valid_labels.ravel(), batchify=True,\n                             cache=cache, batch_size=CFG.batch_size*REPLICAS, repeat=False,\n                             shuffle=False, augment=False,)\n    \n    # calculate steps per epoch after sampling\n    steps_per_epoch = len(train_labels[train_labels==0])/((1. - CFG.pos_prob)*CFG.batch_size*REPLICAS)\n    \n    \n    print('#'*40)   \n    \n    # callbacks\n    callbacks = []\n    ## save best model after each fold\n    sv = tf.keras.callbacks.ModelCheckpoint(\n        'fold-%i.h5'%fold, monitor='val_pF1', verbose=CFG.verbose, save_best_only=True,\n        save_weights_only=False, mode='max', save_freq='epoch')\n    callbacks +=[sv]\n    ## lr-scheduler\n    callbacks += [get_lr_callback(CFG.batch_size)]\n    ## wandb callback\n    if CFG.wandb:\n        callbacks.append(WandbCallback)\n        \n    # train\n    print('Training...')\n    history = model.fit(\n        train_ds, \n        epochs=CFG.epochs if not CFG.debug else 2, \n        callbacks = callbacks, \n        steps_per_epoch=steps_per_epoch,\n        validation_data=valid_ds, \n        #class_weight = {0:1,1:2},\n        verbose=CFG.verbose\n    )\n    \n    # load best model for inference\n    print('Loading best model...')\n    model.load_weights('fold-%i.h5'%fold)  \n    \n    # predict on valid data\n    print('Predicting OOF with TTA...')\n    ds_valid = build_dataset(valid_paths, labels=None, batchify=True, cache=False, \n                             batch_size=CFG.batch_size*REPLICAS*2, repeat=True, \n                             shuffle=False, augment=CFG.tta>1)\n    ct_valid = len(valid_paths); STEPS = CFG.tta * ct_valid/CFG.batch_size/2/REPLICAS\n    pred = model.predict(ds_valid,steps=STEPS,verbose=CFG.verbose)[:CFG.tta*ct_valid,] \n    oof_pred.append(np.mean(pred.reshape((CFG.tta, ct_valid,-1)),axis=0))                 \n    \n    # get id and target for valid data\n    oof_tar.append(valid_df[CFG.target_col].values[:(min_samples if CFG.debug else len(valid_df))])\n    oof_folds.append(np.ones_like(oof_tar[-1],dtype='int8')*fold)\n    oof_ids.append(valid_df.image_path.tolist()[:(min_samples if CFG.debug else len(valid_df))])\n    \n#     # predict on test data\n#     print('Predicting Test...')\n#     ds_test = build_dataset(test_paths, labels=None, cache=False, \n#                     batch_size=(CFG.batch_size*2 if len(test_df)>4 else 1)*REPLICAS,\n#                    repeat=True, shuffle=False, augment=CFG.tta>1)\n#     ct_test = len(test_paths); STEPS = 1 if len(test_df)<=4 else (CFG.tta * ct_test/CFG.batch_size/2/REPLICAS)\n#     pred = model.predict(ds_test,steps=STEPS,verbose=CFG.verbose)[:CFG.tta*ct_test,] \n#     preds[:ct_test, :] += np.mean(pred.reshape((CFG.tta, ct_test,-1)),axis=0) / CFG.folds # not meaningful for DIBUG = True\n    \n    # store best results\n    y_true = oof_tar[-1].astype(np.float32); y_pred = oof_pred[-1]\n    pF1 = pfbeta(y_true, y_pred)\n    pF1_thr = pfbeta_thr(y_true, y_pred).numpy() # tf metric\n    auc = roc_auc_score(y_true, y_pred)\n    oof_val.append(np.max(history.history['val_pF1'] ))\n    print('>>>> FOLD %i OOF pF1_batch = %.3f, pF1 = %.3f, pF1_thr = %.3f, auc = %.3f\\n'%(fold,oof_val[-1], \n                                                                                         pF1, \n                                                                                         pF1_thr,\n                                                                                         auc))  # pF1_batch => pF1 batchwise canculated then aggregated\n    \n    # log best result on wandb & plot\n    if CFG.wandb:\n        log_wandb(fold) # log\n        wandb.run.finish() # finish the run\n        display(ipd.IFrame(run.url, width=1080, height=720)) # show wandb dashboard","metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2022-12-13T20:20:47.368861Z","iopub.execute_input":"2022-12-13T20:20:47.369164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualize Grad-CAM\nMake sure `CFG.wandb=True` for Grad-CAM","metadata":{}},{"cell_type":"code","source":"if CFG.wandb:\n    from mpl_toolkits.axes_grid1 import ImageGrid\n\n    imgs = [cv2.imread(path) for path in glob('/kaggle/working/gradcam/*')[:8]]\n    fig = plt.figure(figsize=(20., 30.))\n    grid = ImageGrid(fig, 111,  # similar to subplot(111)\n                     nrows_ncols=(4, 2),  # creates 2x2 grid of axes\n                     axes_pad=0.25,  # pad between axes in inch.\n                     )\n\n    for ax, im in zip(grid, imgs):\n        # Iterating over the grid returns the Axes.\n        ax.imshow(im)\n        ax.set_xticks([])\n        ax.set_yticks([])\n\n    plt.show()","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Calculate OOF Score","metadata":{}},{"cell_type":"code","source":"# overall oof pF1\noof = np.concatenate(oof_pred); true = np.concatenate(oof_tar);\nids = np.concatenate(oof_ids); folds = np.concatenate(oof_folds)\npF1 = pfbeta(true.astype(np.float32),oof)\nprint('Overall OOF pF1 = %.3f'%pF1)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# save oof\ncolumns = ['image_path', 'true', 'pred']\ndf_oof = pd.DataFrame(np.concatenate([ids[:,None], true, oof], axis=1), columns=columns)\ndf_oof = df_oof.merge(df, on=['image_path'], how='left') # merge with train data\ndf_oof.to_csv('oof.csv',index=False)\ndf_oof.head(2)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# `pF1` Maximize","metadata":{}},{"cell_type":"code","source":"# calculate pf1 for multiple thresholds\nthrs = np.arange(0,1,0.05)\nscores = []\nfor thr in tqdm(thrs):\n    scores+=[pfbeta(df_oof.true.astype('float32'), df_oof.pred.astype('float32')>thr)]\n    \n# get best thr for max pF1\nbest_score_idx = np.argmax(scores)\nbest_score = np.max(scores)\nbest_thr = thrs[best_score_idx]\nprint(f'\\n## MAX pF1 = {best_score: 0.3f} @ {best_thr:0.3f}\\n')\n\n# plot thr vs pF1 graph\nfig, ax = plt.subplots(figsize=(10,6))\nax.fill_between(thrs, scores,\n                color='red', alpha=0.3, )\nax.plot(thrs, scores, '-ok');\nax.axvline(x=best_thr, color='blue', ls='--')\nax.plot(best_thr, best_score, color='blue', marker='o', markersize=12)\nax.set_xlabel('Threshold');\nax.set_ylabel('pF1');\nax.set_title('Threshold Vs pF1');","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction Distribution of OOF & Train \nCheck **Cancer** distribution of `train` and `oof`. ","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nsns.set(style='dark')\n\nplt.figure(figsize=(10*2,6))\n\nplt.subplot(1, 2, 1)\nsns.kdeplot(x=train_df[CFG.target_col[0]], color='b',fill=True);\nsns.kdeplot(x=df_oof.pred.values.astype('float32'), color='r',fill=True);\nplt.grid('ON')\nplt.xlabel(CFG.target_col[0]);plt.ylabel('freq');plt.title('KDE')\nplt.legend(['train', 'oof'])\n\nplt.subplot(1, 2, 2)\nsns.histplot(x=train_df[CFG.target_col[0]], color='b');\nsns.histplot(x=df_oof.pred.values.astype('float32'), color='r');\nplt.grid('ON')\nplt.xlabel(CFG.target_col[0]);plt.ylabel('freq');plt.title('Histogram')\nplt.legend(['train', 'oof'])\n\nplt.tight_layout()\nplt.show()","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Remove Files","metadata":{}},{"cell_type":"code","source":"!rm -r /kaggle/working/wandb","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Reference:\n1. [RANZCR: EfficientNet TPU Training](https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training)\n1. [Triple Stratified KFold with TFRecords](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords)","metadata":{}}]}