{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# [RSNA Screening Mammography Breast Cancer Detection](https://www.kaggle.com/c/petfinder-pawpularity-score)\n> Find breast cancers in screening mammograms\n\n![](https://storage.googleapis.com/kaggle-competitions/kaggle/39272/logos/header.png?t=2022-11-28-17-29-35)","metadata":{}},{"cell_type":"markdown","source":"# Idea:\n* In this notebook will generate **bounding box (bbox)** from **gradcam** images.\n* BBoxes can be then used with Object Detection models for **Cancer Detection**\n* **Wandb** is integrated hence we can use this notebook for visual analysis.\n* **Image**, **Mask**, **GradCAM** & **BBox** is logged in **WandB**","metadata":{}},{"cell_type":"markdown","source":"# Notebooks\n* Only Image:\n    * ROI:\n        * train: [RSNA-BCD: EfficientNet [TF][TPU-1VM][Train]](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train)\n        * infer: [RSNA-BCD: EfficientNet [TF][TPU-1VM][Infer]](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-infer)\n    * NoROI + KerasCV: \n        * train: [RSNA-BCD: NoROI KerasCV [TF][Train]](https://www.kaggle.com/awsaf49/rsna-bcd-noroi-kerascv-tf-train/)\n        * infer: [RSNA-BCD: NoROI KerasCV [TF][Infer]](https://www.kaggle.com/awsaf49/rsna-bcd-noroi-kerascv-tf-infer/)\n* Dataset:\n    * ROI:\n        * [RSNA-BCD: ROI 1024x PNG Dataset](https://www.kaggle.com/datasets/awsaf49/rsna-bcd-roi-1024x-png-dataset)\n    * NoROI:\n        * [RSNA-BCD: 512 PNG v2 PNG Dataset](https://www.kaggle.com/datasets/awsaf49/rsnabcd-512-png-v2-dataset)","metadata":{}},{"cell_type":"markdown","source":"# Install Libraries","metadata":{}},{"cell_type":"markdown","source":"## `Clean up called ...` Issue\nCheck out [here](https://www.kaggle.com/discussions/questions-and-answers/326206#2056118)","metadata":{}},{"cell_type":"code","source":"!pip install tensorflow tensorflow_addons --upgrade --quiet\n!yes | apt install --allow-change-held-packages libcudnn8=8.1.0.77-1+cuda11.2","metadata":{"execution":{"iopub.status.busy":"2022-12-14T09:12:36.240493Z","iopub.execute_input":"2022-12-14T09:12:36.240933Z","iopub.status.idle":"2022-12-14T09:14:56.067218Z","shell.execute_reply.started":"2022-12-14T09:12:36.240848Z","shell.execute_reply":"2022-12-14T09:14:56.065999Z"},"_kg_hide-input":true,"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Libraries","metadata":{}},{"cell_type":"code","source":"!pip install -q efficientnet >> /dev/null\n!pip install -qU wandb\n!pip install -qU scikit-learn\n!pip install -q bbox-utility # source code: https://github.com/awsaf49/bbox","metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-14T09:15:23.392347Z","iopub.execute_input":"2022-12-14T09:15:23.392749Z","iopub.status.idle":"2022-12-14T09:16:06.593206Z","shell.execute_reply.started":"2022-12-14T09:15:23.392713Z","shell.execute_reply":"2022-12-14T09:16:06.592016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import Libraries","metadata":{}},{"cell_type":"code","source":"import os\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3'  # to avoid too many logging messages\nimport pandas as pd, numpy as np, random, shutil\nimport tensorflow as tf, re, math\nimport tensorflow.keras.backend as K\nimport sklearn\nimport matplotlib.pyplot as plt\nimport tensorflow_addons as tfa\nimport efficientnet.tfkeras as efn\nimport wandb\nimport yaml\nimport gc\ngc.collect()\n\nfrom IPython import display as ipd\nfrom glob import glob\nfrom tqdm.notebook import tqdm\nfrom sklearn.model_selection import KFold, StratifiedKFold, GroupKFold, StratifiedGroupKFold\nfrom sklearn.metrics import roc_auc_score","metadata":{"execution":{"iopub.status.busy":"2022-12-14T09:17:12.792449Z","iopub.execute_input":"2022-12-14T09:17:12.793323Z","iopub.status.idle":"2022-12-14T09:17:13.008946Z","shell.execute_reply.started":"2022-12-14T09:17:12.793283Z","shell.execute_reply":"2022-12-14T09:17:13.007627Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Version Check","metadata":{}},{"cell_type":"code","source":"print('np:', np.__version__)\nprint('pd:', pd.__version__)\nprint('sklearn:', sklearn.__version__)\nprint('tf:',tf.__version__)\nprint('tfa:', tfa.__version__)\nprint('w&b:', wandb.__version__)","metadata":{"execution":{"iopub.status.busy":"2022-12-14T09:17:14.645115Z","iopub.execute_input":"2022-12-14T09:17:14.645483Z","iopub.status.idle":"2022-12-14T09:17:15.021139Z","shell.execute_reply.started":"2022-12-14T09:17:14.64545Z","shell.execute_reply":"2022-12-14T09:17:15.019413Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Wandb","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://camo.githubusercontent.com/dd842f7b0be57140e68b2ab9cb007992acd131c48284eaf6b1aca758bfea358b/68747470733a2f2f692e696d6775722e636f6d2f52557469567a482e706e67\" width=\"400\" alt=\"Weights & Biases\" />\n\nWeights & Biases (W&B) is MLOps platform for tracking our experiemnts. We can use it to Build better models faster with experiment tracking, dataset versioning, and model management\n\nSome of the cool features of **W&B**:\n\n* Track, compare, and visualize ML experiments\n* Get live metrics, terminal logs, and system stats streamed to the centralized dashboard.\n* Explain how your model works, show graphs of how model versions improved, discuss bugs, and demonstrate progress towards milestones.\n\n> **Note:** `kaggle_secrets` has not been added to `TPU-1VM` yet, so run anonymously then go to the dashbaord and simply `claim` the run.","metadata":{}},{"cell_type":"code","source":"import wandb\n\ntry:\n    from kaggle_secrets import UserSecretsClient\n    user_secrets = UserSecretsClient()\n    api_key = user_secrets.get_secret(\"WANDB\")\n\n    wandb.login(key=api_key)\n    anonymous = None\nexcept:\n    anonymous = \"must\"\n    print('To use your W&B account,\\nGo to Add-ons -> Secrets and provide your W&B access token. Use the Label name as WANDB. \\nGet your W&B access token from here: https://wandb.ai/authorize')","metadata":{"execution":{"iopub.status.busy":"2022-12-14T09:17:26.112002Z","iopub.execute_input":"2022-12-14T09:17:26.112372Z","iopub.status.idle":"2022-12-14T09:17:29.801455Z","shell.execute_reply.started":"2022-12-14T09:17:26.11234Z","shell.execute_reply":"2022-12-14T09:17:29.800349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Configuration\nModify `conf_thr` and `mask_thr` to control bbox generation","metadata":{}},{"cell_type":"code","source":"class CFG:\n    wandb         = True\n    competition   = 'rsna-bcd-od' \n    _wandb_kernel = 'awsaf49'\n    debug         = False\n    exp_name      = 'efnb4-1024x512-cv=0.140' # name of the experiment, folds will be grouped using 'exp_name'\n    \n    # use verbose=0 for silent, vebose=1 for interactive,\n    verbose      = 1 if debug else 0\n    display_plot = True\n\n    # device\n    device = \"GPU\" #or \"GPU\"\n\n    model_name = 'EfficientNetB4'\n\n    # seed for data-split, layer init, augs\n    seed = 42\n\n    # number of folds for data-split\n    folds = 5\n    selected_folds = [0, 1, 2, 3, 4]\n\n    # size of the image\n    img_size = [1024, 512]\n\n    # batch_size and epochs\n    batch_size = 28\n\n    # augmentation\n    augment   = False\n\n    # test-time augs\n    tta = 1\n    \n    # threshold\n    conf_thr = 0.45  # which image to consider for bbox\n    mask_thr = 0.5  # to create mask from gradcam\n    \n    # target column\n    target_col  = ['cancer']","metadata":{"execution":{"iopub.status.busy":"2022-12-14T09:17:44.172402Z","iopub.execute_input":"2022-12-14T09:17:44.172804Z","iopub.status.idle":"2022-12-14T09:17:44.180579Z","shell.execute_reply.started":"2022-12-14T09:17:44.172769Z","shell.execute_reply":"2022-12-14T09:17:44.179422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Reproducibility","metadata":{}},{"cell_type":"code","source":"def seeding(SEED):\n    np.random.seed(SEED)\n    random.seed(SEED)\n    os.environ['PYTHONHASHSEED'] = str(SEED)\n#     os.environ['TF_CUDNN_DETERMINISTIC'] = str(SEED)\n    tf.random.set_seed(SEED)\n    print('seeding done!!!')\nseeding(CFG.seed)","metadata":{"execution":{"iopub.status.busy":"2022-12-14T09:17:45.455529Z","iopub.execute_input":"2022-12-14T09:17:45.455897Z","iopub.status.idle":"2022-12-14T09:17:45.462435Z","shell.execute_reply.started":"2022-12-14T09:17:45.455864Z","shell.execute_reply":"2022-12-14T09:17:45.461356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Device Configs\nThis notebook is compatible for **remote-tpu**, **local-tpu**, **multi-gpu** and **single-gpu**. Simple change to `device=\"TPU\"` for **remote-tpu** and `device=\"TPU-1VM\"` for **local-tpu** and finally, `device=\"GPU\"` for single or multi-gpu.","metadata":{}},{"cell_type":"code","source":"if \"TPU\" in CFG.device:\n    tpu = 'local' if CFG.device=='TPU-1VM' else None\n    print(\"connecting to TPU...\")\n    try:\n        tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=tpu)\n        strategy = tf.distribute.TPUStrategy(tpu)\n    except:\n        CFG.device = \"GPU\"\n        \nif CFG.device == \"GPU\"  or CFG.device==\"CPU\":\n    ngpu = len(tf.config.experimental.list_physical_devices('GPU'))\n    if ngpu>1:\n        print(\"Using multi GPU\")\n        strategy = tf.distribute.MirroredStrategy()\n    elif ngpu==1:\n        print(\"Using single GPU\")\n        strategy = tf.distribute.get_strategy()\n    else:\n        print(\"Using CPU\")\n        strategy = tf.distribute.get_strategy()\n        CFG.device = \"CPU\"\n\nif CFG.device == \"GPU\":\n    print(\"Num GPUs Available: \", ngpu)\n    \n\nAUTO     = tf.data.experimental.AUTOTUNE\nREPLICAS = strategy.num_replicas_in_sync\nprint(f'REPLICAS: {REPLICAS}')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-12-14T09:17:47.689345Z","iopub.execute_input":"2022-12-14T09:17:47.689725Z","iopub.status.idle":"2022-12-14T09:17:47.87424Z","shell.execute_reply.started":"2022-12-14T09:17:47.689691Z","shell.execute_reply":"2022-12-14T09:17:47.873065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Path for Images\n* Remote-TPU requires **GCS** path. Luckily Kaggle Provides that for us :)\n* Local-TPU aka TPU-1VM doesn't need GCS path =)","metadata":{}},{"cell_type":"code","source":"BASE_PATH = '/kaggle/input/rsna-bcd-roi-1024x512-png-v2-dataset'\nCKPT_PATH = '/kaggle/input/rsna-bcd-efnb4-1024x512-v2-ckpt-ds'\n\nif CFG.device==\"TPU\":\n    from kaggle_datasets import KaggleDatasets\n    GCS_PATH = KaggleDatasets().get_gcs_path(BASE_PATH.split('/')[-1])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-14T09:17:49.696948Z","iopub.execute_input":"2022-12-14T09:17:49.697348Z","iopub.status.idle":"2022-12-14T09:17:49.702905Z","shell.execute_reply.started":"2022-12-14T09:17:49.697315Z","shell.execute_reply":"2022-12-14T09:17:49.701677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Meta Data","metadata":{}},{"cell_type":"code","source":"# use gcs_path for remote-tpu\nif CFG.device==\"TPU\":\n    BASE_PATH = GCS_PATH\n\n# train\ndf = pd.read_csv(f'{BASE_PATH}/train.csv')\ndf['image_path'] = f'{BASE_PATH}/train_images'\\\n                    + '/' + df.patient_id.astype(str)\\\n                    + '/' + df.image_id.astype(str)\\\n                    + '.png'\nprint('Train:')\ndisplay(df.head(2))\n\n# test\ntest_df = pd.read_csv(f'{BASE_PATH}/test.csv')\ntest_df['image_path'] = f'{BASE_PATH}/test_images'\\\n                    + '/' + test_df.patient_id.astype(str)\\\n                    + '/' + test_df.image_id.astype(str)\\\n                    + '.png'\nprint('\\nTest:')\ndisplay(test_df.head(2))","metadata":{"execution":{"iopub.status.busy":"2022-12-14T09:17:50.875615Z","iopub.execute_input":"2022-12-14T09:17:50.876228Z","iopub.status.idle":"2022-12-14T09:17:51.200397Z","shell.execute_reply.started":"2022-12-14T09:17:50.87619Z","shell.execute_reply":"2022-12-14T09:17:51.199378Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Check If Data Exist?","metadata":{}},{"cell_type":"code","source":"tf.io.gfile.exists(df.image_path.iloc[0]), tf.io.gfile.exists(test_df.image_path.iloc[0])","metadata":{"execution":{"iopub.status.busy":"2022-12-14T09:17:52.471642Z","iopub.execute_input":"2022-12-14T09:17:52.472032Z","iopub.status.idle":"2022-12-14T09:17:52.489903Z","shell.execute_reply.started":"2022-12-14T09:17:52.471997Z","shell.execute_reply":"2022-12-14T09:17:52.488945Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train-Test Ditribution","metadata":{}},{"cell_type":"code","source":"print('train_files:',df.shape[0])\nprint('test_files:',test_df.shape[0])","metadata":{"execution":{"iopub.status.busy":"2022-12-14T09:17:53.598792Z","iopub.execute_input":"2022-12-14T09:17:53.599804Z","iopub.status.idle":"2022-12-14T09:17:53.606142Z","shell.execute_reply.started":"2022-12-14T09:17:53.599765Z","shell.execute_reply":"2022-12-14T09:17:53.604815Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Split\n* Data is splited while stratifying, 'laterality', 'view', 'age','cancer', 'biopsy', 'invasive', 'BIRADS', 'implant', 'density','machine_id', 'difficult_negative_case','cancer'.\n* To avoid leakage, data is also split keeping images from same patient in either train or valid not in both.\n* `StratifiedGroupKFold` does the both job **stratifying** and **group** split.\n\n","metadata":{}},{"cell_type":"code","source":"num_bins = 5\ndf[\"age_bin\"] = pd.cut(df['age'].values.reshape(-1), bins=num_bins, labels=False)\n\nstrat_cols = [\n    'laterality', 'view', 'biopsy','invasive', 'BIRADS', 'age_bin',\n    'implant', 'density','machine_id', 'difficult_negative_case',\n    'cancer',\n]\n\ndf['stratify'] = ''\nfor col in strat_cols:\n    df['stratify'] += df[col].astype(str)\n\nskf = StratifiedGroupKFold(n_splits=CFG.folds, shuffle=True, random_state=CFG.seed)\nfor fold, (train_idx, val_idx) in enumerate(skf.split(df, df['stratify'], df[\"patient_id\"])):\n    df.loc[val_idx, 'fold'] = fold\ndisplay(df.groupby(['fold', \"cancer\"]).size())","metadata":{"execution":{"iopub.status.busy":"2022-12-14T09:17:54.988197Z","iopub.execute_input":"2022-12-14T09:17:54.990282Z","iopub.status.idle":"2022-12-14T09:18:02.800598Z","shell.execute_reply.started":"2022-12-14T09:17:54.990233Z","shell.execute_reply":"2022-12-14T09:18:02.799589Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Pipeline\n* Reads the raw file and then decodes it to tf.Tensor\n* Resizes the image in desired size\n* Chages the datatype to **float32**\n* Caches the Data for boosting up the speed.\n* Uses Augmentations to reduce overfitting and make model more robust.\n* Finally, splits the data into batches.\n","metadata":{}},{"cell_type":"code","source":"def build_decoder(with_labels=True, target_size=CFG.img_size, ext='png'):\n    def decode(path):\n        file_bytes = tf.io.read_file(path)\n        if ext == 'png':\n            img = tf.image.decode_png(file_bytes, channels=3)\n        elif ext in ['jpg', 'jpeg']:\n            img = tf.image.decode_jpeg(file_bytes, channels=3)\n        else:\n            raise ValueError(\"Image extension not supported\")\n\n        img = tf.image.resize(img, target_size, method='bilinear')\n        img = tf.cast(img, tf.float32) / 255.0\n        img = tf.reshape(img, [*target_size, 3])\n\n        return img\n    \n    def decode_with_labels(path, label):\n        return decode(path), tf.cast(label, tf.float32)\n    \n    return decode_with_labels if with_labels else decode\n\n\ndef build_augmenter(with_labels=True, dim=CFG.img_size):\n    def augment(img, dim=dim):\n        img = tf.clip_by_value(img, 0, 1)  if CFG.clip else img         \n        img = tf.reshape(img, [*dim, 3])\n        return img\n    \n    def augment_with_labels(img, label):    \n        return augment(img), label\n    \n    return augment_with_labels if with_labels else augment\n\n\ndef build_dataset(paths, labels=None, batch_size=32, cache=True,\n                  decode_fn=None, augment_fn=None,\n                  augment=True, repeat=True, shuffle=1024, \n                  cache_dir=\"\", drop_remainder=False):\n    if cache_dir != \"\" and cache is True:\n        os.makedirs(cache_dir, exist_ok=True)\n    \n    if decode_fn is None:\n        decode_fn = build_decoder(labels is not None)\n    \n    if augment_fn is None:\n        augment_fn = build_augmenter(labels is not None)\n    \n    AUTO = tf.data.experimental.AUTOTUNE\n    slices = paths if labels is None else (paths, labels)\n    \n    ds = tf.data.Dataset.from_tensor_slices(slices)\n    ds = ds.map(decode_fn, num_parallel_calls=AUTO)\n    ds = ds.cache(cache_dir) if cache else ds\n    ds = ds.repeat() if repeat else ds\n    if shuffle: \n        ds = ds.shuffle(shuffle, seed=CFG.seed)\n        opt = tf.data.Options()\n        opt.experimental_deterministic = False\n        ds = ds.with_options(opt)\n    ds = ds.map(augment_fn, num_parallel_calls=AUTO) if augment else ds\n    ds = ds.batch(batch_size, drop_remainder=drop_remainder)\n    ds = ds.prefetch(AUTO)\n    return ds","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-14T09:18:02.803985Z","iopub.execute_input":"2022-12-14T09:18:02.804276Z","iopub.status.idle":"2022-12-14T09:18:02.819852Z","shell.execute_reply.started":"2022-12-14T09:18:02.804249Z","shell.execute_reply":"2022-12-14T09:18:02.818899Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualization\n* Check if augmentation is working properly or not.","metadata":{}},{"cell_type":"code","source":"def display_batch(batch, size=2):\n    if isinstance(batch, tuple):\n        imgs, tars = batch\n    else:\n        imgs = batch\n        tars = None\n    tars = tars.numpy().squeeze()\n    plt.figure(figsize=(size*4, 10))\n    for img_idx in range(size):\n        plt.subplot(1, size, img_idx+1)\n        if tars is not None:\n            plt.title(f'{CFG.target_col[0]}: {tars[img_idx]:0.3f}', fontsize=10)\n        plt.imshow(imgs[img_idx,:, :, :])\n        plt.xticks([]); plt.yticks([])\n    plt.tight_layout()\n    plt.show() ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-14T09:18:02.821483Z","iopub.execute_input":"2022-12-14T09:18:02.821849Z","iopub.status.idle":"2022-12-14T09:18:02.832424Z","shell.execute_reply.started":"2022-12-14T09:18:02.821813Z","shell.execute_reply":"2022-12-14T09:18:02.8314Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Generate Image","metadata":{}},{"cell_type":"code","source":"fold = 0\nfold_df = df.groupby('cancer').head(16)\npaths  = fold_df.image_path.tolist()\nlabels = fold_df[CFG.target_col].values\nds = build_dataset(paths, labels, cache=False, batch_size=32,\n                   repeat=True, shuffle=True, augment=False)\nds = ds.unbatch().batch(20)\nbatch = next(iter(ds))\ndisplay_batch(batch, 5);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-14T09:18:02.836223Z","iopub.execute_input":"2022-12-14T09:18:02.836497Z","iopub.status.idle":"2022-12-14T09:18:06.28734Z","shell.execute_reply.started":"2022-12-14T09:18:02.836471Z","shell.execute_reply":"2022-12-14T09:18:06.286137Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# GradCAM Utils","metadata":{}},{"cell_type":"code","source":"import matplotlib.cm as cm, cv2\nfrom tensorflow import keras\n\ndef gen_gradcam_heatmap(img, model, last_conv='top_conv', pred_index=0):\n    # First, we create a model that maps the input image to the activations\n    # of the last conv layer as well as the output predictions\n    grad_model = tf.keras.models.Model(\n        [model.inputs], [model.get_layer(last_conv).output, model.output]\n    )\n\n    # Then, we compute the gradient of the top predicted class for our input image\n    # with respect to the activations of the last conv layer\n    with tf.GradientTape() as tape:\n        features, preds = grad_model(img)\n        if pred_index is None:\n            pred_index = tf.argmax(preds[0])\n        class_channel = preds[:, pred_index]\n\n    # This is the gradient of the output neuron (top predicted or chosen)\n    # with regard to the output feature map of the last conv layer\n    grads = tape.gradient(class_channel, features)\n\n    # This is a vector where each entry is the mean intensity of the gradient\n    # over a specific feature map channel\n    pooled_grads = tf.reduce_mean(grads, axis=(0, 1, 2))\n\n    # We multiply each channel in the feature map array\n    # by \"how important this channel is\" with regard to the top predicted class\n    # then sum all the channels to obtain the heatmap class activation\n    features = features[0]\n    heatmap = features @ pooled_grads[..., tf.newaxis]\n    heatmap = tf.squeeze(heatmap)\n\n    # For visualization purpose, we will also normalize the heatmap between 0 & 1\n    heatmap = tf.maximum(heatmap, 0) / tf.math.reduce_max(heatmap)\n    return heatmap.numpy(), preds[0]\n\n\ndef get_gradcam(img, model):\n\n    heatmap, pred = gen_gradcam_heatmap(img, model, last_conv='top_conv', pred_index=0)\n    heatmap = cv2.resize(heatmap, dsize=(img.shape[2], img.shape[1]))\n        \n    return heatmap, pred\n\ndef gray2rgb(img, gradcam, alpha=0.55, cmap='inferno'):\n    \"\"\"Convert grayscale heatmap to jet gradcam\"\"\"\n    # Use jet colormap to colorize heatmap\n    cmap = cm.get_cmap(cmap)\n\n    # Use RGB values of the colormap\n    colors  = cmap(np.arange(256))[:, :3]\n    heatmap = colors[gradcam]\n\n    # Superimpose the heatmap on original image\n    sup_img = heatmap * alpha + (img/255.0)\n    sup_img = tf.keras.preprocessing.image.array_to_img(sup_img) # avoid saturated colors\n    return np.asarray(sup_img)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-14T09:18:10.652866Z","iopub.execute_input":"2022-12-14T09:18:10.65341Z","iopub.status.idle":"2022-12-14T09:18:10.815307Z","shell.execute_reply.started":"2022-12-14T09:18:10.653369Z","shell.execute_reply":"2022-12-14T09:18:10.814326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# BBox Utils","metadata":{}},{"cell_type":"code","source":"import cv2\nfrom bbox.utils import draw_bboxes\nimport matplotlib.cm as cm\n\ndef gradcam2bbox(gradcam, thr=0.6):\n    # Binarize the image\n    mask = (gradcam>thr).astype('uint8')\n\n    # Make contours around the binarized image, keep only the largest contour\n    contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_NONE)\n    contour = max(contours, key=cv2.contourArea)\n\n    # Find ROI from largest contour\n    ys = contour.squeeze()[:, 0]\n    xs = contour.squeeze()[:, 1]\n    \n    y1 = np.min(xs); y2 = np.max(xs);\n    x1 = np.min(ys); x2 = np.max(ys);\n    \n    return [x1, y1, x2, y2]\n\nnp.random.seed(32)\ncolors = [(np.random.randint(255), np.random.randint(255), np.random.randint(255))\\\n          for idx in range(1)]\nfont_param = {\"fontname\":\"Times New Roman\",\"fontweight\":\"bold\", \"fontsize\":16}","metadata":{"execution":{"iopub.status.busy":"2022-12-14T09:18:12.521899Z","iopub.execute_input":"2022-12-14T09:18:12.522303Z","iopub.status.idle":"2022-12-14T09:18:13.227079Z","shell.execute_reply.started":"2022-12-14T09:18:12.522267Z","shell.execute_reply":"2022-12-14T09:18:13.226099Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Wandb** Logger\nLog:\n* Image\n* Mask\n* BBox","metadata":{}},{"cell_type":"code","source":"class_labels = {\n    1: \"cancer\",\n}\n\ndef wandb_init(fold):\n    config = {k:v for k,v in dict(vars(CFG)).items() if '__' not in k}\n    config.update({\"fold\":int(fold)})\n    run    = wandb.init(project=\"rsna-bcd-gradcam2bbox\",\n               name=f\"fold-{fold}|dim-{CFG.img_size[0]}x{CFG.img_size[1]}|model-{CFG.model_name}\",\n               config=config,\n               anonymous=anonymous,\n                    )\n    return run\n\ndef log_wandb(fold, data):\n    # Create wandb table with gradcam, bbox, mask\n    wandb_data = []\n    for idx, row in enumerate(tqdm(data, desc='wandb ')):\n        tab = row[0]\n        img = row[1]\n        gradcam = gray2rgb(img, row[2])\n        mask = (row[2]>(CFG.mask_thr*255.0)).astype('uint8')\n        bbox = tab[-2][0]\n        \n        # Add row to table\n        wandb_data+=[[*tab,\n                wandb.Image(img), \n                wandb.Image(gradcam),\n                wandb.Image(img, masks={\n                    \"ground_truth\": {\"mask_data\": mask,\n                                     \"class_labels\": class_labels,}\n                }),\n                wandb.Image(img, boxes={\"ground_truth\": {\"box_data\": [{\"position\": {\"minX\": int(bbox[0]),\n                                                                                    \"maxX\": int(bbox[2]),\n                                                                                    \"minY\": int(bbox[1]),\n                                                                                    \"maxY\": int(bbox[3])},\n                                                                       \"class_id\" : 1,\n                                                                       \"box_caption\": \"cancer\",\n                                                                       \"domain\" : \"pixel\",}],\n                                                         \"class_labels\": class_labels,},\n                                       })\n               ]]\n        \n    # Create Table\n    wandb_table = wandb.Table(data=wandb_data,\n                              columns=[*cols, 'image', 'gradcam','mask_img','bbox_img'])\n    \n    # Log data to wandb\n    wandb.log({'viz_table':wandb_table,})","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-14T09:36:57.620009Z","iopub.execute_input":"2022-12-14T09:36:57.620389Z","iopub.status.idle":"2022-12-14T09:36:57.632444Z","shell.execute_reply.started":"2022-12-14T09:36:57.620357Z","shell.execute_reply":"2022-12-14T09:36:57.631454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Generate BBox from GradCAM","metadata":{}},{"cell_type":"code","source":"oof_pred = []; oof_tar = []; oof_val = []; oof_ids = []; oof_folds = []\npreds = np.zeros((test_df.shape[0],1))\n\nfor fold in np.arange(CFG.folds):\n    \n    # Ignore not selected folds\n    if fold not in CFG.selected_folds:\n        continue\n        \n    # Init wandb\n    if CFG.wandb:\n        run = wandb_init(fold)\n\n    # Train and valid dataframe\n    valid_df = df.query(\"fold==@fold\")\n    \n    # Get image_paths and labels\n    valid_paths = valid_df.image_path.values; valid_labels = valid_df[CFG.target_col].values.astype(np.float32)\n    \n    # Min samples in debug mode\n    min_samples = CFG.batch_size*REPLICAS*2\n    \n    # For debug model run on small portion\n    if CFG.debug:\n        valid_paths = valid_paths[:min_samples]; valid_labels = valid_labels[:min_samples]\n    \n    # Sample count\n    num_valid = len(valid_paths)\n    num_valid_pos = (valid_labels==1).sum()\n    \n    # Show message\n    print('#'*40); print('#### FOLD: ',fold)\n    print('#### NUM_SAMPLES: {:,} | NUM_CANCER: {:,}'.format(num_valid, num_valid_pos))\n    print('#'*40)\n    \n    # Build model\n    print('Loading model...')\n    K.clear_session()\n    with strategy.scope():\n        model = tf.keras.models.load_model(f'{CKPT_PATH}/fold-%i.h5'%fold,\n                                          compile=False)\n    \n    # Celect top and worst 10 cases for each class\n    gradcam_df  = valid_df.query(\"cancer==1\").reset_index(drop=True)\n    gradcam_ds  = build_dataset(gradcam_df.image_path.values, labels=None, cache=False, batch_size=1,\n                   repeat=False, shuffle=False, augment=False)\n    \n    # Columns\n    noimg_cols  = ['image_id','patient_id','cancer']\n    cols = noimg_cols + ['bbox','pred']\n    \n    # Create wandb table for upload\n    data = []\n    for idx, img in enumerate(tqdm(gradcam_ds, total=len(gradcam_df), desc='gradcam ')):\n        gradcam, pred = get_gradcam(img, model)\n        bbox = [gradcam2bbox(gradcam, thr=CFG.mask_thr)]\n        tab = gradcam_df.loc[idx, noimg_cols].tolist()+[bbox, pred[0].numpy()]\n        img = img.numpy()[0]\n        img = (img*255.0).astype('uint8')\n        gradcam = (gradcam*255.0).astype('uint8')\n        if pred>=CFG.conf_thr:\n            data+=[[tab, img, gradcam]]\n    \n    # Save bbox data for each fold\n    bbox_df = pd.DataFrame([x[0] for x in data], columns=cols)\n    bbox_df.to_csv(f'bbox{fold}.csv',index=False)\n            \n    # Log bbox, mask, gradcam on wandb & plot\n    if CFG.wandb:\n        log_wandb(fold, data) # log\n        wandb.run.finish() # finish the run\n        display(ipd.IFrame(run.url, width=1080, height=720)) # show wandb dashboard\n        \n    # Clean unless last fold (last fold for local plotting)\n    if fold!=CFG.selected_folds[-1]:\n        del data; gc.collect()\n        \n    print();","metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"scrolled":true,"execution":{"iopub.status.busy":"2022-12-14T09:26:05.750786Z","iopub.execute_input":"2022-12-14T09:26:05.751185Z","iopub.status.idle":"2022-12-14T09:36:05.276498Z","shell.execute_reply.started":"2022-12-14T09:26:05.75114Z","shell.execute_reply":"2022-12-14T09:36:05.274025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualize Gradcam & BBox","metadata":{}},{"cell_type":"code","source":"for idx in range(5):\n    row = data[idx]\n    img = row[1].copy()\n    gradcam = gray2rgb(img, row[2])\n    mask = (row[2]>(CFG.mask_thr*255.0)).astype('uint8')\n    \n    tab = row[0]\n    iid = tab[0]; pid = tab[1]; view = tab[-4]; lat = tab[-5]; pred=tab[-1]\n    bboxes = tab[-2]  # voc pascal style\n    \n    plt.figure(figsize=(20,10))\n    print(\"#\"*20)\n    print(f\"ImageID: {iid}\\nPatientID: {pid}\\nView: {view}\\nLaterality: {lat}\")\n    print(f\"Pred: {pred:0.3f}\")\n    print(\"#\"*20)\n\n    # Plot image\n    plt.subplot(141)\n    plt.imshow(img)\n    plt.title(\"Image\", **font_param); plt.axis('off')\n\n    # Plot GradCAM\n    plt.subplot(142)\n    plt.imshow(gradcam)\n    plt.title(\"GradCAM\", **font_param); plt.axis('off')\n\n    # Plot Mask\n    plt.subplot(143)\n    plt.imshow(img)\n    plt.imshow(np.ma.masked_where(mask==0, mask), alpha=0.3, cmap='autumn')\n    plt.title(\"Mask\", **font_param); plt.axis('off')\n\n    # Plot BBox\n    plt.subplot(144)\n    labels = [0]*1\n    names = ['cancer']*1\n    plt.imshow(draw_bboxes(img = img,\n                               bboxes = bboxes, \n                               classes = names,\n                               class_ids = labels,\n                               class_name = True, \n                               colors = colors, \n                               bbox_format = 'voc',\n                               line_thickness = 4))\n    plt.title(\"BBox\", **font_param); plt.axis('off')\n\n    plt.tight_layout()\n    plt.show()\n    print()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-14T09:37:03.53984Z","iopub.execute_input":"2022-12-14T09:37:03.540548Z","iopub.status.idle":"2022-12-14T09:37:08.304589Z","shell.execute_reply.started":"2022-12-14T09:37:03.540509Z","shell.execute_reply":"2022-12-14T09:37:08.303654Z"},"jupyter":{"source_hidden":true,"outputs_hidden":true},"collapsed":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Store BBox","metadata":{}},{"cell_type":"code","source":"bbox_df = pd.concat([pd.read_csv(f'bbox{fold}.csv') for fold in CFG.selected_folds], axis=0, ignore_index=True)\nprint(f'BBox Size: {len(bbox_df)}')\nbbox_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-14T09:37:52.072858Z","iopub.execute_input":"2022-12-14T09:37:52.073575Z","iopub.status.idle":"2022-12-14T09:37:52.094189Z","shell.execute_reply.started":"2022-12-14T09:37:52.073538Z","shell.execute_reply":"2022-12-14T09:37:52.09311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Remove Files","metadata":{}},{"cell_type":"code","source":"!rm -r /kaggle/working/wandb","metadata":{"execution":{"iopub.status.busy":"2022-12-14T06:48:46.893463Z","iopub.status.idle":"2022-12-14T06:48:46.894178Z","shell.execute_reply.started":"2022-12-14T06:48:46.893919Z","shell.execute_reply":"2022-12-14T06:48:46.893943Z"},"trusted":true},"execution_count":null,"outputs":[]}]}