{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# [RSNA Screening Mammography Breast Cancer Detection](https://www.kaggle.com/c/petfinder-pawpularity-score)\n> Find breast cancers in screening mammograms\n\n![](https://storage.googleapis.com/kaggle-competitions/kaggle/39272/logos/header.png?t=2022-11-28-17-29-35)","metadata":{}},{"cell_type":"markdown","source":"# Idea:\n* In this notebook will generate **bounding box (bbox)** from **gradcam** images.\n* BBoxes can be then used with Object Detection models for **Cancer Detection**\n* **Wandb** is integrated hence we can use this notebook for visual analysis.\n* **Image**, **Mask**, **GradCAM** & **BBox** is logged in **WandB**","metadata":{}},{"cell_type":"markdown","source":"# Notebooks\n* Only Image:\n    * ROI:\n        * train: [RSNA-BCD: EfficientNet [TF][TPU-1VM][Train]](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train)\n        * infer: [RSNA-BCD: EfficientNet [TF][TPU-1VM][Infer]](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-infer)\n    * NoROI + KerasCV: \n        * train: [RSNA-BCD: NoROI KerasCV [TF][Train]](https://www.kaggle.com/awsaf49/rsna-bcd-noroi-kerascv-tf-train/)\n        * infer: [RSNA-BCD: NoROI KerasCV [TF][Infer]](https://www.kaggle.com/awsaf49/rsna-bcd-noroi-kerascv-tf-infer/)\n* Dataset:\n    * ROI:\n        * [RSNA-BCD: ROI 1024x PNG Dataset](https://www.kaggle.com/datasets/awsaf49/rsna-bcd-roi-1024x-png-dataset)\n    * NoROI:\n        * [RSNA-BCD: 512 PNG v2 PNG Dataset](https://www.kaggle.com/datasets/awsaf49/rsnabcd-512-png-v2-dataset)","metadata":{}},{"cell_type":"markdown","source":"# Install Libraries","metadata":{}},{"cell_type":"markdown","source":"## `Clean up called ...` Issue\nCheck out [here](https://www.kaggle.com/discussions/questions-and-answers/326206#2056118)","metadata":{}},{"cell_type":"code","source":"!pip install tensorflow tensorflow_addons --upgrade --quiet\n!yes | apt install --allow-change-held-packages libcudnn8=8.1.0.77-1+cuda11.2","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-02-11T08:00:27.153224Z","iopub.execute_input":"2023-02-11T08:00:27.153877Z","iopub.status.idle":"2023-02-11T08:02:45.901995Z","shell.execute_reply.started":"2023-02-11T08:00:27.153786Z","shell.execute_reply":"2023-02-11T08:02:45.900825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Libraries","metadata":{}},{"cell_type":"code","source":"!pip install -q efficientnet >> /dev/null\n!pip install -qU wandb\n!pip install -qU scikit-learn\n!pip install -q bbox-utility # source code: https://github.com/awsaf49/bbox","metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:02:52.271422Z","iopub.execute_input":"2023-02-11T08:02:52.271861Z","iopub.status.idle":"2023-02-11T08:03:38.514963Z","shell.execute_reply.started":"2023-02-11T08:02:52.271822Z","shell.execute_reply":"2023-02-11T08:03:38.513606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import Libraries","metadata":{}},{"cell_type":"code","source":"import os\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3'  # to avoid too many logging messages\nimport pandas as pd, numpy as np, random, shutil\nimport tensorflow as tf, re, math\nimport tensorflow.keras.backend as K\nimport sklearn\nimport matplotlib.pyplot as plt\nimport tensorflow_addons as tfa\nimport efficientnet.tfkeras as efn\nimport wandb\nimport yaml\nimport gc\ngc.collect()\n\nfrom IPython import display as ipd\nfrom glob import glob\nfrom tqdm.notebook import tqdm\nfrom sklearn.model_selection import KFold, StratifiedKFold, GroupKFold, StratifiedGroupKFold\nfrom sklearn.metrics import roc_auc_score","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:03:43.074433Z","iopub.execute_input":"2023-02-11T08:03:43.075474Z","iopub.status.idle":"2023-02-11T08:03:48.451874Z","shell.execute_reply.started":"2023-02-11T08:03:43.075434Z","shell.execute_reply":"2023-02-11T08:03:48.450567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Version Check","metadata":{}},{"cell_type":"code","source":"print('np:', np.__version__)\nprint('pd:', pd.__version__)\nprint('sklearn:', sklearn.__version__)\nprint('tf:',tf.__version__)\nprint('tfa:', tfa.__version__)\nprint('w&b:', wandb.__version__)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:05:42.045995Z","iopub.execute_input":"2023-02-11T08:05:42.04644Z","iopub.status.idle":"2023-02-11T08:05:42.053185Z","shell.execute_reply.started":"2023-02-11T08:05:42.046407Z","shell.execute_reply":"2023-02-11T08:05:42.052101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Wandb","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://camo.githubusercontent.com/dd842f7b0be57140e68b2ab9cb007992acd131c48284eaf6b1aca758bfea358b/68747470733a2f2f692e696d6775722e636f6d2f52557469567a482e706e67\" width=\"400\" alt=\"Weights & Biases\" />\n\nWeights & Biases (W&B) is MLOps platform for tracking our experiemnts. We can use it to Build better models faster with experiment tracking, dataset versioning, and model management\n\nSome of the cool features of **W&B**:\n\n* Track, compare, and visualize ML experiments\n* Get live metrics, terminal logs, and system stats streamed to the centralized dashboard.\n* Explain how your model works, show graphs of how model versions improved, discuss bugs, and demonstrate progress towards milestones.\n\n> **Note:** `kaggle_secrets` has not been added to `TPU-1VM` yet, so run anonymously then go to the dashbaord and simply `claim` the run.","metadata":{}},{"cell_type":"code","source":"import wandb\n\ntry:\n    from kaggle_secrets import UserSecretsClient\n    user_secrets = UserSecretsClient()\n#     api_key = user_secrets.get_secret(\"WANDB\")\n    api_key='546c7dbc247f4a0d9c125ac58f90b8afa1252861'\n\n    wandb.login(key=api_key)\n    anonymous = None\nexcept:\n    anonymous = \"must\"\n    print('To use your W&B account,\\nGo to Add-ons -> Secrets and provide your W&B access token. Use the Label name as WANDB. \\nGet your W&B access token from here: https://wandb.ai/authorize')","metadata":{"execution":{"iopub.status.busy":"2023-02-11T08:07:30.504754Z","iopub.execute_input":"2023-02-11T08:07:30.505406Z","iopub.status.idle":"2023-02-11T08:07:33.004836Z","shell.execute_reply.started":"2023-02-11T08:07:30.505363Z","shell.execute_reply":"2023-02-11T08:07:33.003633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Configuration\nModify `conf_thr` and `mask_thr` to control bbox generation","metadata":{}},{"cell_type":"code","source":"class CFG:\n    wandb         = True\n    competition   = 'rsna-bcd-od' \n    _wandb_kernel = 'awsaf49'\n    debug         = False\n    exp_name      = 'efnb4-1024x512-cv=0.140' # name of the experiment, folds will be grouped using 'exp_name'\n    \n    # use verbose=0 for silent, vebose=1 for interactive,\n    verbose      = 1 if debug else 0\n    display_plot = True\n\n    # device\n    device = \"GPU\" #or \"GPU\"\n\n    model_name = 'EfficientNetB4'\n\n    # seed for data-split, layer init, augs\n    seed = 42\n\n    # number of folds for data-split\n    folds = 5\n    selected_folds = [0, 1, 2, 3, 4]\n\n    # size of the image\n    img_size = [1024, 512]\n\n    # batch_size and epochs\n    batch_size = 28\n\n    # augmentation\n    augment   = False\n\n    # test-time augs\n    tta = 1\n    \n    # threshold\n    conf_thr = 0.45  # which image to consider for bbox\n    mask_thr = 0.5  # to create mask from gradcam\n    \n    # target column\n    target_col  = ['cancer']","metadata":{"execution":{"iopub.status.busy":"2023-02-11T08:07:41.839211Z","iopub.execute_input":"2023-02-11T08:07:41.839681Z","iopub.status.idle":"2023-02-11T08:07:41.848685Z","shell.execute_reply.started":"2023-02-11T08:07:41.839641Z","shell.execute_reply":"2023-02-11T08:07:41.847366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Reproducibility","metadata":{}},{"cell_type":"code","source":"def seeding(SEED):\n    np.random.seed(SEED)\n    random.seed(SEED)\n    os.environ['PYTHONHASHSEED'] = str(SEED)\n#     os.environ['TF_CUDNN_DETERMINISTIC'] = str(SEED)\n    tf.random.set_seed(SEED)\n    print('seeding done!!!')\nseeding(CFG.seed)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T08:07:48.162712Z","iopub.execute_input":"2023-02-11T08:07:48.163399Z","iopub.status.idle":"2023-02-11T08:07:48.170119Z","shell.execute_reply.started":"2023-02-11T08:07:48.163364Z","shell.execute_reply":"2023-02-11T08:07:48.168867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Device Configs\nThis notebook is compatible for **remote-tpu**, **local-tpu**, **multi-gpu** and **single-gpu**. Simple change to `device=\"TPU\"` for **remote-tpu** and `device=\"TPU-1VM\"` for **local-tpu** and finally, `device=\"GPU\"` for single or multi-gpu.","metadata":{}},{"cell_type":"code","source":"if \"TPU\" in CFG.device:\n    tpu = 'local' if CFG.device=='TPU-1VM' else None\n    print(\"connecting to TPU...\")\n    try:\n        tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=tpu)\n        strategy = tf.distribute.TPUStrategy(tpu)\n    except:\n        CFG.device = \"GPU\"\n        \nif CFG.device == \"GPU\"  or CFG.device==\"CPU\":\n    ngpu = len(tf.config.experimental.list_physical_devices('GPU'))\n    if ngpu>1:\n        print(\"Using multi GPU\")\n        strategy = tf.distribute.MirroredStrategy()\n    elif ngpu==1:\n        print(\"Using single GPU\")\n        strategy = tf.distribute.get_strategy()\n    else:\n        print(\"Using CPU\")\n        strategy = tf.distribute.get_strategy()\n        CFG.device = \"CPU\"\n\nif CFG.device == \"GPU\":\n    print(\"Num GPUs Available: \", ngpu)\n    \n\nAUTO     = tf.data.experimental.AUTOTUNE\nREPLICAS = strategy.num_replicas_in_sync\nprint(f'REPLICAS: {REPLICAS}')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-02-11T08:07:51.318228Z","iopub.execute_input":"2023-02-11T08:07:51.318659Z","iopub.status.idle":"2023-02-11T08:07:51.512535Z","shell.execute_reply.started":"2023-02-11T08:07:51.318617Z","shell.execute_reply":"2023-02-11T08:07:51.511037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Path for Images\n* Remote-TPU requires **GCS** path. Luckily Kaggle Provides that for us :)\n* Local-TPU aka TPU-1VM doesn't need GCS path =)","metadata":{}},{"cell_type":"code","source":"BASE_PATH = '/kaggle/input/rsna-bcd-roi-1024x512-png-v2-dataset'\nCKPT_PATH = '/kaggle/input/rsna-bcd-efnb4-1024x512-v2-ckpt-ds'\n\nif CFG.device==\"TPU\":\n    from kaggle_datasets import KaggleDatasets\n    GCS_PATH = KaggleDatasets().get_gcs_path(BASE_PATH.split('/')[-1])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:07:58.074103Z","iopub.execute_input":"2023-02-11T08:07:58.074523Z","iopub.status.idle":"2023-02-11T08:07:58.081414Z","shell.execute_reply.started":"2023-02-11T08:07:58.074489Z","shell.execute_reply":"2023-02-11T08:07:58.080119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Meta Data","metadata":{}},{"cell_type":"code","source":"# use gcs_path for remote-tpu\nif CFG.device==\"TPU\":\n    BASE_PATH = GCS_PATH\n\n# train\ndf = pd.read_csv(f'{BASE_PATH}/train.csv')\ndf['image_path'] = f'{BASE_PATH}/train_images'\\\n                    + '/' + df.patient_id.astype(str)\\\n                    + '/' + df.image_id.astype(str)\\\n                    + '.png'\nprint('Train:')\ndisplay(df.head(2))\n\n# test\ntest_df = pd.read_csv(f'{BASE_PATH}/test.csv')\ntest_df['image_path'] = f'{BASE_PATH}/test_images'\\\n                    + '/' + test_df.patient_id.astype(str)\\\n                    + '/' + test_df.image_id.astype(str)\\\n                    + '.png'\nprint('\\nTest:')\ndisplay(test_df.head(2))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:08:00.511541Z","iopub.execute_input":"2023-02-11T08:08:00.512452Z","iopub.status.idle":"2023-02-11T08:08:01.023459Z","shell.execute_reply.started":"2023-02-11T08:08:00.512409Z","shell.execute_reply":"2023-02-11T08:08:01.022347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nfrom PIL import Image\n# Read a PIL image  \nimg = Image.open(df['image_path'][0])\nprint(img)\n# Convert the PIL image to Tensor\nimg_to_tensor = tf.convert_to_tensor(img)\n# print the converted Torch tensor\nprint(img_to_tensor)\nprint(\"dtype of tensor:\",img_to_tensor.shape)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-11T10:23:34.712658Z","iopub.execute_input":"2023-02-11T10:23:34.713156Z","iopub.status.idle":"2023-02-11T10:23:34.738709Z","shell.execute_reply.started":"2023-02-11T10:23:34.713103Z","shell.execute_reply":"2023-02-11T10:23:34.737639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Check If Data Exist?","metadata":{}},{"cell_type":"code","source":"tf.io.gfile.exists(df.image_path.iloc[0]), tf.io.gfile.exists(test_df.image_path.iloc[0])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:08:04.510372Z","iopub.execute_input":"2023-02-11T08:08:04.510757Z","iopub.status.idle":"2023-02-11T08:08:04.530131Z","shell.execute_reply.started":"2023-02-11T08:08:04.510726Z","shell.execute_reply":"2023-02-11T08:08:04.529242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train-Test Ditribution","metadata":{}},{"cell_type":"code","source":"print('train_files:',df.shape[0])\nprint('test_files:',test_df.shape[0])","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:08:06.734245Z","iopub.execute_input":"2023-02-11T08:08:06.734656Z","iopub.status.idle":"2023-02-11T08:08:06.741019Z","shell.execute_reply.started":"2023-02-11T08:08:06.734625Z","shell.execute_reply":"2023-02-11T08:08:06.739856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Split\n* Data is splited while stratifying, 'laterality', 'view', 'age','cancer', 'biopsy', 'invasive', 'BIRADS', 'implant', 'density','machine_id', 'difficult_negative_case','cancer'.\n* To avoid leakage, data is also split keeping images from same patient in either train or valid not in both.\n* `StratifiedGroupKFold` does the both job **stratifying** and **group** split.\n\n","metadata":{}},{"cell_type":"code","source":"num_bins = 5\ndf[\"age_bin\"] = pd.cut(df['age'].values.reshape(-1), bins=num_bins, labels=False)\n\nstrat_cols = [\n    'laterality', 'view', 'biopsy','invasive', 'BIRADS', 'age_bin',\n    'implant', 'density','machine_id', 'difficult_negative_case',\n    'cancer',\n]\n\ndf['stratify'] = ''\nfor col in strat_cols:\n    df['stratify'] += df[col].astype(str)\n\nskf = StratifiedGroupKFold(n_splits=CFG.folds, shuffle=True, random_state=CFG.seed)\nfor fold, (train_idx, val_idx) in enumerate(skf.split(df, df['stratify'], df[\"patient_id\"])):\n    df.loc[val_idx, 'fold'] = fold\ndisplay(df.groupby(['fold', \"cancer\"]).size())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:08:09.410523Z","iopub.execute_input":"2023-02-11T08:08:09.411256Z","iopub.status.idle":"2023-02-11T08:08:19.43561Z","shell.execute_reply.started":"2023-02-11T08:08:09.41122Z","shell.execute_reply":"2023-02-11T08:08:19.434141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Pipeline\n* Reads the raw file and then decodes it to tf.Tensor\n* Resizes the image in desired size\n* Chages the datatype to **float32**\n* Caches the Data for boosting up the speed.\n* Uses Augmentations to reduce overfitting and make model more robust.\n* Finally, splits the data into batches.\n","metadata":{}},{"cell_type":"code","source":"def build_decoder(with_labels=True, target_size=CFG.img_size, ext='png'):\n    def decode(path):\n        file_bytes = tf.io.read_file(path)\n        if ext == 'png':\n            img = tf.image.decode_png(file_bytes, channels=3)\n        elif ext in ['jpg', 'jpeg']:\n            img = tf.image.decode_jpeg(file_bytes, channels=3)\n        else:\n            raise ValueError(\"Image extension not supported\")\n\n        img = tf.image.resize(img, target_size, method='bilinear')\n        img = tf.cast(img, tf.float32) / 255.0\n        img = tf.reshape(img, [*target_size, 3])\n\n        return img\n    \n    def decode_with_labels(path, label):\n        return decode(path), tf.cast(label, tf.float32)\n    \n    return decode_with_labels if with_labels else decode\n\n\ndef build_augmenter(with_labels=True, dim=CFG.img_size):\n    def augment(img, dim=dim):\n        img = tf.clip_by_value(img, 0, 1)  if CFG.clip else img         \n        img = tf.reshape(img, [*dim, 3])\n        return img\n    \n    def augment_with_labels(img, label):    \n        return augment(img), label\n    \n    return augment_with_labels if with_labels else augment\n\n\ndef build_dataset(paths, labels=None, batch_size=32, cache=True,\n                  decode_fn=None, augment_fn=None,\n                  augment=True, repeat=True, shuffle=1024, \n                  cache_dir=\"\", drop_remainder=False):\n    if cache_dir != \"\" and cache is True:\n        os.makedirs(cache_dir, exist_ok=True)\n    \n    if decode_fn is None:\n        decode_fn = build_decoder(labels is not None)\n    \n    if augment_fn is None:\n        augment_fn = build_augmenter(labels is not None)\n    \n    AUTO = tf.data.experimental.AUTOTUNE\n    slices = paths if labels is None else (paths, labels)\n    \n    ds = tf.data.Dataset.from_tensor_slices(slices)\n    ds = ds.map(decode_fn, num_parallel_calls=AUTO)\n    ds = ds.cache(cache_dir) if cache else ds\n    ds = ds.repeat() if repeat else ds\n    if shuffle: \n        ds = ds.shuffle(shuffle, seed=CFG.seed)\n        opt = tf.data.Options()\n        opt.experimental_deterministic = False\n        ds = ds.with_options(opt)\n    ds = ds.map(augment_fn, num_parallel_calls=AUTO) if augment else ds\n    ds = ds.batch(batch_size, drop_remainder=drop_remainder)\n    ds = ds.prefetch(AUTO)\n    return ds","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:08:28.897373Z","iopub.execute_input":"2023-02-11T08:08:28.897772Z","iopub.status.idle":"2023-02-11T08:08:28.912996Z","shell.execute_reply.started":"2023-02-11T08:08:28.897739Z","shell.execute_reply":"2023-02-11T08:08:28.911828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualization\n* Check if augmentation is working properly or not.","metadata":{}},{"cell_type":"code","source":"def display_batch(batch, size=2):\n    if isinstance(batch, tuple):\n        imgs, tars = batch\n    else:\n        imgs = batch\n        tars = None\n    tars = tars.numpy().squeeze()\n    plt.figure(figsize=(size*4, 10))\n    for img_idx in range(size):\n        plt.subplot(1, size, img_idx+1)\n        if tars is not None:\n            plt.title(f'{CFG.target_col[0]}: {tars[img_idx]:0.3f}', fontsize=10)\n        plt.imshow(imgs[img_idx,:, :, :])\n        plt.xticks([]); plt.yticks([])\n    plt.tight_layout()\n    plt.show() ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:08:36.249236Z","iopub.execute_input":"2023-02-11T08:08:36.249647Z","iopub.status.idle":"2023-02-11T08:08:36.259069Z","shell.execute_reply.started":"2023-02-11T08:08:36.249575Z","shell.execute_reply":"2023-02-11T08:08:36.257901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Generate Image","metadata":{}},{"cell_type":"code","source":"fold = 0\nfold_df = df.groupby('cancer').head(16)\npaths  = fold_df.image_path.tolist()\nlabels = fold_df[CFG.target_col].values\nds = build_dataset(paths, labels, cache=False, batch_size=32,\n                   repeat=True, shuffle=True, augment=False)\nds = ds.unbatch().batch(20)\nbatch = next(iter(ds))\ndisplay_batch(batch, 5);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:08:44.366936Z","iopub.execute_input":"2023-02-11T08:08:44.367418Z","iopub.status.idle":"2023-02-11T08:08:49.968731Z","shell.execute_reply.started":"2023-02-11T08:08:44.367378Z","shell.execute_reply":"2023-02-11T08:08:49.967632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# GradCAM Utils","metadata":{}},{"cell_type":"code","source":"import matplotlib.cm as cm, cv2\nfrom tensorflow import keras\n\ndef gen_gradcam_heatmap(img, model, last_conv='top_conv', pred_index=0):\n    # First, we create a model that maps the input image to the activations\n    # of the last conv layer as well as the output predictions\n    grad_model = tf.keras.models.Model(\n        [model.inputs], [model.get_layer(last_conv).output, model.output]\n    )\n\n    # Then, we compute the gradient of the top predicted class for our input image\n    # with respect to the activations of the last conv layer\n    with tf.GradientTape() as tape:\n        features, preds = grad_model(img)\n        print(img.shape)\n        if pred_index is None:\n            pred_index = tf.argmax(preds[0])\n        class_channel = preds[:, pred_index]\n\n    # This is the gradient of the output neuron (top predicted or chosen)\n    # with regard to the output feature map of the last conv layer\n    grads = tape.gradient(class_channel, features)\n\n    # This is a vector where each entry is the mean intensity of the gradient\n    # over a specific feature map channel\n    pooled_grads = tf.reduce_mean(grads, axis=(0, 1, 2))\n\n    # We multiply each channel in the feature map array\n    # by \"how important this channel is\" with regard to the top predicted class\n    # then sum all the channels to obtain the heatmap class activation\n    features = features[0]\n    heatmap = features @ pooled_grads[..., tf.newaxis]\n    heatmap = tf.squeeze(heatmap)\n\n    # For visualization purpose, we will also normalize the heatmap between 0 & 1\n    heatmap = tf.maximum(heatmap, 0) / tf.math.reduce_max(heatmap)\n    return heatmap.numpy(), preds[0]\n\n\ndef get_gradcam(img, model):\n\n    heatmap, pred = gen_gradcam_heatmap(img, model, last_conv='top_conv', pred_index=0)\n    heatmap = cv2.resize(heatmap, dsize=(img.shape[2], img.shape[1]))\n        \n    return heatmap, pred\n\ndef gray2rgb(img, gradcam, alpha=0.55, cmap='inferno'):\n    \"\"\"Convert grayscale heatmap to jet gradcam\"\"\"\n    # Use jet colormap to colorize heatmap\n    cmap = cm.get_cmap(cmap)\n\n    # Use RGB values of the colormap\n    colors  = cmap(np.arange(256))[:, :3]\n    heatmap = colors[gradcam]\n\n    # Superimpose the heatmap on original image\n    sup_img = heatmap * alpha + (img/255.0)\n    sup_img = tf.keras.preprocessing.image.array_to_img(sup_img) # avoid saturated colors\n    return np.asarray(sup_img)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:09:00.345996Z","iopub.execute_input":"2023-02-11T08:09:00.346414Z","iopub.status.idle":"2023-02-11T08:09:00.51986Z","shell.execute_reply.started":"2023-02-11T08:09:00.34638Z","shell.execute_reply":"2023-02-11T08:09:00.518829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# BBox Utils","metadata":{}},{"cell_type":"code","source":"import cv2\nfrom bbox.utils import draw_bboxes\nimport matplotlib.cm as cm\n\ndef gradcam2bbox(gradcam, thr=0.6):\n    # Binarize the image\n    mask = (gradcam>thr).astype('uint8')\n\n    # Make contours around the binarized image, keep only the largest contour\n    contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_NONE)\n    contour = max(contours, key=cv2.contourArea)\n\n    # Find ROI from largest contour\n    ys = contour.squeeze()[:, 0]\n    xs = contour.squeeze()[:, 1]\n    \n    y1 = np.min(xs); y2 = np.max(xs);\n    x1 = np.min(ys); x2 = np.max(ys);\n    \n    return [x1, y1, x2, y2]\n\nnp.random.seed(32)\ncolors = [(np.random.randint(255), np.random.randint(255), np.random.randint(255))\\\n          for idx in range(1)]\nfont_param = {\"fontname\":\"Times New Roman\",\"fontweight\":\"bold\", \"fontsize\":16}","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:09:08.373973Z","iopub.execute_input":"2023-02-11T08:09:08.374609Z","iopub.status.idle":"2023-02-11T08:09:09.081979Z","shell.execute_reply.started":"2023-02-11T08:09:08.374549Z","shell.execute_reply":"2023-02-11T08:09:09.080924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Wandb** Logger\nLog:\n* Image\n* Mask\n* BBox","metadata":{}},{"cell_type":"code","source":"class_labels = {\n    1: \"cancer\",\n}\n\ndef wandb_init(fold):\n    config = {k:v for k,v in dict(vars(CFG)).items() if '__' not in k}\n    config.update({\"fold\":int(fold)})\n    run    = wandb.init(project=\"rsna-bcd-gradcam2bbox\",\n               name=f\"fold-{fold}|dim-{CFG.img_size[0]}x{CFG.img_size[1]}|model-{CFG.model_name}\",\n               config=config,\n               anonymous=anonymous,\n                    )\n    return run\n\ndef log_wandb(fold, data):\n    # Create wandb table with gradcam, bbox, mask\n    wandb_data = []\n    for idx, row in enumerate(tqdm(data, desc='wandb ')):\n        tab = row[0]\n        img = row[1]\n        gradcam = gray2rgb(img, row[2])\n        mask = (row[2]>(CFG.mask_thr*255.0)).astype('uint8')\n        bbox = tab[-2][0]\n        \n        # Add row to table\n        wandb_data+=[[*tab,\n                wandb.Image(img), \n                wandb.Image(gradcam),\n                wandb.Image(img, masks={\n                    \"ground_truth\": {\"mask_data\": mask,\n                                     \"class_labels\": class_labels,}\n                }),\n                wandb.Image(img, boxes={\"ground_truth\": {\"box_data\": [{\"position\": {\"minX\": int(bbox[0]),\n                                                                                    \"maxX\": int(bbox[2]),\n                                                                                    \"minY\": int(bbox[1]),\n                                                                                    \"maxY\": int(bbox[3])},\n                                                                       \"class_id\" : 1,\n                                                                       \"box_caption\": \"cancer\",\n                                                                       \"domain\" : \"pixel\",}],\n                                                         \"class_labels\": class_labels,},\n                                       })\n               ]]\n        \n    # Create Table\n    wandb_table = wandb.Table(data=wandb_data,\n                              columns=[*cols, 'image', 'gradcam','mask_img','bbox_img'])\n    \n    # Log data to wandb\n    wandb.log({'viz_table':wandb_table,})","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:09:16.081497Z","iopub.execute_input":"2023-02-11T08:09:16.082513Z","iopub.status.idle":"2023-02-11T08:09:16.095726Z","shell.execute_reply.started":"2023-02-11T08:09:16.082476Z","shell.execute_reply":"2023-02-11T08:09:16.094315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Generate BBox from GradCAM","metadata":{}},{"cell_type":"code","source":"oof_pred = []; oof_tar = []; oof_val = []; oof_ids = []; oof_folds = []\npreds = np.zeros((test_df.shape[0],1))\n\nfor fold in np.arange(CFG.folds):\n    \n    # Ignore not selected folds\n    if fold not in CFG.selected_folds:\n        continue\n        \n    # Init wandb\n    if CFG.wandb:\n        run = wandb_init(fold)\n\n    # Train and valid dataframe\n    valid_df = df.query(\"fold==@fold\")\n    \n    # Get image_paths and labels\n    valid_paths = valid_df.image_path.values; valid_labels = valid_df[CFG.target_col].values.astype(np.float32)\n    \n    # Min samples in debug mode\n    min_samples = CFG.batch_size*REPLICAS*2\n    \n    # For debug model run on small portion\n    if CFG.debug:\n        valid_paths = valid_paths[:min_samples]; valid_labels = valid_labels[:min_samples]\n    \n    # Sample count\n    num_valid = len(valid_paths)\n    num_valid_pos = (valid_labels==1).sum()\n    \n    # Show message\n    print('#'*40); print('#### FOLD: ',fold)\n    print('#### NUM_SAMPLES: {:,} | NUM_CANCER: {:,}'.format(num_valid, num_valid_pos))\n    print('#'*40)\n    \n    # Build model\n    print('Loading model...')\n    K.clear_session()\n    with strategy.scope():\n        model = tf.keras.models.load_model(f'{CKPT_PATH}/fold-%i.h5'%fold,\n                                          compile=False)\n        # Save the model weights to a H5 file\n        model.save(\"model.h5\")\n\n        # Log the H5 file as an artifact in W&B\n        wandb.save(\"model.h5\")\n        \n    \n    # Celect top and worst 10 cases for each class\n    gradcam_df  = valid_df.query(\"cancer==1\").reset_index(drop=True)\n    gradcam_ds  = build_dataset(gradcam_df.image_path.values, labels=None, cache=False, batch_size=1,\n                   repeat=False, shuffle=False, augment=False)\n    \n    # Columns\n    noimg_cols  = ['image_id','patient_id','cancer']\n    cols = noimg_cols + ['bbox','pred']\n    \n    # Create wandb table for upload\n    data = []\n    for idx, img in enumerate(tqdm(gradcam_ds, total=len(gradcam_df), desc='gradcam ')):\n        gradcam, pred = get_gradcam(img, model)\n        bbox = [gradcam2bbox(gradcam, thr=CFG.mask_thr)]\n        tab = gradcam_df.loc[idx, noimg_cols].tolist()+[bbox, pred[0].numpy()]\n        img = img.numpy()[0]\n        img = (img*255.0).astype('uint8')\n        gradcam = (gradcam*255.0).astype('uint8')\n        if pred>=CFG.conf_thr:\n            data+=[[tab, img, gradcam]]\n    \n    # Save bbox data for each fold\n    bbox_df = pd.DataFrame([x[0] for x in data], columns=cols)\n    bbox_df.to_csv(f'bbox{fold}.csv',index=False)\n            \n    # Log bbox, mask, gradcam on wandb & plot\n    if CFG.wandb:\n        log_wandb(fold, data) # log\n        wandb.run.finish() # finish the run\n        display(ipd.IFrame(run.url, width=1080, height=720)) # show wandb dashboard\n        \n    # Clean unless last fold (last fold for local plotting)\n    if fold!=CFG.selected_folds[-1]:\n        del data; gc.collect()\n        \n    print();","metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"scrolled":true,"execution":{"iopub.status.busy":"2023-02-11T08:43:30.291768Z","iopub.execute_input":"2023-02-11T08:43:30.292171Z","iopub.status.idle":"2023-02-11T09:02:59.926398Z","shell.execute_reply.started":"2023-02-11T08:43:30.292138Z","shell.execute_reply":"2023-02-11T09:02:59.925247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualize Gradcam & BBox","metadata":{}},{"cell_type":"code","source":"for idx in range(5):\n    row = data[idx]\n    img = row[1].copy()\n    gradcam = gray2rgb(img, row[2])\n    mask = (row[2]>(CFG.mask_thr*255.0)).astype('uint8')\n    \n    tab = row[0]\n    iid = tab[0]; pid = tab[1]; view = tab[-4]; lat = tab[-5]; pred=tab[-1]\n    bboxes = tab[-2]  # voc pascal style\n    \n    plt.figure(figsize=(20,10))\n    print(\"#\"*20)\n    print(f\"ImageID: {iid}\\nPatientID: {pid}\\nView: {view}\\nLaterality: {lat}\")\n    print(f\"Pred: {pred:0.3f}\")\n    print(\"#\"*20)\n\n    # Plot image\n    plt.subplot(141)\n    plt.imshow(img)\n    plt.title(\"Image\", **font_param); plt.axis('off')\n\n    # Plot GradCAM\n    plt.subplot(142)\n    plt.imshow(gradcam)\n    plt.title(\"GradCAM\", **font_param); plt.axis('off')\n\n    # Plot Mask\n    plt.subplot(143)\n    plt.imshow(img)\n    plt.imshow(np.ma.masked_where(mask==0, mask), alpha=0.3, cmap='autumn')\n    plt.title(\"Mask\", **font_param); plt.axis('off')\n\n    # Plot BBox\n    plt.subplot(144)\n    labels = [0]*1\n    names = ['cancer']*1\n    plt.imshow(draw_bboxes(img = img,\n                               bboxes = bboxes, \n                               classes = names,\n                               class_ids = labels,\n                               class_name = True, \n                               colors = colors, \n                               bbox_format = 'voc',\n                               line_thickness = 4))\n    plt.title(\"BBox\", **font_param); plt.axis('off')\n\n    plt.tight_layout()\n    plt.show()\n    print()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-11T08:29:25.439289Z","iopub.execute_input":"2023-02-11T08:29:25.439794Z","iopub.status.idle":"2023-02-11T08:29:31.317636Z","shell.execute_reply.started":"2023-02-11T08:29:25.439753Z","shell.execute_reply":"2023-02-11T08:29:31.316665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# import shutil\n# shutil.make_archive(models, 'zip', '/kaggle/input/rsna-bcd-efnb4-1024x512-v2-ckpt-ds')","metadata":{"execution":{"iopub.status.busy":"2023-02-11T10:50:34.895097Z","iopub.execute_input":"2023-02-11T10:50:34.895541Z","iopub.status.idle":"2023-02-11T10:50:34.900336Z","shell.execute_reply.started":"2023-02-11T10:50:34.895507Z","shell.execute_reply":"2023-02-11T10:50:34.899116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Store BBox","metadata":{}},{"cell_type":"code","source":"bbox_df = pd.concat([pd.read_csv(f'bbox{fold}.csv') for fold in CFG.selected_folds], axis=0, ignore_index=True)\nprint(f'BBox Size: {len(bbox_df)}')\nbbox_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-14T09:37:52.072858Z","iopub.execute_input":"2022-12-14T09:37:52.073575Z","iopub.status.idle":"2022-12-14T09:37:52.094189Z","shell.execute_reply.started":"2022-12-14T09:37:52.073538Z","shell.execute_reply":"2022-12-14T09:37:52.09311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3'  # to avoid too many logging messages\nimport pandas as pd, numpy as np, random, shutil\nimport tensorflow as tf, re, math\nimport sklearn\nimport matplotlib.pyplot as plt\nimport yaml\nimport gc\ngc.collect()\nfrom PIL import Image\nfrom IPython import display as ipd\nfrom glob import glob\nfrom tqdm.notebook import tqdm\nfrom sklearn.model_selection import KFold, StratifiedKFold, GroupKFold, StratifiedGroupKFold\nfrom sklearn.metrics import roc_auc_score","metadata":{"execution":{"iopub.status.busy":"2023-02-11T10:01:19.442505Z","iopub.execute_input":"2023-02-11T10:01:19.443272Z","iopub.status.idle":"2023-02-11T10:01:19.773925Z","shell.execute_reply.started":"2023-02-11T10:01:19.443232Z","shell.execute_reply":"2023-02-11T10:01:19.772842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.cm as cm, cv2\nfrom tensorflow import keras\n\ndef gen_gradcam_heatmap(img, model, last_conv='top_conv', pred_index=0):\n    # First, we create a model that maps the input image to the activations\n    # of the last conv layer as well as the output predictions\n    grad_model = tf.keras.models.Model(\n        [model.inputs], [model.get_layer(last_conv).output, model.output]\n    )\n\n    # Then, we compute the gradient of the top predicted class for our input image\n    # with respect to the activations of the last conv layer\n    with tf.GradientTape() as tape:\n        features, preds = grad_model(img)\n        if pred_index is None:\n            pred_index = tf.argmax(preds[0])\n        class_channel = preds[:, pred_index]\n        \n        \n\n    # This is the gradient of the output neuron (top predicted or chosen)\n    # with regard to the output feature map of the last conv layer\n    grads = tape.gradient(class_channel, features)\n    \n\n    # This is a vector where each entry is the mean intensity of the gradient\n    # over a specific feature map channel\n    pooled_grads = tf.reduce_mean(grads, axis=(0, 1, 2))\n    \n    \n\n    # We multiply each channel in the feature map array\n    # by \"how important this channel is\" with regard to the top predicted class\n    # then sum all the channels to obtain the heatmap class activation\n    features = features[0]\n    heatmap = features @ pooled_grads[..., tf.newaxis]\n    heatmap = tf.squeeze(heatmap)\n    \n    # For visualization purpose, we will also normalize the heatmap between 0 & 1\n    heatmap = tf.maximum(heatmap, 0) / tf.math.reduce_max(heatmap)\n    return heatmap.numpy(), preds[0]\n\n\ndef get_gradcam(img, model):\n\n    heatmap, pred = gen_gradcam_heatmap(img, model, last_conv='top_conv', pred_index=0)\n    heatmap = cv2.resize(heatmap, dsize=(img.shape[2], img.shape[1]))\n        \n    return heatmap, pred\n\ndef gray2rgb(img, gradcam, alpha=0.55, cmap='inferno'):\n    \"\"\"Convert grayscale heatmap to jet gradcam\"\"\"\n    # Use jet colormap to colorize heatmap\n    cmap = cm.get_cmap(cmap)\n\n    # Use RGB values of the colormap\n    colors  = cmap(np.arange(256))[:, :3]\n    heatmap = colors[gradcam]\n\n    # Superimpose the heatmap on original image\n    sup_img = heatmap * alpha + (img/255.0)\n    sup_img = tf.keras.preprocessing.image.array_to_img(sup_img) # avoid saturated colors\n    return np.asarray(sup_img)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T10:57:24.724152Z","iopub.execute_input":"2023-02-11T10:57:24.724674Z","iopub.status.idle":"2023-02-11T10:57:24.739788Z","shell.execute_reply.started":"2023-02-11T10:57:24.724634Z","shell.execute_reply":"2023-02-11T10:57:24.738616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import cv2\nfrom bbox.utils import draw_bboxes\nimport matplotlib.cm as cm\n\ndef gradcam2bbox(gradcam, thr=0.6):\n    # Binarize the image\n    mask = (gradcam>thr).astype('uint8')\n\n    # Make contours around the binarized image, keep only the largest contour\n    contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_NONE)\n    contour = max(contours, key=cv2.contourArea)\n\n    # Find ROI from largest contour\n    ys = contour.squeeze()[:, 0]\n    xs = contour.squeeze()[:, 1]\n    \n    y1 = np.min(xs); y2 = np.max(xs);\n    x1 = np.min(ys); x2 = np.max(ys);\n    \n    return [x1, y1, x2, y2]\n\nnp.random.seed(32)\ncolors = [(np.random.randint(255), np.random.randint(255), np.random.randint(255))\\\n          for idx in range(1)]\nfont_param = {\"fontname\":\"Times New Roman\",\"fontweight\":\"bold\", \"fontsize\":16}","metadata":{"execution":{"iopub.status.busy":"2023-02-11T10:57:26.795971Z","iopub.execute_input":"2023-02-11T10:57:26.796386Z","iopub.status.idle":"2023-02-11T10:57:26.809427Z","shell.execute_reply.started":"2023-02-11T10:57:26.79635Z","shell.execute_reply":"2023-02-11T10:57:26.808265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import PIL.Image\n\nfp = open(\"/kaggle/input/images/0.png\",\"rb\")\nimg = PIL.Image.open(fp).resize((512,1024))\npath='/kaggle/input/rsna-bcd-efnb4-1024x512-v2-ckpt-ds/fold-0.h5'\nmodel = tf.keras.models.load_model(path,compile=False)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T10:57:26.993222Z","iopub.execute_input":"2023-02-11T10:57:26.993622Z","iopub.status.idle":"2023-02-11T10:57:31.292997Z","shell.execute_reply.started":"2023-02-11T10:57:26.993583Z","shell.execute_reply":"2023-02-11T10:57:31.29185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# image = np.array(Image.open(\"/kaggle/input/images/0.png\").resize((512,1024)))\nimg = tf.convert_to_tensor(image)\nimg=np.expand_dims(img,axis=0)\n# img = tf.convert_to_tensor(img_to_tensor)\ngradcam, pred = get_gradcam(img, model)\nprint(pred,gradcam)\n# bbox = [gradcam2bbox(gradcam, thr=CFG.mask_thr)]\n# tab = gradcam_df.loc[idx, noimg_cols].tolist()+[bbox, pred[0].numpy()]\n# img = img.numpy()[0]\n# img = (img*255.0).astype('uint8')\n# gradcam = (gradcam*255.0).astype('uint8')\n\n","metadata":{"execution":{"iopub.status.busy":"2023-02-11T10:57:31.295582Z","iopub.execute_input":"2023-02-11T10:57:31.295932Z","iopub.status.idle":"2023-02-11T10:57:31.876864Z","shell.execute_reply.started":"2023-02-11T10:57:31.295903Z","shell.execute_reply":"2023-02-11T10:57:31.875652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" with strategy.scope():\n        model = tf.keras.models.load_model(f'{CKPT_PATH}/fold-%i.h5'%fold,\n                                          compile=False)\ngradcam, pred = get_gradcam(img, model)\nprint(pred,gradcam)","metadata":{"execution":{"iopub.status.busy":"2023-02-11T10:56:55.989726Z","iopub.execute_input":"2023-02-11T10:56:55.990479Z","iopub.status.idle":"2023-02-11T10:57:00.859205Z","shell.execute_reply.started":"2023-02-11T10:56:55.990436Z","shell.execute_reply":"2023-02-11T10:57:00.858037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Remove Files","metadata":{}},{"cell_type":"code","source":"!rm -r /kaggle/working/wandb","metadata":{"execution":{"iopub.status.busy":"2022-12-14T06:48:46.893463Z","iopub.status.idle":"2022-12-14T06:48:46.894178Z","shell.execute_reply.started":"2022-12-14T06:48:46.893919Z","shell.execute_reply":"2022-12-14T06:48:46.893943Z"},"trusted":true},"execution_count":null,"outputs":[]}]}