{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":52254,"databundleVersionId":6863140,"sourceType":"competition"},{"sourceId":6211844,"sourceType":"datasetVersion","datasetId":3567114},{"sourceId":6526093,"sourceType":"datasetVersion","datasetId":3772985}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<center><img src=\"https://keras.io/img/logo-small.png\" alt=\"Keras logo\" width=\"100\"><br/>\nThis starter notebook is provided by the Keras team.</center>","metadata":{}},{"cell_type":"markdown","source":"# Training Notebook\n\n# RSNA 2023 Abdominal Trauma Detection with [KerasCV](https://github.com/keras-team/keras-cv) and [KerasCore](https://github.com/keras-team/keras-core)\n\nThis notebook walks you through how to train a **Convolutional Neural Network (CNN)** model using Keras (Core and CV) on the RSNA 2023 Abdominal Trauma Detection dataset made available for this competition.\n\nFun fact: This notebook is backend (tensorflow, pytorch, jax) agnostic. Using KerasCV and KerasCore we can choose a backend of our choise! Feel free to read [Keras Core](https://keras.io/keras_core/announcement/) announcement to know more about Keras.\n\nIn this notebook you will learn:\n\n* Loading the data using [`tf.data`](https://www.tensorflow.org/guide/data).\n* Applying augmentations inside the data pipeline.\n* Create the model using KerasCV presets.\n* Train the model.\n* Visualize the training plots.\n\n## Notebooks\n\nFor this competition we have two starter notebook. This notebook (you are reading) trains the model on the dataset, while there lies another notebook that performs inference and submits to the competition.\n\n1. [**Training Kernel**](https://www.kaggle.com/code/aritrag/kerascv-starter-notebook-train)\n2. [**Inference Kernel**](https://www.kaggle.com/code/aritrag/kerascv-starter-notebook-infer)\n\n**Note**: [KerasCV guides](https://keras.io/guides/keras_cv/) is the place to go for a deeper understanding of KerasCV individually.","metadata":{}},{"cell_type":"markdown","source":"# Setup and Imports\n\nWe will need KerasCV for this notebook.\n\nFeel free to use `pip install keras-cv` instead of the installation from github.","metadata":{}},{"cell_type":"code","source":"! pip install -q git+https://github.com/keras-team/keras-cv\n#!pip install keras-cv","metadata":{"_kg_hide-output":true,"_kg_hide-input":false,"execution":{"iopub.status.busy":"2023-10-28T13:46:59.19015Z","iopub.execute_input":"2023-10-28T13:46:59.190389Z","iopub.status.idle":"2023-10-28T13:47:29.553784Z","shell.execute_reply.started":"2023-10-28T13:46:59.190365Z","shell.execute_reply":"2023-10-28T13:47:29.552766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\n# You can use `tensorflow`, `pytorch`, `jax` here\n# KerasCore makes the notebook backend agnostic :)\nos.environ[\"KERAS_BACKEND\"] = \"tensorflow\"\n\nimport keras_cv\nimport keras_core as keras\nfrom keras_core import layers\n\nimport numpy as np\nimport pandas as pd\nimport tensorflow as tf\nfrom matplotlib import pyplot as plt\nfrom sklearn.model_selection import train_test_split\nfrom tqdm.notebook import tqdm\nimport gc\nimport pandas.api.types\nimport sklearn.metrics\nimport tensorflow_addons as tfa","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:47:34.330223Z","iopub.execute_input":"2023-10-28T13:47:34.330591Z","iopub.status.idle":"2023-10-28T13:47:47.516048Z","shell.execute_reply.started":"2023-10-28T13:47:34.330556Z","shell.execute_reply":"2023-10-28T13:47:47.515184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Configuration\n\nA particularly good practise is to have a configuration class for your notebooks. This not only keeps your configurations all at a single place but also becomes handy to map the configs to the performance of the model.\n\nPlease play around with the configurations and see how the performance of the model changes.","metadata":{}},{"cell_type":"markdown","source":"## Note on some observations\n\nReference Notebook: https://www.kaggle.com/code/aritrag/eda-train-csv\n\n1. Class Dependencies: Refers to inherent relationships between classes in the analysis.\n2. Complementarity: `bowel_injury` and `bowel_healthy`, as well as `extravasation_injury` and `extravasation_healthy`, are perfectly complementary, with their sum always equal to 1.0.\n3. Simplification: For the model, only `{bowel/extravasation}_injury` will be included, and the corresponding healthy status can be calculated using a sigmoid function.\n4. Softmax: `{kidney/liver/spleen}_{healthy/low/high}` classifications are softmaxed, ensuring their combined probabilities sum up to 1.0 for each organ, simplifying the model while preserving essential information.","metadata":{}},{"cell_type":"code","source":"class Config:\n    SEED = 42\n    IMAGE_SIZE = [260, 260]\n    BATCH_SIZE = 64\n    EPOCHS = 40\n    TARGET_COLS  = [\n        \"bowel_injury\", \"extravasation_injury\",\n        \"kidney_healthy\", \"kidney_low\", \"kidney_high\",\n        \"liver_healthy\", \"liver_low\", \"liver_high\",\n        \"spleen_healthy\", \"spleen_low\", \"spleen_high\",\n    ]\n    PRED_COLS  = [\n        'bowel_healthy', \"bowel_injury\",  \n        'extravasation_healthy', \"extravasation_injury\",\n        \"kidney_healthy\", \"kidney_low\", \"kidney_high\",\n        \"liver_healthy\", \"liver_low\", \"liver_high\",\n        \"spleen_healthy\", \"spleen_low\", \"spleen_high\",\n    ]\n    AUTOTUNE = tf.data.AUTOTUNE\n\nconfig = Config()","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:49:48.708964Z","iopub.execute_input":"2023-10-28T13:49:48.709356Z","iopub.status.idle":"2023-10-28T13:49:48.715929Z","shell.execute_reply.started":"2023-10-28T13:49:48.709325Z","shell.execute_reply":"2023-10-28T13:49:48.714864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Reproducibility\n\nWe would want this notebook to have reproducible results. Here we set the seed for all the random algorithms so that we can reproduce the experiments each time exactly the same way.","metadata":{}},{"cell_type":"code","source":"keras.utils.set_random_seed(seed=config.SEED)","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:47:53.639036Z","iopub.execute_input":"2023-10-28T13:47:53.639393Z","iopub.status.idle":"2023-10-28T13:47:53.644099Z","shell.execute_reply.started":"2023-10-28T13:47:53.639364Z","shell.execute_reply":"2023-10-28T13:47:53.643043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Dataset\n\nThe dataset provided in the competition consists of DICOM images. We will not be training on the DICOM images, rather would work on PNG image which are extracted from the DICOM format.\n\n[A helpful resource on the conversion of DICOM to PNG](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs)","metadata":{}},{"cell_type":"code","source":"BASE_PATH = f\"/kaggle/input/rsna-atd-512x512-png-v2-dataset\"","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-28T13:47:58.699574Z","iopub.execute_input":"2023-10-28T13:47:58.70025Z","iopub.status.idle":"2023-10-28T13:47:58.704239Z","shell.execute_reply.started":"2023-10-28T13:47:58.700217Z","shell.execute_reply":"2023-10-28T13:47:58.703164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Meta Data\n\nThe `train.csv` file contains the following meta information:\n\n- `patient_id`: A unique ID code for each patient.\n- `series_id`: A unique ID code for each scan.\n- `instance_number`: The image number within the scan. The lowest instance number for many series is above zero as the original scans were cropped to the abdomen.\n- `[bowel/extravasation]_[healthy/injury]`: The two injury types with binary targets.\n- `[kidney/liver/spleen]_[healthy/low/high]`: The three injury types with three target levels.\n- `any_injury`: Whether the patient had any injury at all.\n","metadata":{}},{"cell_type":"code","source":"# train\ndataframe = pd.read_csv(f\"{BASE_PATH}/train.csv\")\ndataframe[\"image_path\"] = f\"{BASE_PATH}/train_images\"\\\n                   + \"/\" + dataframe.patient_id.astype(str)\\\n                    + \"/\" + dataframe.series_id.astype(str)\\\n                    + \"/\" + dataframe.instance_number.astype(str) +\".png\"\ndataframe = dataframe.drop_duplicates()","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:48:01.436567Z","iopub.execute_input":"2023-10-28T13:48:01.437436Z","iopub.status.idle":"2023-10-28T13:48:01.571911Z","shell.execute_reply.started":"2023-10-28T13:48:01.4374Z","shell.execute_reply":"2023-10-28T13:48:01.570963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We split the training dataset into train and validation. This is a common practise in the Machine Learning pipelines. We not only want to train our model, but also want to validate it's training.\n\nA small catch here is that the training and validation data should have an aligned data distribution. Here we handle that by grouping the lables and then splitting the dataset. This ensures an aligned data distribution between the training and the validation splits.","metadata":{}},{"cell_type":"code","source":"# Function to handle the split for each group\ndef split_group(group, test_size=0.2):\n    if len(group) == 1:\n        return (group, pd.DataFrame()) if np.random.rand() < test_size else (pd.DataFrame(), group)\n    else:\n        return train_test_split(group, test_size=test_size, random_state=42)\n\n# Initialize the train and validation datasets\ntrain_data = pd.DataFrame()\nval_data = pd.DataFrame()\ntest_data = pd.DataFrame()\n\n# Iterate through the groups and split them, handling single-sample groups\nfor _, group in dataframe.groupby(config.TARGET_COLS):\n    train_group, test_group = split_group(group,.3)\n    train_data = pd.concat([train_data, train_group], ignore_index=True)\n    test_data = pd.concat([test_data, test_group], ignore_index=True)\n    \n# Iterate through the groups and split them, handling single-sample groups\nfor _, group in train_data.groupby(config.TARGET_COLS):\n    train_group, val_group = split_group(group)\n    train_data = pd.concat([train_data, train_group], ignore_index=True)\n    val_data = pd.concat([val_data, val_group], ignore_index=True)","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:48:03.36847Z","iopub.execute_input":"2023-10-28T13:48:03.368879Z","iopub.status.idle":"2023-10-28T13:48:03.543342Z","shell.execute_reply.started":"2023-10-28T13:48:03.368815Z","shell.execute_reply":"2023-10-28T13:48:03.542571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.shape, val_data.shape, test_data.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:48:05.631331Z","iopub.execute_input":"2023-10-28T13:48:05.631685Z","iopub.status.idle":"2023-10-28T13:48:05.638882Z","shell.execute_reply.started":"2023-10-28T13:48:05.631656Z","shell.execute_reply":"2023-10-28T13:48:05.637865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Pipeline /w tf.data\n\nHere we build the data pipeline using `tf.data`. Using `tf.data` we can map out data to an augmentation pipeline simple by using the ` map` API.\n\nAdding augmentations to the data pipeline is as simple as adding a layer into the list of layers that the `Augmenter` processes.\n\nReference: https://keras.io/api/keras_cv/layers/augmentation/","metadata":{}},{"cell_type":"code","source":"def decode_image_and_label(image_path, label):\n    file_bytes = tf.io.read_file(image_path)\n    image = tf.io.decode_png(file_bytes, channels=3, dtype=tf.uint8)\n    image = tf.image.resize(image, config.IMAGE_SIZE, method=\"bilinear\")\n    image = tf.cast(image, tf.float32) / 255.0\n    \n    label = tf.cast(label, tf.float32)\n    #         bowel       fluid       kidney      liver       spleen\n    labels = (label[0:1], label[1:2], label[2:5], label[5:8], label[8:11])\n    \n    return (image, labels)\n\n\ndef apply_augmentation(images, labels):\n    augmenter = keras_cv.layers.Augmenter(\n        [\n            keras_cv.layers.RandomFlip(mode=\"horizontal_and_vertical\"),\n            keras_cv.layers.RandomCutout(height_factor=0.2, width_factor=0.2),\n            \n        ]\n    )\n    return (augmenter(images), labels)\n\n\ndef build_dataset(image_paths, labels):\n    ds = (\n        tf.data.Dataset.from_tensor_slices((image_paths, labels))\n        .map(decode_image_and_label, num_parallel_calls=config.AUTOTUNE)\n        .shuffle(config.BATCH_SIZE * 10)\n        .batch(config.BATCH_SIZE)\n        .map(apply_augmentation, num_parallel_calls=config.AUTOTUNE)\n        .prefetch(config.AUTOTUNE)\n    )\n    return ds","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:48:16.482436Z","iopub.execute_input":"2023-10-28T13:48:16.483079Z","iopub.status.idle":"2023-10-28T13:48:16.492744Z","shell.execute_reply.started":"2023-10-28T13:48:16.483044Z","shell.execute_reply":"2023-10-28T13:48:16.491767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"paths  = train_data.image_path.tolist()\nlabels = train_data[config.TARGET_COLS].values\n\nds = build_dataset(image_paths=paths, labels=labels)\nimages, labels = next(iter(ds))\nimages.shape, [label.shape for label in labels]","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:48:48.107919Z","iopub.execute_input":"2023-10-28T13:48:48.108806Z","iopub.status.idle":"2023-10-28T13:49:04.493041Z","shell.execute_reply.started":"2023-10-28T13:48:48.108767Z","shell.execute_reply":"2023-10-28T13:49:04.492127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# No more customizing your plots by hand, KerasCV has your back ;)\nkeras_cv.visualization.plot_image_gallery(\n    images=images,\n    value_range=(0, 1),\n    rows=1,\n    cols=2,\n)","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:49:06.431516Z","iopub.execute_input":"2023-10-28T13:49:06.432232Z","iopub.status.idle":"2023-10-28T13:49:06.81204Z","shell.execute_reply.started":"2023-10-28T13:49:06.4322Z","shell.execute_reply":"2023-10-28T13:49:06.811096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Build Model\n\nWe are going to load a pretrained model from the [list of avaiable backbones in KerasCV](https://keras.io/api/keras_cv/models/backbones/). We are using the `ResNetBackbone` as our backbone. The practise of using a pretrained model and finetuning it to a specific dataset is prevalent in the DL community.\n\nWe use the [Functional API](https://keras.io/guides/functional_api/) of Keras to build the model. The design of the model would be such that we input a single image and we get different heads for the various predictions we need (kidney, spleen...).\n\nWe have also added a Learning Rate scheduler for you to work with. When an athlete trains, the first step is always to warm up. We take a similar approach to training our models. We warm up with model where the learning rate increses from the initial LR to a higher LR. After the warmup stage we provide a decay algorithm (cosine here). A list of all the learning rate scheduler can be found [here](https://keras.io/api/optimizers/learning_rate_schedules/).","metadata":{}},{"cell_type":"code","source":"def build_model(warmup_steps, decay_steps):\n    # Define Input\n    inputs = keras.Input(shape=config.IMAGE_SIZE + [3,], batch_size=config.BATCH_SIZE)\n    \n    # Define Backbone\n    backbone = keras_cv.models.ResNetV2Backbone.from_preset(\"resnet50_v2\")\n    backbone.include_rescaling = False\n    x = backbone(inputs)\n    \n    # GAP to get the activation maps\n    gap = keras.layers.GlobalAveragePooling2D()\n    x = gap(x)\n\n    # Define 'necks' for each head\n    x_bowel = keras.layers.Dense(32, activation='silu')(x)\n    x_extra = keras.layers.Dense(32, activation='silu')(x)\n    x_liver = keras.layers.Dense(32, activation='silu')(x)\n    x_kidney = keras.layers.Dense(32, activation='silu')(x)\n    x_spleen = keras.layers.Dense(32, activation='silu')(x)\n\n    # Define heads\n    out_bowel = keras.layers.Dense(1, name='bowel', activation='sigmoid')(x_bowel) # use sigmoid to convert predictions to [0-1]\n    out_extra = keras.layers.Dense(1, name='extra', activation='sigmoid')(x_extra) # use sigmoid to convert predictions to [0-1]\n    out_liver = keras.layers.Dense(3, name='liver', activation='softmax')(x_liver) # use softmax for the liver head\n    out_kidney = keras.layers.Dense(3, name='kidney', activation='softmax')(x_kidney) # use softmax for the kidney head\n    out_spleen = keras.layers.Dense(3, name='spleen', activation='softmax')(x_spleen) # use softmax for the spleen head\n    \n    # Concatenate the outputs\n    outputs = [out_bowel, out_extra, out_liver, out_kidney, out_spleen]\n\n    # Create model\n    print(\"[INFO] Building the model...\")\n    model = keras.Model(inputs=inputs, outputs=outputs)\n    \n    # Cosine Decay\n    cosine_decay = keras.optimizers.schedules.CosineDecay(\n        initial_learning_rate=1e-4,\n        decay_steps=decay_steps,\n        alpha=0.0,\n        warmup_target=1e-3,\n        warmup_steps=warmup_steps,\n    )\n    \n    # Exponential Decay\n    #exp_decay = keras.optimizers.schedules.ExponentialDecay(\n    #    initial_learning_rate=.1,\n    #    decay_steps=decay_steps,\n    #    decay_rate=0.9,\n    #    staircase=True,\n    #)\n    \n    # Compile the model\n    optimizer = keras.optimizers.Adam(learning_rate=cosine_decay)\n    loss = {\n        \"bowel\":keras.losses.BinaryCrossentropy(),\n        \"extra\":keras.losses.BinaryCrossentropy(),\n        \"liver\":keras.losses.CategoricalCrossentropy(),\n        \"kidney\":keras.losses.CategoricalCrossentropy(),\n        \"spleen\":keras.losses.CategoricalCrossentropy(),\n    }\n    metrics = {\n        \"bowel\":[\"accuracy\"],\n        \"extra\":[\"accuracy\"],\n        \"liver\":[\"accuracy\"],\n        \"kidney\":[\"accuracy\"],\n        \"spleen\":[\"accuracy\"],\n        \n    }\n    print(\"[INFO] Compiling the model...\")\n    model.compile(\n        optimizer=optimizer,\n      loss=loss,\n      metrics=metrics\n    )\n    \n    return model","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:49:22.195608Z","iopub.execute_input":"2023-10-28T13:49:22.196283Z","iopub.status.idle":"2023-10-28T13:49:22.211021Z","shell.execute_reply.started":"2023-10-28T13:49:22.196251Z","shell.execute_reply":"2023-10-28T13:49:22.21004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train the model with \"model.fit\"","metadata":{}},{"cell_type":"code","source":"# get image_paths and labels\nprint(\"[INFO] Building the dataset...\")\ntrain_paths = train_data.image_path.values; train_labels = train_data[config.TARGET_COLS].values.astype(np.float32)\nvalid_paths = val_data.image_path.values; valid_labels = val_data[config.TARGET_COLS].values.astype(np.float32)\n#test_paths  = test_data.image_path.values; test_labels = test_data[config.TARGET_COLS].values.astype(np.float32)\n\n# train and valid dataset\ntrain_ds = build_dataset(image_paths=train_paths, labels=train_labels)\nval_ds = build_dataset(image_paths=valid_paths, labels=valid_labels)\n#test_ds = build_dataset(image_paths=test_paths, labels=test_labels)\n\ntotal_train_steps = train_ds.cardinality().numpy() * config.BATCH_SIZE * config.EPOCHS\nwarmup_steps = int(total_train_steps * 0.10)\ndecay_steps = total_train_steps - warmup_steps\n\nprint(f\"{total_train_steps=}\")\nprint(f\"{warmup_steps=}\")\nprint(f\"{decay_steps=}\")","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:49:29.865213Z","iopub.execute_input":"2023-10-28T13:49:29.865593Z","iopub.status.idle":"2023-10-28T13:49:31.538715Z","shell.execute_reply.started":"2023-10-28T13:49:29.865563Z","shell.execute_reply":"2023-10-28T13:49:31.537789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# build the model\nprint(\"[INFO] Building the model...\")\nmodel = build_model(warmup_steps, decay_steps)\n\n# train\nprint(\"[INFO] Training...\")\nhistory = model.fit(\n    train_ds,\n    epochs=config.EPOCHS,\n    validation_data=val_ds,\n)","metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2023-10-28T13:49:55.284654Z","iopub.execute_input":"2023-10-28T13:49:55.285336Z","iopub.status.idle":"2023-10-28T13:55:57.014812Z","shell.execute_reply.started":"2023-10-28T13:49:55.285301Z","shell.execute_reply":"2023-10-28T13:55:57.014013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualize the training plots","metadata":{}},{"cell_type":"code","source":"# Create a 3x2 grid for the subplots\nfig, axes = plt.subplots(5, 1, figsize=(5, 15))\n\n# Flatten axes to iterate through them\naxes = axes.flatten()\n\n# Iterate through the metrics and plot them\nfor i, name in enumerate([\"bowel\", \"extra\", \"kidney\", \"liver\", \"spleen\"]):\n    # Plot training accuracy\n    axes[i].plot(history.history[name + '_accuracy'], label='Training ' + name)\n    # Plot validation accuracy\n    axes[i].plot(history.history['val_' + name + '_accuracy'], label='Validation ' + name)\n    axes[i].set_title(name)\n    axes[i].set_xlabel('Epoch')\n    axes[i].set_ylabel('Accuracy')\n    axes[i].legend()\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:56:01.116754Z","iopub.execute_input":"2023-10-28T13:56:01.11714Z","iopub.status.idle":"2023-10-28T13:56:02.270569Z","shell.execute_reply.started":"2023-10-28T13:56:01.117107Z","shell.execute_reply":"2023-10-28T13:56:02.26967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(history.history[\"loss\"], label=\"loss\")\nplt.plot(history.history[\"val_loss\"], label=\"val loss\")\nplt.legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:56:09.357247Z","iopub.execute_input":"2023-10-28T13:56:09.3576Z","iopub.status.idle":"2023-10-28T13:56:09.602361Z","shell.execute_reply.started":"2023-10-28T13:56:09.357569Z","shell.execute_reply":"2023-10-28T13:56:09.601449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# store best results\nbest_epoch = np.argmin(history.history['val_loss'])\nbest_loss = history.history['val_loss'][best_epoch]\nbest_acc_bowel = history.history['val_bowel_accuracy'][best_epoch]\nbest_acc_extra = history.history['val_extra_accuracy'][best_epoch]\nbest_acc_liver = history.history['val_liver_accuracy'][best_epoch]\nbest_acc_kidney = history.history['val_kidney_accuracy'][best_epoch]\nbest_acc_spleen = history.history['val_spleen_accuracy'][best_epoch]\n\n# Find mean accuracy\nbest_acc = np.mean(\n    [best_acc_bowel,\n     best_acc_extra,\n     best_acc_liver,\n     best_acc_kidney,\n     best_acc_spleen\n])\n\n\nprint(f'>>>> BEST Loss  : {best_loss:.3f}\\n>>>> BEST Acc   : {best_acc:.3f}\\n>>>> BEST Epoch : {best_epoch}\\n')\nprint('ORGAN Acc:')\nprint(f'  >>>> {\"Bowel\".ljust(15)} : {best_acc_bowel:.3f}')\nprint(f'  >>>> {\"Extravasation\".ljust(15)} : {best_acc_extra:.3f}')\nprint(f'  >>>> {\"Liver\".ljust(15)} : {best_acc_liver:.3f}')\nprint(f'  >>>> {\"Kidney\".ljust(15)} : {best_acc_kidney:.3f}')\nprint(f'  >>>> {\"Spleen\".ljust(15)} : {best_acc_spleen:.3f}')","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:56:12.094268Z","iopub.execute_input":"2023-10-28T13:56:12.094969Z","iopub.status.idle":"2023-10-28T13:56:12.103413Z","shell.execute_reply.started":"2023-10-28T13:56:12.094936Z","shell.execute_reply":"2023-10-28T13:56:12.102542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Store the model for inference","metadata":{}},{"cell_type":"code","source":"'''\n# Save the model\nmodel = keras.models.load_model(MODEL_PATH)\nmodel.summary()\n'''\n\nmodel.save(\"myModel.keras\")","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:56:14.83125Z","iopub.execute_input":"2023-10-28T13:56:14.83186Z","iopub.status.idle":"2023-10-28T13:56:16.319918Z","shell.execute_reply.started":"2023-10-28T13:56:14.831814Z","shell.execute_reply":"2023-10-28T13:56:16.318894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def decode_image(image_path):\n    file_bytes = tf.io.read_file(image_path)\n    image = tf.io.decode_png(file_bytes, channels=3, dtype=tf.uint8)\n    image = tf.image.resize(image, config.IMAGE_SIZE, method=\"bilinear\")\n    image = tf.cast(image, tf.float32) / 255.0\n    return image\n\ndef build_dataset(image_paths):\n    ds = (\n        tf.data.Dataset.from_tensor_slices(image_paths)\n        .map(decode_image, num_parallel_calls=config.AUTOTUNE)\n        .shuffle(config.BATCH_SIZE * 10)\n        .batch(config.BATCH_SIZE)\n        .prefetch(config.AUTOTUNE)\n    )\n    return ds","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:56:19.754245Z","iopub.execute_input":"2023-10-28T13:56:19.755129Z","iopub.status.idle":"2023-10-28T13:56:19.761606Z","shell.execute_reply.started":"2023-10-28T13:56:19.755097Z","shell.execute_reply":"2023-10-28T13:56:19.760671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def post_proc(pred):\n    proc_pred = np.empty((pred.shape[0], 2*2 + 3*3), dtype=\"float32\")\n    print(pred)\n    in_bow = pred[:, 0]\n    h_bow = 1 - in_bow\n    in_ex = pred[:, 1]\n    h_ex = 1 - in_ex\n    \n    # bowel, extravasation\n    proc_pred[:, 0] = h_bow\n    proc_pred[:, 1] = in_bow\n    proc_pred[:, 2] = h_ex\n    proc_pred[:, 3] = in_ex\n    \n    # liver, kidney, sneel\n    proc_pred[:, 4:7] = pred[:, 2:5]\n    proc_pred[:, 7:10] = pred[:, 5:8]\n    proc_pred[:, 10:13] = pred[:, 8:11]\n\n    return proc_pred","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:56:23.633579Z","iopub.execute_input":"2023-10-28T13:56:23.633933Z","iopub.status.idle":"2023-10-28T13:56:23.641142Z","shell.execute_reply.started":"2023-10-28T13:56:23.633905Z","shell.execute_reply":"2023-10-28T13:56:23.639995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Getting unique patient IDs from test dataset\npatient_ids = test_data[\"patient_id\"].unique()\n#patient_ids = patient_ids[1:3]\n\n# Initializing array to store predictions\npatient_preds = np.zeros(\n    shape=(len(patient_ids), 2*2 + 3*3),\n    dtype=\"float32\"\n)\n\n# Iterating over each patient\nfor pidx, patient_id in tqdm(enumerate(patient_ids), total=len(patient_ids), desc=\"Patients \"):\n    print(f\"Patient ID: {patient_id}\")\n    \n    # Query the dataframe for a particular patient\n    #patient_df = test_df.query(\"patient_id == patient_id\")\n    patient_df = test_data[test_data['patient_id'] == patient_id]\n    \n    # Getting image paths for a patient\n    patient_paths = patient_df.image_path.tolist()\n\n    # Building dataset for prediction\n    dtest = build_dataset(patient_paths)\n    \n    # Predicting with the model\n    pred = model.predict(dtest)\n    pred = np.concatenate(pred, axis=-1).astype(\"float32\")\n    pred = pred[:len(patient_paths), :]\n    pred = np.mean(pred.reshape(1, len(patient_paths), 11), axis=0)\n    pred = np.max(pred, axis=0, keepdims=True)\n    \n    patient_preds[pidx, :] += post_proc(pred)[0]\n\n    # Deleting variables to free up memory \n    del patient_df, patient_paths, dtest, pred; gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:56:26.991888Z","iopub.execute_input":"2023-10-28T13:56:26.992817Z","iopub.status.idle":"2023-10-28T13:56:34.485095Z","shell.execute_reply.started":"2023-10-28T13:56:26.99278Z","shell.execute_reply":"2023-10-28T13:56:34.484107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df = pd.DataFrame({\"patient_id\":patient_ids,})\npred_df[config.PRED_COLS] = patient_preds.astype('float32')","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:56:37.854344Z","iopub.execute_input":"2023-10-28T13:56:37.854716Z","iopub.status.idle":"2023-10-28T13:56:37.864757Z","shell.execute_reply.started":"2023-10-28T13:56:37.854686Z","shell.execute_reply":"2023-10-28T13:56:37.863729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:56:44.796975Z","iopub.execute_input":"2023-10-28T13:56:44.797775Z","iopub.status.idle":"2023-10-28T13:56:44.816674Z","shell.execute_reply.started":"2023-10-28T13:56:44.797736Z","shell.execute_reply":"2023-10-28T13:56:44.815778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def normalize_probabilities_to_one(df: pd.DataFrame, group_columns: list) -> pd.DataFrame:\n    # Normalize the sum of each row's probabilities to 100%.\n    # 0.75, 0.75 => 0.5, 0.5\n    # 0.1, 0.1 => 0.5, 0.5\n    row_totals = df[group_columns].sum(axis=1)\n    if row_totals.min() == 0:\n        print('All rows must contain at least one non-zero prediction')\n    for col in group_columns:\n        df[col] /= row_totals\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:56:52.692901Z","iopub.execute_input":"2023-10-28T13:56:52.69324Z","iopub.status.idle":"2023-10-28T13:56:52.698913Z","shell.execute_reply.started":"2023-10-28T13:56:52.693213Z","shell.execute_reply":"2023-10-28T13:56:52.697961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Run basic QC checks on the inputs\nif not pandas.api.types.is_numeric_dtype(pred_df.values):\n    print('All submission values must be numeric')\n\nif not np.isfinite(pred_df.values).all():\n    print('All submission values must be finite')\n\nif pred_df.min().min() < 0:\n    print('All predictions must be at least zero')","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:56:54.642067Z","iopub.execute_input":"2023-10-28T13:56:54.642781Z","iopub.status.idle":"2023-10-28T13:56:54.650524Z","shell.execute_reply.started":"2023-10-28T13:56:54.642751Z","shell.execute_reply":"2023-10-28T13:56:54.649598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"true_df = test_data.iloc[:,:14]\ntrue_df = true_df.drop_duplicates()\n\ntrue_df.shape\n\n\n#true_df2 = pd.DataFrame()\n#for i in range(len(pred_df['patient_id'])):\n#    tmp = true_df[true_df['patient_id'] == pred_df.iloc[i,0]]\n#    true_df2 = true_df2._append(tmp)\n    \n#true_df = true_df2\n#true_df","metadata":{"execution":{"iopub.status.busy":"2023-10-28T13:57:03.562742Z","iopub.execute_input":"2023-10-28T13:57:03.563516Z","iopub.status.idle":"2023-10-28T13:57:03.583946Z","shell.execute_reply.started":"2023-10-28T13:57:03.563484Z","shell.execute_reply":"2023-10-28T13:57:03.583078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the label group log losses\nbinary_targets = ['bowel', 'extravasation']\ntriple_level_targets = ['kidney', 'liver', 'spleen']\nall_target_categories = binary_targets + triple_level_targets\n\nlabel_group_losses = []\nlabel_group_precision = []\nlabel_group_recall = []\nlabel_group_accuracy = []\nlabel_group_precision_multi = []\nlabel_group_recall_multi = []\nlabel_group_accuracy_multi = []\nlabel_group_f1 = []\nprecision = tf.keras.metrics.Precision()\nrecall = tf.keras.metrics.Recall()\nacc = tf.keras.metrics.BinaryAccuracy(threshold=.5)\n\nfor category in all_target_categories:\n    \n    if category in binary_targets:\n        col_group = [f'{category}_healthy', f'{category}_injury']\n    else:\n        col_group = [f'{category}_healthy', f'{category}_low', f'{category}_high']\n\n    true_df = normalize_probabilities_to_one(true_df, col_group)\n\n    for col in col_group:\n        if col not in pred_df.columns:\n            print(f'Missing submission column {col}')\n    pred_df = normalize_probabilities_to_one(pred_df, col_group)\n    \n    y_true = true_df[col_group].values\n    y_pred = pred_df[col_group].values\n    \n    label_group_losses.append(\n        sklearn.metrics.log_loss(\n            y_true,\n            y_pred\n        )    \n    )\n    #print(y_true)\n    #print(y_pred)\n    precision.update_state(y_true,y_pred)\n    label_group_precision_multi.append(precision.result().numpy())\n    \n    recall.update_state(y_true,y_pred)\n    label_group_recall_multi.append(recall.result().numpy())\n    \n    acc.update_state(y_true,y_pred)\n    label_group_accuracy_multi.append(acc.result().numpy())\n    \n    f1 = tfa.metrics.F1Score(num_classes=len(col_group))\n    f1.update_state(y_true, y_pred)\n    label_group_f1.append(f1.result().numpy())\n    \nfor col in config.PRED_COLS:\n    y_true = true_df[col].values\n    y_pred = pred_df[col].values\n    \n    precision.update_state(y_true,y_pred)\n    label_group_precision.append(precision.result().numpy())\n    \n    recall.update_state(y_true,y_pred)\n    label_group_recall.append(recall.result().numpy())\n    \n    acc.update_state(y_true,y_pred)\n    label_group_accuracy.append(acc.result().numpy())","metadata":{"execution":{"iopub.status.busy":"2023-10-28T14:00:21.319869Z","iopub.execute_input":"2023-10-28T14:00:21.320799Z","iopub.status.idle":"2023-10-28T14:00:21.68216Z","shell.execute_reply.started":"2023-10-28T14:00:21.320764Z","shell.execute_reply":"2023-10-28T14:00:21.681369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"losses \", label_group_losses)\nprint(\"precision \", label_group_precision)\nprint(\"recall \", label_group_recall)\nprint(\"accuracy \", label_group_accuracy)\nprint(\"precision multiclass\", label_group_precision_multi)\nprint(\"recall multiclass\", label_group_recall_multi)\nprint(\"accuracy multiclass\", label_group_accuracy_multi)\nprint(\"F1 score \", label_group_f1)","metadata":{"execution":{"iopub.status.busy":"2023-10-28T14:01:16.922885Z","iopub.execute_input":"2023-10-28T14:01:16.92323Z","iopub.status.idle":"2023-10-28T14:01:16.930679Z","shell.execute_reply.started":"2023-10-28T14:01:16.923204Z","shell.execute_reply":"2023-10-28T14:01:16.929501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}