{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<center><img src=\"https://keras.io/img/logo-small.png\" alt=\"Keras logo\" width=\"100\"><br/>\nThis starter notebook is provided by the Keras team.</center>","metadata":{}},{"cell_type":"markdown","source":"# Inference Notebook\n\n# RSNA 2023 Abdominal Trauma Detection with [KerasCV](https://github.com/keras-team/keras-cv) and [KerasCore](https://github.com/keras-team/keras-core)\n\nThis notebook walks you through how to submit to the RSNA 2023 Abdominal Trauma Detection competition.\n\n## Notebooks\n\nFor this competition we have two starter notebook. This notebook (you are reading) infers on the trained model. To understand how the model was trained from scratch please visit the training kernel.\n\n1. [**Training Kernel**](https://www.kaggle.com/code/aritrag/kerascv-starter-notebook-train)\n2. [**Inference Kernel**](https://www.kaggle.com/code/aritrag/kerascv-starter-notebook-infer)\n\n**Note**: [KerasCV guides](https://keras.io/guides/keras_cv/) is the place to go for a deeper understanding of KerasCV individually.","metadata":{}},{"cell_type":"markdown","source":"# Imports and Setup\n\nPlease install namex, keras-cv and keras-core for this notebook to work!\n\nNote: We will very soon have keras and keras models accessible from offline notebook =D ","metadata":{}},{"cell_type":"code","source":"! pip install -q /kaggle/input/keras-cv-core-namex/namex-0.0.7-py3-none-any.whl\n! pip install -q /kaggle/input/keras-cv-core-namex/keras_core-0.1.4-py3-none-any.whl\n! pip install -q /kaggle/input/keras-cv-core-namex/keras_cv-0.6.1-py3-none-any.whl","metadata":{"execution":{"iopub.status.busy":"2023-10-02T14:00:06.896152Z","iopub.execute_input":"2023-10-02T14:00:06.896384Z","iopub.status.idle":"2023-10-02T14:01:47.412767Z","shell.execute_reply.started":"2023-10-02T14:00:06.896363Z","shell.execute_reply":"2023-10-02T14:01:47.411491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nos.environ[\"KERAS_BACKEND\"] = \"tensorflow\"\nimport keras_core as keras\nimport keras_cv\n\nimport gc\nimport cv2\nimport pydicom\nfrom joblib import Parallel, delayed\n\nimport tensorflow as tf\nimport numpy as np\nimport pandas as pd\nfrom tqdm.notebook import tqdm\nfrom glob import glob","metadata":{"execution":{"iopub.status.busy":"2023-10-02T14:01:47.415042Z","iopub.execute_input":"2023-10-02T14:01:47.415999Z","iopub.status.idle":"2023-10-02T14:01:59.629925Z","shell.execute_reply.started":"2023-10-02T14:01:47.415943Z","shell.execute_reply":"2023-10-02T14:01:59.629046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"BASE_PATH = \"/kaggle/input/rsna-2023-abdominal-trauma-detection\"\nIMAGE_DIR = \"/tmp/dataset/rsna-atd\"\nINPUT_MODEL_PATH = \"/kaggle/input/kerascv-starter-notebook-train-1baaf6/rsna-atd.keras\"\nMODEL_PATH = \"/kaggle/working/rsna-atd.keras\"\nSTRIDE = 10","metadata":{"execution":{"iopub.status.busy":"2023-10-02T14:05:16.497721Z","iopub.execute_input":"2023-10-02T14:05:16.498195Z","iopub.status.idle":"2023-10-02T14:05:16.503518Z","shell.execute_reply.started":"2023-10-02T14:05:16.498161Z","shell.execute_reply":"2023-10-02T14:05:16.502502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Config:\n    IMAGE_SIZE = [256, 256]\n    RESIZE_DIM = 256\n    BATCH_SIZE = 64\n    AUTOTUNE = tf.data.AUTOTUNE\n    TARGET_COLS  = [\"bowel_healthy\", \"bowel_injury\", \"extravasation_healthy\",\n                   \"extravasation_injury\", \"kidney_healthy\", \"kidney_low\",\n                   \"kidney_high\", \"liver_healthy\", \"liver_low\", \"liver_high\",\n                   \"spleen_healthy\", \"spleen_low\", \"spleen_high\"]\n\nconfig = Config()","metadata":{"execution":{"iopub.status.busy":"2023-10-02T14:05:17.890305Z","iopub.execute_input":"2023-10-02T14:05:17.890632Z","iopub.status.idle":"2023-10-02T14:05:17.895918Z","shell.execute_reply.started":"2023-10-02T14:05:17.890605Z","shell.execute_reply":"2023-10-02T14:05:17.894957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Initialize the Trained Model\n\nRefer to the [training kernel](https://www.kaggle.com/code/aritrag/kerascv-starter-notebook-train) to understand how we have trained the model.\n\nHere we initialize the trained model.","metadata":{}},{"cell_type":"code","source":"! cp {INPUT_MODEL_PATH} ./\n\nmodel = keras.models.load_model(MODEL_PATH)\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2023-10-02T14:05:20.92526Z","iopub.execute_input":"2023-10-02T14:05:20.925591Z","iopub.status.idle":"2023-10-02T14:05:21.977838Z","shell.execute_reply.started":"2023-10-02T14:05:20.925564Z","shell.execute_reply":"2023-10-02T14:05:21.975398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Pipeline\n\nWe have a model that accepts image in the PNG format. We will now convert the DICOM images to PNG for the model to infer.","metadata":{}},{"cell_type":"code","source":"meta_df = pd.read_csv(f\"{BASE_PATH}/test_series_meta.csv\")\n\n# Checking if patients are repeated by finding the number of unique patient IDs\nnum_rows = meta_df.shape[0]\nunique_patients = meta_df[\"patient_id\"].nunique()\n\nprint(f\"{num_rows=}\")\nprint(f\"{unique_patients=}\")","metadata":{"execution":{"iopub.status.busy":"2023-10-02T14:05:37.256656Z","iopub.execute_input":"2023-10-02T14:05:37.257132Z","iopub.status.idle":"2023-10-02T14:05:37.288326Z","shell.execute_reply.started":"2023-10-02T14:05:37.257093Z","shell.execute_reply":"2023-10-02T14:05:37.287342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"meta_df[\"dicom_folder\"] = BASE_PATH + \"/\" + \"test_images\"\\\n                                    + \"/\" + meta_df.patient_id.astype(str)\\\n                                    + \"/\" + meta_df.series_id.astype(str)\n\ntest_folders = meta_df.dicom_folder.tolist()\ntest_paths = []\nfor folder in tqdm(test_folders):\n    test_paths += sorted(glob(os.path.join(folder, \"*dcm\")))[::STRIDE]","metadata":{"execution":{"iopub.status.busy":"2023-10-02T14:05:47.827433Z","iopub.execute_input":"2023-10-02T14:05:47.827757Z","iopub.status.idle":"2023-10-02T14:05:47.869371Z","shell.execute_reply.started":"2023-10-02T14:05:47.827731Z","shell.execute_reply":"2023-10-02T14:05:47.868522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = pd.DataFrame(test_paths, columns=[\"dicom_path\"])\ntest_df[\"patient_id\"] = test_df.dicom_path.map(lambda x: x.split(\"/\")[-3]).astype(int)\ntest_df[\"series_id\"] = test_df.dicom_path.map(lambda x: x.split(\"/\")[-2]).astype(int)\ntest_df[\"instance_number\"] = test_df.dicom_path.map(lambda x: x.split(\"/\")[-1].replace(\".dcm\",\"\")).astype(int)\n\ntest_df[\"image_path\"] = f\"{IMAGE_DIR}/test_images\"\\\n                    + \"/\" + test_df.patient_id.astype(str)\\\n                    + \"/\" + test_df.series_id.astype(str)\\\n                    + \"/\" + test_df.instance_number.astype(str) +\".png\"\n\ntest_df.head(2)","metadata":{"execution":{"iopub.status.busy":"2023-10-02T14:05:52.762963Z","iopub.execute_input":"2023-10-02T14:05:52.763323Z","iopub.status.idle":"2023-10-02T14:05:52.783692Z","shell.execute_reply.started":"2023-10-02T14:05:52.763296Z","shell.execute_reply":"2023-10-02T14:05:52.782573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking if patients are repeated by finding the number of unique patient IDs\nnum_rows = test_df.shape[0]\nunique_patients = test_df[\"patient_id\"].nunique()\n\nprint(f\"{num_rows=}\")\nprint(f\"{unique_patients=}\")","metadata":{"execution":{"iopub.status.busy":"2023-10-02T14:05:57.930581Z","iopub.execute_input":"2023-10-02T14:05:57.930901Z","iopub.status.idle":"2023-10-02T14:05:57.937421Z","shell.execute_reply.started":"2023-10-02T14:05:57.930875Z","shell.execute_reply":"2023-10-02T14:05:57.936479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## DICOM to PNG pipeline","metadata":{}},{"cell_type":"code","source":"!rm -r {IMAGE_DIR}\nos.makedirs(f\"{IMAGE_DIR}/train_images\", exist_ok=True)\nos.makedirs(f\"{IMAGE_DIR}/test_images\", exist_ok=True)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-02T14:06:03.861237Z","iopub.execute_input":"2023-10-02T14:06:03.861856Z","iopub.status.idle":"2023-10-02T14:06:04.833191Z","shell.execute_reply.started":"2023-10-02T14:06:03.861826Z","shell.execute_reply":"2023-10-02T14:06:04.832106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def standardize_pixel_array(dcm):\n    # Correct DICOM pixel_array if PixelRepresentation == 1.\n    pixel_array = dcm.pixel_array\n    if dcm.PixelRepresentation == 1:\n        bit_shift = dcm.BitsAllocated - dcm.BitsStored\n        dtype = pixel_array.dtype \n        new_array = (pixel_array << bit_shift).astype(dtype) >>  bit_shift\n        pixel_array = pydicom.pixel_data_handlers.util.apply_modality_lut(new_array, dcm)\n    return pixel_array\n\ndef read_xray(path, fix_monochrome=True):\n    dicom = pydicom.dcmread(path)\n    data = standardize_pixel_array(dicom)\n    data = data - np.min(data)\n    data = data / (np.max(data) + 1e-5)\n    if fix_monochrome and dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = 1.0 - data\n    return data\n\ndef resize_and_save(file_path):\n    img = read_xray(file_path)\n    h, w = img.shape[:2]  # orig hw\n    img = cv2.resize(img, (config.RESIZE_DIM, config.RESIZE_DIM), cv2.INTER_LINEAR)\n    img = (img * 255).astype(np.uint8)\n    \n    sub_path = file_path.split(\"/\",4)[-1].split(\".dcm\")[0] + \".png\"\n    infos = sub_path.split(\"/\")\n    sub_path = file_path.split(\"/\",4)[-1].split(\".dcm\")[0] + \".png\"\n    infos = sub_path.split(\"/\")\n    pid = infos[-3]\n    sid = infos[-2]\n    iid = infos[-1]; iid = iid.replace(\".png\",\"\")\n    new_path = os.path.join(IMAGE_DIR, sub_path)\n    os.makedirs(new_path.rsplit(\"/\",1)[0], exist_ok=True)\n    cv2.imwrite(new_path, img)\n    return","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-02T14:06:46.82489Z","iopub.execute_input":"2023-10-02T14:06:46.825918Z","iopub.status.idle":"2023-10-02T14:06:46.836602Z","shell.execute_reply.started":"2023-10-02T14:06:46.825873Z","shell.execute_reply":"2023-10-02T14:06:46.835713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nfile_paths = test_df.dicom_path.tolist()\n_ = Parallel(n_jobs=2, backend=\"threading\")(\n    delayed(resize_and_save)(file_path) for file_path in tqdm(file_paths, leave=True, position=0)\n)\n\ndel _; gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-10-02T14:06:55.221203Z","iopub.execute_input":"2023-10-02T14:06:55.221525Z","iopub.status.idle":"2023-10-02T14:06:55.635709Z","shell.execute_reply.started":"2023-10-02T14:06:55.221497Z","shell.execute_reply":"2023-10-02T14:06:55.634822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Building the tf.data pipeline","metadata":{}},{"cell_type":"code","source":"def decode_image(image_path):\n    file_bytes = tf.io.read_file(image_path)\n    image = tf.io.decode_png(file_bytes, channels=3, dtype=tf.uint8)\n    image = tf.image.resize(image, config.IMAGE_SIZE, method=\"bilinear\")\n    image = tf.cast(image, tf.float32) / 255.0\n    return image\n\ndef build_dataset(image_paths):\n    ds = (\n        tf.data.Dataset.from_tensor_slices(image_paths)\n        .map(decode_image, num_parallel_calls=config.AUTOTUNE)\n        .shuffle(config.BATCH_SIZE * 10)\n        .batch(config.BATCH_SIZE)\n        .prefetch(config.AUTOTUNE)\n    )\n    return ds","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-02T14:07:00.75995Z","iopub.execute_input":"2023-10-02T14:07:00.760402Z","iopub.status.idle":"2023-10-02T14:07:00.769963Z","shell.execute_reply.started":"2023-10-02T14:07:00.760369Z","shell.execute_reply":"2023-10-02T14:07:00.768914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"paths  = test_df.image_path.tolist()\n\nds = build_dataset(paths)\nimages = next(iter(ds))\n\nimages.shape","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-02T14:07:05.227306Z","iopub.execute_input":"2023-10-02T14:07:05.227619Z","iopub.status.idle":"2023-10-02T14:07:10.975234Z","shell.execute_reply.started":"2023-10-02T14:07:05.227594Z","shell.execute_reply":"2023-10-02T14:07:10.97432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"keras_cv.visualization.plot_image_gallery(\n    images=images,\n    value_range=(0, 1),\n    rows=1,\n    cols=3,\n)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-02T14:07:23.194153Z","iopub.execute_input":"2023-10-02T14:07:23.194506Z","iopub.status.idle":"2023-10-02T14:07:23.456062Z","shell.execute_reply.started":"2023-10-02T14:07:23.194471Z","shell.execute_reply":"2023-10-02T14:07:23.452637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Inference","metadata":{}},{"cell_type":"code","source":"def post_proc(pred):\n    proc_pred = np.empty((pred.shape[0], 2*2 + 3*3), dtype=\"float32\")\n\n    # bowel, extravasation\n    proc_pred[:, 0] = pred[:, 0]\n    proc_pred[:, 1] = 1 - proc_pred[:, 0]\n    proc_pred[:, 2] = pred[:, 1]\n    proc_pred[:, 3] = 1 - proc_pred[:, 2]\n    \n    # liver, kidney, sneel\n    proc_pred[:, 4:7] = pred[:, 2:5]\n    proc_pred[:, 7:10] = pred[:, 5:8]\n    proc_pred[:, 10:13] = pred[:, 8:11]\n\n    return proc_pred","metadata":{"execution":{"iopub.status.busy":"2023-10-02T14:07:29.968378Z","iopub.execute_input":"2023-10-02T14:07:29.968704Z","iopub.status.idle":"2023-10-02T14:07:29.975394Z","shell.execute_reply.started":"2023-10-02T14:07:29.968678Z","shell.execute_reply":"2023-10-02T14:07:29.974496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Getting unique patient IDs from test dataset\npatient_ids = test_df[\"patient_id\"].unique()\n\n# Initializing array to store predictions\npatient_preds = np.zeros(\n    shape=(len(patient_ids), 2*2 + 3*3),\n    dtype=\"float32\"\n)\n\n# Iterating over each patient\nfor pidx, patient_id in tqdm(enumerate(patient_ids), total=len(patient_ids), desc=\"Patients \"):\n    print(f\"Patient ID: {patient_id}\")\n    \n    # Query the dataframe for a particular patient\n    patient_df = test_df.query(\"patient_id == @patient_id\")\n    \n    # Getting image paths for a patient\n    patient_paths = patient_df.image_path.tolist()\n\n    # Building dataset for prediction\n    dtest = build_dataset(patient_paths)\n    \n    # Predicting with the model\n    pred = model.predict(dtest)\n    pred = np.concatenate(pred, axis=-1).astype(\"float32\")\n    pred = pred[:len(patient_paths), :]\n    pred = np.mean(pred.reshape(1, len(patient_paths), 11), axis=0)\n    pred = np.max(pred, axis=0, keepdims=True)\n    \n    patient_preds[pidx, :] += post_proc(pred)[0]\n    \n\n    # Deleting variables to free up memory \n    del patient_df, patient_paths, dtest, pred; gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-10-02T14:07:34.662805Z","iopub.execute_input":"2023-10-02T14:07:34.663152Z","iopub.status.idle":"2023-10-02T14:07:35.214367Z","shell.execute_reply.started":"2023-10-02T14:07:34.663123Z","shell.execute_reply":"2023-10-02T14:07:35.21304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"!rm -rf {MODEL_PATH}","metadata":{"execution":{"iopub.status.busy":"2023-10-02T14:07:55.359901Z","iopub.execute_input":"2023-10-02T14:07:55.360288Z","iopub.status.idle":"2023-10-02T14:07:56.582702Z","shell.execute_reply.started":"2023-10-02T14:07:55.360262Z","shell.execute_reply":"2023-10-02T14:07:56.581399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create Submission\npred_df = pd.DataFrame({\"patient_id\":patient_ids,})\npred_df[config.TARGET_COLS] = patient_preds.astype(\"float32\")\n\n# Align with sample submission\nsub_df = pd.read_csv(f\"{BASE_PATH}/sample_submission.csv\")\nsub_df = sub_df[[\"patient_id\"]]\nsub_df = sub_df.merge(pred_df, on=\"patient_id\", how=\"left\")\n\n# Store submission\nsub_df.to_csv(\"submission.csv\",index=False)\nsub_df.head(2)","metadata":{"execution":{"iopub.status.busy":"2023-10-02T14:08:00.000961Z","iopub.execute_input":"2023-10-02T14:08:00.001345Z","iopub.status.idle":"2023-10-02T14:08:00.045314Z","shell.execute_reply.started":"2023-10-02T14:08:00.001315Z","shell.execute_reply":"2023-10-02T14:08:00.044361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}