{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This is the optimized version of Mark's [Notebook](https://www.kaggle.com/code/markwijkhuizen/rsna-convnextv2-inference-tensorflow).\nThe changes I made are the followings.\n\n  - Decode jpeg image using Nvidia DALI instead of DicomSDL\n  - Augment image using torch tensors instead of numpy ndarrays\n\nReference:  \n  [SE-ResNeXt50 GPU optimized](https://www.kaggle.com/code/tivfrvqhs5/se-resnext50-gpu-optimized)  \n  [SE-ResNeXt50 full GPU decoding](https://www.kaggle.com/code/christofhenkel/se-resnext50-full-gpu-decoding)\n\n\n---------------------------------------------------","metadata":{}},{"cell_type":"markdown","source":"Hello fellow Kagglers,\n\nThis notebook demonstrates the inference process using an EfficientNetV2T model training on a TPU.\n\nThe inference process consists of 2 steps.\n\n1) Preprocess and save all images in parallel, images are cropped, padded to correct aspect ratio and resized to the target size.\n\n2) Prediction on preprocessed images in a simple loop, final cancer score is mean of all images.\n\n**V3**\n\n* Switch to ConvNextV2 models\n\n**V6**\n\nThis will be the final update of my notebooks for this competitions, which should achieve a LB score in the low 0.50s. I will continue participating in this competition, however, I will not share my progress anymore.\n\n* Mainly updated preprocessing procedure, see examples under Example Preprocessing.\n* Correct linear/sigmoid normalization of images, many thanks to [Bob de Graaf](https://www.kaggle.com/bobdegraaf) which shared this amazing notebook: [DicomSDL & VOI-LUT](https://www.kaggle.com/code/bobdegraaf/dicomsdl-voi-lut)\n* Inference should be faster now due to actually using dicomsdl and not just installing it...\n\n**Training Notebook:** [RSNA ConvNextV2 Training Tensorflow TPU](https://www.kaggle.com/code/markwijkhuizen/rsna-efficientnetv2-training-tensorflow-tpu/notebook)\n\n**Data Preprocessing Notebook:** [RSNA Cropped TFRecords 768x1344 Dataset](https://www.kaggle.com/code/markwijkhuizen/rsna-cropped-tfrecords-768x1344-dataset/notebook)\n\n\nGood luck to all of you in the last month of this excisting competition!","metadata":{}},{"cell_type":"code","source":"%%capture\n# Source: https://www.kaggle.com/code/remekkinas/fast-dicom-processing-1-6-2x-faster?scriptVersionId=113360473\n!pip install /kaggle/input/rsnamodules/dicomsdl-0.109.1-cp37-cp37m-manylinux_2_12_x86_64.manylinux2010_x86_64.whl \n\ntry:\n    import pylibjpeg\nexcept:\n   !pip install /kaggle/input/rsna-2022-whl/{pylibjpeg-1.4.0-py3-none-any.whl,python_gdcm-3.0.15-cp37-cp37m-manylinux_2_17_x86_64.manylinux2014_x86_64.whl}","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:17:12.412464Z","iopub.execute_input":"2023-02-19T12:17:12.412832Z","iopub.status.idle":"2023-02-19T12:17:41.579552Z","shell.execute_reply.started":"2023-02-19T12:17:12.412801Z","shell.execute_reply":"2023-02-19T12:17:41.578249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Install Keras CV Attention Model Pip Package for ConvNextV2 Models\n!pip install --no-deps /kaggle/input/keras-cv-attention-models/keras_cv_attention_models-1.3.9-py3-none-any.whl","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:17:41.58265Z","iopub.execute_input":"2023-02-19T12:17:41.583103Z","iopub.status.idle":"2023-02-19T12:18:03.62659Z","shell.execute_reply.started":"2023-02-19T12:17:41.583062Z","shell.execute_reply":"2023-02-19T12:18:03.62536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install -q /kaggle/input/nvidia-dali-nightly-cuda110-1230dev/nvidia_dali_nightly_cuda110-1.23.0.dev20230203-7187866-py3-none-manylinux2014_x86_64.whl","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:18:03.628752Z","iopub.execute_input":"2023-02-19T12:18:03.629196Z","iopub.status.idle":"2023-02-19T12:18:46.484223Z","shell.execute_reply.started":"2023-02-19T12:18:03.629134Z","shell.execute_reply":"2023-02-19T12:18:46.483028Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport pylibjpeg\nimport pydicom\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\n\nimport torch\nimport torch.nn.functional as F\n\nfrom joblib import Parallel, delayed\nfrom tqdm.notebook import tqdm\nfrom multiprocessing import cpu_count\nfrom keras_cv_attention_models import convnext\n\nimport cv2\nimport glob\nimport importlib\nimport os\nimport joblib\nimport time\nimport dicomsdl\nimport gc\n\n# Tensorflow and CV2 set number of threads to 1 for speedup in parallell function mapping\ntf.config.threading.set_inter_op_parallelism_threads(num_threads=1)\ncv2.setNumThreads(1)\n\n# Pandas DataFrame Display Options\npd.options.display.max_colwidth = 99","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:20:29.74212Z","iopub.execute_input":"2023-02-19T12:20:29.742747Z","iopub.status.idle":"2023-02-19T12:20:29.754363Z","shell.execute_reply.started":"2023-02-19T12:20:29.742701Z","shell.execute_reply":"2023-02-19T12:20:29.753227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Config","metadata":{}},{"cell_type":"code","source":"IS_INTERACTIVE = os.environ['KAGGLE_KERNEL_RUN_TYPE'] == 'Interactive'\n\nTARGET_HEIGHT = 1344\nTARGET_WIDTH = 768\nN_CHANNELS = 1\nINPUT_SHAPE = (TARGET_HEIGHT, TARGET_WIDTH, N_CHANNELS)\nTARGET_HEIGHT_WIDTH_RATIO = TARGET_HEIGHT / TARGET_WIDTH\nTHRESHOLD_BEST = 0.50\n\nCLAHE = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(32, 32))\n\nCROP_IMAGE = True\nAPPLY_CLAHE = False\nAPPLY_EQ_HIST = False\n\nIMAGE_FORMAT = 'jpg'","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:18:46.497836Z","iopub.execute_input":"2023-02-19T12:18:46.498245Z","iopub.status.idle":"2023-02-19T12:18:46.515161Z","shell.execute_reply.started":"2023-02-19T12:18:46.498207Z","shell.execute_reply":"2023-02-19T12:18:46.512937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# VOI LUT","metadata":{}},{"cell_type":"code","source":"# Source: https://www.kaggle.com/code/bobdegraaf/dicomsdl-voi-lut\ndef voi_lut(image, dicom, util=np):\n    # Additional Checks\n    if 'WindowWidth' not in dicom.getPixelDataInfo() or 'WindowWidth' not in dicom.getPixelDataInfo():\n        return image\n    \n    # Load only the variables we need\n    center = dicom['WindowCenter']\n    width = dicom['WindowWidth']\n    bits_stored = dicom['BitsStored']\n    voi_lut_function = dicom['VOILUTFunction']\n\n    # For sigmoid it's a list, otherwise a single value\n    if isinstance(center, list):\n        center = center[0]\n    if isinstance(width, list):\n        width = width[0]\n\n    # Set y_min, max & range\n    y_min = 0\n    y_max = float(2**bits_stored - 1)\n    y_range = y_max\n\n    # Function with default LINEAR (so for Nan, it will use linear)\n    if voi_lut_function == 'SIGMOID':\n        image = y_range / (1 + util.exp(-4 * (image - center) / width)) + y_min\n    else:\n        # Checks width for < 1 (in our case not necessary, always >= 750)\n        center -= 0.5\n        width -= 1\n\n        below = image <= (center - width / 2)\n        above = image > (center + width / 2)\n        between = util.logical_and(~below, ~above)\n\n        image[below] = y_min\n        image[above] = y_max\n        if between.any():\n            image[between] = (\n                ((image[between] - center) / width + 0.5) * y_range + y_min\n            )\n\n    return image","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:23:47.961819Z","iopub.execute_input":"2023-02-19T12:23:47.962266Z","iopub.status.idle":"2023-02-19T12:23:47.972227Z","shell.execute_reply.started":"2023-02-19T12:23:47.962229Z","shell.execute_reply":"2023-02-19T12:23:47.970969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Crop Image","metadata":{}},{"cell_type":"code","source":"# Smooth vector used to smoothen sums/stds of axes\ndef smooth(l, util):\n    # kernel size is 1% of vector\n    kernel_size = int(len(l) * 0.01)\n    if hasattr(util, 'convolve'):\n        kernel = util.ones(kernel_size) / kernel_size\n        return util.convolve(l, kernel, mode='same')\n    else:\n        kernel = util.ones(kernel_size,device='cuda') / kernel_size\n        l = l.view(1, 1, l.size(0))\n        kernel = kernel.view(1, 1, kernel.size(0))\n        return F.conv1d(l.type(util.float), kernel, padding='same')[0][0]\n\n# X Crop offset based on first column with sum below 5% of maximum column sums*std\ndef get_x_offset(image, max_col_sum_ratio_threshold=0.05, debug=None, util=np):\n    # Image Dimensions\n    H, W = image.shape\n    # Percentual margin added to offset\n    margin = int(image.shape[1] * 0.00)\n    # Threshold values based on smoothed sum x std to capture varying intensity columns\n    vv = smooth(image.sum(axis=0).squeeze(), util) * smooth(image.std(axis=0).squeeze(), util)\n    # Find maximum sum in first 75% of columns\n    vv_argmax = vv[:int(image.shape[1] * 0.75)].argmax()\n    # Threshold value\n    vv_threshold = vv.max() * max_col_sum_ratio_threshold\n    \n    # Find first column after maximum column below threshold value\n    for offset, v in enumerate(vv):\n        # Start searching from vv_argmax\n        if offset < vv_argmax:\n            continue\n        \n        # Column below threshold value found\n        if v < vv_threshold:\n            offset = min(W, offset + margin)\n            break\n            \n    if isinstance(debug, np.ndarray):\n        if isinstance(image, torch.Tensor):\n            debug[1].imshow(image.cpu().numpy())\n            vv = vv.cpu().numpy()\n            vv_argmax = vv_argmax.cpu().numpy()\n            vv_threshold = vv_threshold.cpu().numpy()\n        else:\n            debug[1].imshow(image)\n        debug[1].set_title('X Offset')\n        vv_scale = H / vv.max() * 0.90\n        # Values\n        debug[1].plot(H - vv * vv_scale , c='red', label='vv')\n        # Threshold\n        debug[1].hlines(H - vv_threshold * vv_scale, 0, W -1, colors='orange', label='threshold')\n        # Max Value\n        debug[1].scatter(vv_argmax, H - vv[vv_argmax] * vv_scale, c='blue', s=100, label='Max', zorder=np.PINF)\n        # First Column Below Threshold\n        debug[1].scatter(offset, H - vv[offset] * vv_scale, c='purple', s=100, label='Offset', zorder=np.PINF)\n        debug[1].set_ylim(H, 0)\n        debug[1].legend()\n        debug[1].axis('off')\n        \n    return offset\n\n# Y Crop offset based on first bottom and top rows with sum below 10% of maximum row sum*std\ndef get_y_offsets(image, max_row_sum_ratio_threshold=0.10, debug=None, util=np):\n    # Image Dimensions\n    H, W = image.shape\n    # Margin to add to offsets\n    margin = 0\n    # Threshold values based on smoothed sum x std to capture varying intensity columns\n    vv = smooth(image.sum(axis=1).squeeze(), util) * smooth(image.std(axis=1).squeeze(), util)\n    # Find maximum sum * std row in inter quartile rows\n    vv_argmax = int(image.shape[0] * 0.25) + vv[int(image.shape[0] * 0.25):int(image.shape[0] * 0.75)].argmax()\n    # Threshold value\n    vv_threshold = vv.max() * max_row_sum_ratio_threshold\n    # Default crop offsets\n    offset_bottom = 0\n    offset_top = H\n\n    # Bottom offset, search from argmax to bottom\n    for offset in reversed(range(0, vv_argmax)):\n        v = vv[offset]\n        if v < vv_threshold:\n            offset_bottom = offset\n            break\n    \n    if isinstance(debug, np.ndarray):\n        if isinstance(image, torch.Tensor):\n            debug[2].imshow(image.cpu().numpy())\n            vv = vv.cpu().numpy()\n            vv_argmax = vv_argmax.cpu().numpy()\n            vv_threshold = vv_threshold.cpu().numpy()\n        else:\n            debug[2].imshow(image)\n        debug[2].set_title('Y Bottom Offset')\n        vv_scale = W / vv.max() * 0.90\n        # Values\n        debug[2].plot(vv * vv_scale, np.arange(H), c='red', label='vv')\n        # Threshold\n        debug[2].vlines(vv_threshold * vv_scale, 0, H -1, colors='orange', label='threshold')\n        # Max Value\n        debug[2].scatter(vv[vv_argmax] * vv_scale, vv_argmax, c='blue', s=100, label='Max', zorder=np.PINF)\n        # First Column Below Threshold\n        debug[2].scatter(vv[offset_bottom] * vv_scale, offset_bottom, c='purple', s=100, label='Offset', zorder=np.PINF)\n        debug[2].set_ylim(H, 0)\n        debug[2].legend()\n        debug[2].axis('off')\n            \n    # Top offset, search from argmax to top\n    for offset in range(vv_argmax, H):\n        v = vv[offset]\n        if v < vv_threshold:\n            offset_top = offset\n            break\n            \n    if isinstance(debug, np.ndarray):\n        if isinstance(image, torch.Tensor):\n            debug[3].imshow(image.cpu().numpy())\n#             vv = vv.cpu().numpy()\n#             vv_argmax = vv_argmax.cpu().numpy()\n#             vv_threshold = vv_threshold.cpu().numpy()\n        else:\n            debug[3].imshow(image)\n        debug[3].set_title('Y Top Offset')\n        vv_scale = W / vv.max() * 0.90\n        # Values\n        debug[3].plot(vv * vv_scale, np.arange(H) , c='red', label='vv')\n        # Threshold\n        debug[3].vlines(vv_threshold * vv_scale, 0, H -1, colors='orange', label='threshold')\n        # Max Value\n        debug[3].scatter(vv[vv_argmax] * vv_scale, vv_argmax, c='blue', s=100, label='Max', zorder=np.PINF)\n        # First Column Below Threshold\n        debug[3].scatter(vv[offset_top] * vv_scale, offset_top, c='purple', s=100, label='Offset', zorder=np.PINF)\n        debug[2].set_ylim(H, 0)\n        debug[3].legend()\n        debug[3].axis('off')\n            \n    return max(0, offset_bottom - margin), min(image.shape[0], offset_top + margin)\n\n# Crop image and pad offsets to target image height/width ratio to preserve information\ndef crop(image, size=None, debug=False, util=np):\n    # Image dimensions\n    H, W = image.shape\n    # Compute x/bottom/top offsets\n    x_offset = get_x_offset(image, debug=debug, util=util)\n    offset_bottom, offset_top = get_y_offsets(image[:,:x_offset], debug=debug, util=util)\n    # Crop Height and Width\n    h_crop = offset_top - offset_bottom\n    w_crop = x_offset\n    \n    # Pad crop offsets to target aspect ratio\n    if size is not None:\n        # Height too large, pad x offset\n        if (h_crop / w_crop) > TARGET_HEIGHT_WIDTH_RATIO:\n            x_offset += int(h_crop / TARGET_HEIGHT_WIDTH_RATIO - w_crop)\n        else:\n            # Height too small, pad bottom/top offsets\n            offset_bottom -= int(0.50 * (w_crop * TARGET_HEIGHT_WIDTH_RATIO - h_crop))\n            offset_bottom_correction = max(0, -offset_bottom)\n            offset_bottom += offset_bottom_correction\n\n            offset_top += int(0.50 * (w_crop * TARGET_HEIGHT_WIDTH_RATIO - h_crop))\n            offset_top += offset_bottom_correction\n        \n    # Crop Image\n    image = image[offset_bottom:offset_top:,:x_offset]\n        \n    return image","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:18:46.529179Z","iopub.execute_input":"2023-02-19T12:18:46.529729Z","iopub.status.idle":"2023-02-19T12:18:46.561206Z","shell.execute_reply.started":"2023-02-19T12:18:46.529695Z","shell.execute_reply":"2023-02-19T12:18:46.560326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# DALI","metadata":{}},{"cell_type":"code","source":"from nvidia.dali import fn, math, pipeline_def, types\nfrom nvidia.dali.plugin.pytorch import feed_ndarray as feed_ndarray\n\n@pipeline_def\ndef jpeg_pipeline():\n    jpegs = fn.external_source(device=\"cpu\", name=\"jpeg\")\n    images = fn.experimental.decoders.image(jpegs, device='mixed', output_type=types.ANY_DATA, dtype=types.UINT16)\n    images = fn.cast(images, dtype=types.FLOAT)\n    return images\n\n\ndef read_encoded_stream(filename):\n    dcmfile = pydicom.dcmread(filename)   \n    if dcmfile.file_meta.TransferSyntaxUID == '1.2.840.10008.1.2.4.90':\n        offset = dcmfile.PixelData.find(b\"\\x00\\x00\\x00\\x0C\")   #<---- the jpeg2000 header\n    else:\n        offset = dcmfile.PixelData.find(b\"\\xff\\xd8\") #<---- the jpeg lossless header    \n\n    buff = np.array(bytearray(dcmfile.PixelData[offset:]), dtype=np.uint8)\n    return buff\n\ndef decode_jpg(file_path, dicom, pipe):\n    buff = read_encoded_stream(file_path)\n    pipe.feed_input(\"jpeg\", [buff])\n    out = pipe.run()\n    dali_img = out[0][0]\n\n    image = torch.empty(dali_img.shape(), dtype=torch.float, device=\"cuda\")\n    feed_ndarray(dali_img, image, cuda_stream=torch.cuda.current_stream(device=0))\n\n    return image.squeeze()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:18:46.562986Z","iopub.execute_input":"2023-02-19T12:18:46.564163Z","iopub.status.idle":"2023-02-19T12:18:48.268089Z","shell.execute_reply.started":"2023-02-19T12:18:46.564101Z","shell.execute_reply":"2023-02-19T12:18:48.267036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Process Image","metadata":{}},{"cell_type":"code","source":"def process(file_path, size=(TARGET_WIDTH, TARGET_HEIGHT), crop_image=CROP_IMAGE, apply_clahe=APPLY_CLAHE, apply_eq_hist=APPLY_EQ_HIST, debug=False, save=True, pipe=None):\n    gpu = pipe is not None\n    # Read Dicom File\n    dicom = dicomsdl.open(file_path)\n    if gpu:\n        image = decode_jpg(file_path, dicom, pipe)\n    else:\n        image = dicom.pixelData()\n\n    if gpu:\n        util = torch\n    else:\n        util = np\n    \n    # Save original image for debug purposes\n    if debug:\n        fig, axes = plt.subplots(1, 5, figsize=(20,10))\n        if gpu:\n            image0 = image.cpu().numpy()\n        else:\n            image0 = np.copy(image)\n        axes[0].imshow(image0)\n        axes[0].set_title('Original Image')\n        axes[0].axis('off')\n    else:\n        axes = False\n    \n    # voi_lut\n    try:\n        image = voi_lut(image, dicom, util)\n    except:\n        pass\n    \n    # Some images have 0 values as highest intensity and need to be inverted\n    if dicom.getPixelDataInfo()['PhotometricInterpretation'] == 'MONOCHROME1':\n        image = util.max(image) - image\n\n    # Normalize [0,1] range\n    image = (image - image.min()) / (image.max() - image.min())\n\n    # Convert to uint8 image in range [0, 255]\n    if hasattr(image,'astype'):\n        image = (image * 255).astype(util.uint8)\n    else:\n        image = (image * 255).type(util.uint8).type(util.float)\n    \n    # Flip T0 Left/Right Orientation\n    h0, w0 = image.shape\n    if image[:,int(-w0 * 0.10):].sum() > image[:,:int(w0 * 0.10)].sum():\n        image = util.flip(image, (1,))\n    \n    # Crop Image\n    if crop_image:\n        image = crop(image, debug=axes, util=util)\n        \n    # Resize\n    if size is not None:\n        # Pad black pixels to make square image\n        h, w = image.shape\n        if (h / w) > TARGET_HEIGHT_WIDTH_RATIO:\n            pad = int(h / TARGET_HEIGHT_WIDTH_RATIO - w)\n            if hasattr(util, 'pad'):\n                image = util.pad(image, [[0,0], [0, pad]])\n            else:\n                image = F.pad(image,(0,pad))\n            h, w = image.shape\n        else:\n            pad = int(0.50 * (w * TARGET_HEIGHT_WIDTH_RATIO - h))\n            if hasattr(util, 'pad'):\n                image = util.pad(image, [[pad, pad], [0,0]])\n            else:\n                image = F.pad(image,(pad,pad,0,0))\n            h, w = image.shape\n        # Resize\n        if gpu:\n            # The result of torch's interpolate and cv2 are slightly different.\n            # If you take this change serious,\n            # you may want to convert image to numpy ndarray before resizing.\n            image = F.interpolate(image.view(1, 1, h, w), (size[1],size[0]), mode=\"area\")[0, 0]\n        else:\n            image = cv2.resize(image, size, interpolation=cv2.INTER_AREA)\n\n    if gpu:\n        image = image.type(util.uint8).cpu().numpy()\n        \n    # Apply CLAHE contrast enhancement\n    if apply_clahe:\n        image = CLAHE.apply(image)\n        \n     # Apply Histogram Equalization\n    if apply_eq_hist:\n        image = cv2.equalizeHist(image)\n        \n    # Show Processed Image    \n    if debug:\n        axes[4].imshow(image)\n        axes[4].set_title('Processed Image')\n        axes[4].axis('off')\n        plt.show()\n        \n    # Save Only\n    if save:\n        image_id = file_path.split('/')[-1].split('.')[0]\n        if IMAGE_FORMAT == 'png':\n            cv2.imwrite(f'{image_id}.png', image)\n        else:\n            cv2.imwrite(f'{image_id}.jpg', image, [cv2.IMWRITE_JPEG_QUALITY, 95])","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:25:58.891892Z","iopub.execute_input":"2023-02-19T12:25:58.892351Z","iopub.status.idle":"2023-02-19T12:25:58.91093Z","shell.execute_reply.started":"2023-02-19T12:25:58.892314Z","shell.execute_reply":"2023-02-19T12:25:58.909739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Example Preprocessing","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\n    \ndef get_file_path(args):\n    patient_id, image_id = args\n    return f'/kaggle/input/rsna-breast-cancer-detection/train_images/{patient_id}/{image_id}.dcm'\n    \ntrain['file_path'] = train[['patient_id', 'image_id']].apply(get_file_path, axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:18:48.29622Z","iopub.execute_input":"2023-02-19T12:18:48.296593Z","iopub.status.idle":"2023-02-19T12:18:48.856696Z","shell.execute_reply.started":"2023-02-19T12:18:48.296561Z","shell.execute_reply":"2023-02-19T12:18:48.855724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# CPU\n\nN = 8\n\nfor fp in tqdm(train['file_path'].head(N)):\n    process(fp, crop_image=True, size=(TARGET_WIDTH, TARGET_HEIGHT), debug=True, save=False)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:22:32.657602Z","iopub.execute_input":"2023-02-19T12:22:32.657994Z","iopub.status.idle":"2023-02-19T12:23:03.324807Z","shell.execute_reply.started":"2023-02-19T12:22:32.657962Z","shell.execute_reply":"2023-02-19T12:23:03.323986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# GPU\n\nN = 8\n\np = jpeg_pipeline(batch_size=1, num_threads=2, device_id=0, prefetch_queue_depth=1)\np.build()\n\nfor fp in tqdm(train['file_path'].head(N)):\n    process(fp, crop_image=True, size=(TARGET_WIDTH, TARGET_HEIGHT), debug=True, save=False, pipe=p)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:26:04.878585Z","iopub.execute_input":"2023-02-19T12:26:04.87895Z","iopub.status.idle":"2023-02-19T12:26:29.542742Z","shell.execute_reply.started":"2023-02-19T12:26:04.878919Z","shell.execute_reply":"2023-02-19T12:26:29.541688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# GPU Image Processing","metadata":{}},{"cell_type":"code","source":"test = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/test.csv')\n\ndef get_file_path(args):\n    patient_id, image_id = args\n    return f'/kaggle/input/rsna-breast-cancer-detection/test_images/{patient_id}/{image_id}.dcm'\n    \ntest['file_path'] = test[['patient_id', 'image_id']].apply(get_file_path, axis=1)\n\ndisplay(test.info())\ndisplay(test.head())","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:21:13.919357Z","iopub.execute_input":"2023-02-19T12:21:13.919792Z","iopub.status.idle":"2023-02-19T12:21:13.970649Z","shell.execute_reply.started":"2023-02-19T12:21:13.919757Z","shell.execute_reply":"2023-02-19T12:21:13.969765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\np = jpeg_pipeline(batch_size=1, num_threads=2, device_id=0, prefetch_queue_depth=1)\np.build()\n\nfor fp in tqdm(test['file_path']):\n    try:\n        process(fp,pipe=p)\n    except Exception as e:\n        print('Fall back to CPU')\n        process(fp)\n        p = jpeg_pipeline(batch_size=1, num_threads=2, device_id=0, prefetch_queue_depth=1)\n        p.build()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:21:13.974597Z","iopub.execute_input":"2023-02-19T12:21:13.976719Z","iopub.status.idle":"2023-02-19T12:21:15.103573Z","shell.execute_reply.started":"2023-02-19T12:21:13.976682Z","shell.execute_reply":"2023-02-19T12:21:15.102458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model","metadata":{}},{"cell_type":"code","source":"def normalize(image):\n    # Repeat channels to create 3 channel images required by pretrained ConvNextV2 models\n    image = tf.repeat(image, repeats=3, axis=3)\n    # Cast to float 32\n    image = tf.cast(image, tf.float32)\n    # Normalize with respect to ImageNet mean/std\n    image = tf.keras.applications.imagenet_utils.preprocess_input(image, mode='torch')\n\n    return image","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:21:15.106474Z","iopub.execute_input":"2023-02-19T12:21:15.107556Z","iopub.status.idle":"2023-02-19T12:21:15.113794Z","shell.execute_reply.started":"2023-02-19T12:21:15.107515Z","shell.execute_reply":"2023-02-19T12:21:15.112536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_model():\n    # Inputs, note the names are equal to the dictionary keys in the dataset\n    image = tf.keras.layers.Input(INPUT_SHAPE, name='image', dtype=tf.uint8)\n\n    # Normalize Input\n    image_norm = normalize(image)\n\n    # CNN Feature Maps\n    x = convnext.ConvNeXtV2Tiny(\n        input_shape=(TARGET_HEIGHT, TARGET_WIDTH, 3),\n        pretrained=None,\n        num_classes=0,\n    )(image_norm)\n\n    # Average Pooling BxHxWxC -> BxC\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    # Dropout to prevent Overfitting\n    x = tf.keras.layers.Dropout(0.30)(x)\n    # Output value between [0, 1] using Sigmoid function\n    outputs = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n\n    # Define model with inputs and outputs\n    model = tf.keras.models.Model(inputs=image, outputs=outputs)\n\n    # Load pretrained Model Weights\n    model.load_weights('/kaggle/input/rsna-efficientnetv2-training-tensorflow-tpu-ds/model.h5')\n\n    # Set model non-trainable\n    model.trainable = False\n\n    # Compile model\n    model.compile()\n\n    return model","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:21:15.115405Z","iopub.execute_input":"2023-02-19T12:21:15.116461Z","iopub.status.idle":"2023-02-19T12:21:15.132367Z","shell.execute_reply.started":"2023-02-19T12:21:15.116425Z","shell.execute_reply":"2023-02-19T12:21:15.131194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Pretrained File Path: '/kaggle/input/sartorius-training-dataset/model.h5'\ntf.keras.backend.clear_session()\n# enable XLA optmizations\ntf.config.optimizer.set_jit(True)\n\nmodel = get_model()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:21:15.133929Z","iopub.execute_input":"2023-02-19T12:21:15.134744Z","iopub.status.idle":"2023-02-19T12:21:19.783854Z","shell.execute_reply.started":"2023-02-19T12:21:15.134708Z","shell.execute_reply":"2023-02-19T12:21:19.782875Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot model summary\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:21:19.78541Z","iopub.execute_input":"2023-02-19T12:21:19.785778Z","iopub.status.idle":"2023-02-19T12:21:19.799394Z","shell.execute_reply.started":"2023-02-19T12:21:19.785741Z","shell.execute_reply":"2023-02-19T12:21:19.798273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Sample Submission","metadata":{}},{"cell_type":"code","source":"sample_submission = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/sample_submission.csv')\n\ndisplay(sample_submission.info())\ndisplay(sample_submission.head())","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:21:19.801295Z","iopub.execute_input":"2023-02-19T12:21:19.801698Z","iopub.status.idle":"2023-02-19T12:21:19.826827Z","shell.execute_reply.started":"2023-02-19T12:21:19.801662Z","shell.execute_reply":"2023-02-19T12:21:19.825773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Inference","metadata":{}},{"cell_type":"code","source":"SUBMISSION_ROWS = []\n# Iterate over all patient_id/laterality combinations groups\nfor idx, ((patient_id, laterality), g) in enumerate(tqdm(test.groupby(['patient_id', 'laterality']))):\n    # Cancer target is mean of predicted cancer values\n    cancer = 0\n    # Iterate over all scans in group\n    for row_idx, row in g.iterrows():\n        # Load Image\n        image_id = row['image_id']\n        image = cv2.imread(f'{image_id}.{IMAGE_FORMAT}', -1)\n        # Show First Few Images\n        if idx < 16:\n            plt.figure(figsize=(5,8))\n            plt.imshow(image)\n            plt.show()\n        \n        # Expand to Batch HxW -> 1xHxWx1\n        image = np.expand_dims(image, [0, 3])\n        # Make Prediction\n        cancer += model.predict_on_batch(image).squeeze() / len(g)\n        # Remove Image\n        os.remove(f'{image_id}.{IMAGE_FORMAT}')\n        \n    # Add Submission Row\n    SUBMISSION_ROWS.append({\n        'prediction_id': f'{patient_id}_{laterality}',\n        'cancer': np.int8(cancer > THRESHOLD_BEST),\n    })\n    \n    if np.random.rand() > 0.99:\n        gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:21:19.829031Z","iopub.execute_input":"2023-02-19T12:21:19.829969Z","iopub.status.idle":"2023-02-19T12:21:29.14131Z","shell.execute_reply.started":"2023-02-19T12:21:19.829933Z","shell.execute_reply":"2023-02-19T12:21:29.14029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Save Submission","metadata":{}},{"cell_type":"code","source":"# Create DataFrame from submission rows\nsubmission_df = pd.DataFrame(SUBMISSION_ROWS)\n\ndisplay(submission_df.info())\ndisplay(submission_df.head())","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:21:29.144556Z","iopub.execute_input":"2023-02-19T12:21:29.145285Z","iopub.status.idle":"2023-02-19T12:21:29.166624Z","shell.execute_reply.started":"2023-02-19T12:21:29.145243Z","shell.execute_reply":"2023-02-19T12:21:29.165508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save submission as CSV\nsubmission_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:21:29.168199Z","iopub.execute_input":"2023-02-19T12:21:29.16863Z","iopub.status.idle":"2023-02-19T12:21:29.175746Z","shell.execute_reply.started":"2023-02-19T12:21:29.168593Z","shell.execute_reply":"2023-02-19T12:21:29.175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Sanity Check\ndisplay(pd.read_csv('submission.csv').head())","metadata":{"execution":{"iopub.status.busy":"2023-02-19T12:21:29.177226Z","iopub.execute_input":"2023-02-19T12:21:29.177886Z","iopub.status.idle":"2023-02-19T12:21:29.189072Z","shell.execute_reply.started":"2023-02-19T12:21:29.177847Z","shell.execute_reply":"2023-02-19T12:21:29.188156Z"},"trusted":true},"execution_count":null,"outputs":[]}]}