{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Problem Description <a class=\"anchor\" id=\"Problem Description\"></a>\n   * **problem** : In this competition our goal is to predict the presence or absence of cancer in mammography images. \n   * **helped Notebooks**: in our journey i will use some hopefull ideas from other notebook and i will mention them in description \n   * **Tasks we will cover in this section** \n       1. Per-Processing images \n          - understand the data \n          - explore data from diffrent view perspective \n          - Technic to processes data to feed into the mode \n               1. read images in each case Patient_ID \n               2. Resize the image and Crop the ROI ( region of intersted \n               3. save the process the image in npy format \n               4. extrat the image label from Train.Csv file \n               5. Virtualize few sample \n   ","metadata":{}},{"cell_type":"markdown","source":"<div class=\" alert alert-info\"> <strong>Purpose</strong>\n    Following this notebook is about desmonstration of Bullding Fourire Transformation Layer \n</div>\n <div class=\" alert alert-warning\"> <strong>warining </strong>\n    may the resutlts not satisciaed the goal of the competition \n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"\"","metadata":{}},{"cell_type":"code","source":"%%capture\n!pip install python-gdcm\n!pip install pylibjpeg","metadata":{"execution":{"iopub.status.busy":"2023-03-02T17:23:14.478011Z","iopub.execute_input":"2023-03-02T17:23:14.478495Z","iopub.status.idle":"2023-03-02T17:23:41.173Z","shell.execute_reply.started":"2023-03-02T17:23:14.478395Z","shell.execute_reply":"2023-03-02T17:23:41.171571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from IPython.display import display_html\ndef restartkernel() :\n    display_html(\"<script>Jupyter.notebook.kernel.restart()</script>\",raw=True)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:41:54.219187Z","iopub.execute_input":"2023-02-20T02:41:54.219646Z","iopub.status.idle":"2023-02-20T02:41:54.25859Z","shell.execute_reply.started":"2023-02-20T02:41:54.219557Z","shell.execute_reply":"2023-02-20T02:41:54.257534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import platform, sys, pydicom\n\nprint(\n    platform.platform(),\n    \"\\nPython\", sys.version,\n    \"\\npydicom\", pydicom.__version__\n)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:44:56.015217Z","iopub.execute_input":"2023-02-20T02:44:56.015664Z","iopub.status.idle":"2023-02-20T02:44:56.141943Z","shell.execute_reply.started":"2023-02-20T02:44:56.015582Z","shell.execute_reply":"2023-02-20T02:44:56.140914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### import nesecccery packages \nimport matplotlib.pyplot as plt \nimport plotly.express as px\nfrom plotly.subplots import make_subplots\nimport plotly.graph_objects as go\nfrom tqdm.notebook import tqdm\n\nimport numpy as np \nimport math\nfrom collections.abc import Iterable\nimport types\n\nimport pydicom as dc \nimport pandas as pd \nimport cv2\n\nfrom pathlib import Path\nimport os \n\nimport torch.nn as nn\nimport torch\nimport torchvision\nfrom torchvision import transforms\nimport torchmetrics\nimport pytorch_lightning as pl\nfrom pytorch_lightning.callbacks import ModelCheckpoint\nfrom pytorch_lightning.loggers import TensorBoardLogger\n\n#from pytorch_lightning.plugins.training_type.ddp import DDPPlugin\nimport warnings \ndef fxn():\n    warnings.warn(\"deprecated\", DeprecationWarning)\n\nwith warnings.catch_warnings():\n    warnings.simplefilter(\"ignore\")\n    fxn()\n\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-03-02T17:10:40.873886Z","iopub.execute_input":"2023-03-02T17:10:40.874923Z","iopub.status.idle":"2023-03-02T17:10:40.885694Z","shell.execute_reply.started":"2023-03-02T17:10:40.874874Z","shell.execute_reply":"2023-03-02T17:10:40.884266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Exploer data Csv file** <a class=\"anchor\" id =\"Exploer data Csv file\"></a>\n- we will read the file csv file and check if the data is unbalane corresponding Class we have and let us clean and remove informations thses wil not use in our task \n    1. remvoe unecessery Columns \n    2. check the Class balancing \n    3. read few samples from data \n       - **Notation :** keep in mind here each patient has own series diagnois Slices Image Dicom File followed by [Case_study] **-->** [Slices-Images]\n\n     ","metadata":{}},{"cell_type":"code","source":"Path_data = \"../input/rsna-breast-cancer-detection/train.csv\"\ndata_df = pd.read_csv(Path_data)\n### let us get the number samples we have \nnumber_samples =len(data_df.patient_id)\nlist_col = list(data_df.columns)\nprint(f\"number of total samples {number_samples }\\n Columns dataFrame {list_col}\")","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.048543Z","iopub.status.idle":"2023-02-19T18:40:03.04933Z","shell.execute_reply.started":"2023-02-19T18:40:03.049064Z","shell.execute_reply":"2023-02-19T18:40:03.049089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.0509Z","iopub.status.idle":"2023-02-19T18:40:03.051707Z","shell.execute_reply.started":"2023-02-19T18:40:03.051434Z","shell.execute_reply":"2023-02-19T18:40:03.051459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\nNotaion : we will remove the some columes thses will not use to Pipline models and Per-Processing Step \n-  ['site_id','laterality',\n'view', 'age', 'biopsy',\n'invasive', 'BIRADS',\n'implant','machine_id',\n'difficult_negative_case']\n\"\"\"\ndf_Cols = data_df.drop(['site_id','laterality',\n'view', 'age', 'biopsy',\n'invasive', 'BIRADS',\n'implant','machine_id',\n'difficult_negative_case'],axis=1)\ndata_df_col = df_Cols.fillna(0) ## replace Nan with zeroes\ndata_df_col.head(6)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.05316Z","iopub.status.idle":"2023-02-19T18:40:03.053975Z","shell.execute_reply.started":"2023-02-19T18:40:03.053714Z","shell.execute_reply":"2023-02-19T18:40:03.053739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### check labels Counts between data distrubtion \nTemp_before= data_df_col[\"cancer\"].value_counts()\nsave_count_before = pd.DataFrame({ \"Class\":Temp_before.index,\n                              \"Value\": Temp_before.values})\n## let us figure out solution how to balance the data \n\"\"\"\nsolution 1 : we will try to remove some samples from data that has class 0 \nand make the samples equale class 0 == class 1 \nNotation : this will cause a problem is reducing the number\nsamples Low data repesentation to feed into the model to lean from \n\"\"\"\nsub_stract_classes = save_count_before.Value.iloc[0] - save_count_before.Value.iloc[1]\nprint(sub_stract_classes)\n### now will slice from the results substraction we got \nsave_count_before.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.055605Z","iopub.status.idle":"2023-02-19T18:40:03.056481Z","shell.execute_reply.started":"2023-02-19T18:40:03.056181Z","shell.execute_reply":"2023-02-19T18:40:03.056206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alter-","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-warning\" role=\"alert\"> <strong>Important :</strong> here we can see that Class Cancer {0} has a huge amount of data samples greater than Class Cancer {1} which means , insteat we will need to balance the data using Technics such Augmentation that i will not cover in this notebook i wil explore this in another note  <br>\nCancer 97 % <br>\nno Cancer 2.7 %     \n</div>","metadata":{}},{"cell_type":"code","source":"px.pie(save_count_before, values=\"Value\", names=\"Class\", title='Cancer distribution before')","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.057998Z","iopub.status.idle":"2023-02-19T18:40:03.058789Z","shell.execute_reply.started":"2023-02-19T18:40:03.058514Z","shell.execute_reply":"2023-02-19T18:40:03.058538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-info\" role=\"alert\"> <strong>Consideration need to keep in mind</strong>:\n    <ul>\n        <li><strong>First:</strong> if we will balancing the data distrubtion between the labele that will lead to decreas the number of samples at end the mode may couldn't to have enough feature</li>\n        <li><strong>Second:</strong> the Images isn't localize to yet we will need to augment and process data to make it better repesentation to feed to Model</li>\n    </ul>\n    </div>\n","metadata":{}},{"cell_type":"markdown","source":"#### ","metadata":{}},{"cell_type":"code","source":"#### now we will need to sorte the data by cancer Column \ndata_sorted = data_df_col.sort_values(by=\"cancer\")\nbalanced_df = data_sorted[52390:]","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.060275Z","iopub.status.idle":"2023-02-19T18:40:03.061147Z","shell.execute_reply.started":"2023-02-19T18:40:03.060848Z","shell.execute_reply":"2023-02-19T18:40:03.060873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Temp_after= balanced_df[\"cancer\"].value_counts()\nsave_count_after= pd.DataFrame({ \"Class\":Temp_after.index,\n                              \"Value\": Temp_after.values})\nsave_count_after.head(2)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.06272Z","iopub.status.idle":"2023-02-19T18:40:03.063575Z","shell.execute_reply.started":"2023-02-19T18:40:03.063278Z","shell.execute_reply":"2023-02-19T18:40:03.063303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = make_subplots(rows=1, cols=2)\nLabel = [\"no cancer\", \"cancer\"]\ncolors = ['lightslategray',] * 2\ncolors[1] = 'crimson'\nfig.add_trace(\n    go.Bar(x=Label, y=save_count_after.Value, \n           marker_color=colors,\n           name=\"After balancing Labels\"),\n    row=1, col=1\n)\n\nfig.add_trace(\n    go.Bar(x=Label, y=save_count_before.Value,\n            marker_color=colors,\n            name=\"Before balancing Labels\"),\n    row=1, col=2\n)\n\nfig.update_layout(height=500, width=700, title_text=\"Cancer distribution balancing \", showlegend=True)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.065012Z","iopub.status.idle":"2023-02-19T18:40:03.065866Z","shell.execute_reply.started":"2023-02-19T18:40:03.065576Z","shell.execute_reply":"2023-02-19T18:40:03.065619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Total_samples=len(balanced_df.patient_id)\nprint(f\"the total samples we have now : {Total_samples}\")","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.067473Z","iopub.status.idle":"2023-02-19T18:40:03.068372Z","shell.execute_reply.started":"2023-02-19T18:40:03.068085Z","shell.execute_reply":"2023-02-19T18:40:03.06811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## set the seed \nnp.random.seed(2000)\n### here will read few samples from dataset \nPath_train = \"../input/rsna-breast-cancer-detection/train_images/\"\nfig , axis = plt.subplots(2,2,figsize=(5,5))\nrand_index = np.random.randint([900,2316],size=(1,)).item()\nassing_indx = rand_index\nfor row in range(2):\n    for col in range(2):\n        ## get the full path Imae\n        sub_path = str(balanced_df.patient_id.iloc[assing_indx]) + \"/\" + str(balanced_df.image_id.iloc[assing_indx])\n        Label = balanced_df.cancer.iloc[assing_indx]\n        Image_path = os.path.join(Path_train,sub_path) + \".dcm\"\n        Image_array = dc.read_file(Image_path).pixel_array \n        axis[row][col].imshow(Image_array,cmap=\"bone\")\n        if Label == 0:\n            axis[row][col].set_title(\"Negative\")\n        else:\n            axis[row][col].set_title(\"Positive\")\n    assing_indx += 1     \n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-02-19T18:40:03.069738Z","iopub.status.idle":"2023-02-19T18:40:03.07053Z","shell.execute_reply.started":"2023-02-19T18:40:03.070246Z","shell.execute_reply":"2023-02-19T18:40:03.070269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Per-Processing Data**   \n   - **Pipline Steps**\n        1. read images in each case Patient_ID\n        2. Resize the image and Crop the ROI ( region of intersted )\n        3. save the process the image in npy format\n        4. extrat the image label from Train.Csv file\n        5. Virtualize few sample \n        \n   <div class=\"alert alert-warning\" role=\"alert\"> <strong>Hopefull Notebook</strong>:\n    <ul>\n        <li><strong>Resize AND Crop methods </strong> Notebook link: <a  class=\"anchor\" href=\"https://www.kaggle.com/code/snnclsr/roi-extraction-using-opencv\"  id=\"Per-Processing\"> Sinan Calisir </a> Please Upvote</li>\n        <li><strong>DICOM images to PNGs:</strong> Notebook link : <a  class=\"anchor\" href=\"https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\"  id=\"Per-Processing\"> Radek Osmulski </a>  Please Upvote </li>\n    </ul>\n    </div>\n     \n","metadata":{}},{"cell_type":"markdown","source":"### **how to work with Dicom file and understand it better**   \n   - **in this section we will try to go through how to handle dicom files**\n     1. get the tags dicomConfiguration and use important tags \n     2. Rescaling algorithm <br>\n        **Note** The following is an algorithm I did google and i grabed some Stackover flow posts, some mentions in the DICOM specs.\n        1. Rescale type attribute\n        2. Window Center\n        3. RescaleIntercept \n        \n<div class=\"alert alert-light\" role=\"alert\"> <strong>dicomConfiguration  :</strong><br>\n<strong>note : Photometric Interpretation tages relate to data we have</strong><br>\n    \nMONOCHROME1 : indicates that the greyscale ranges from bright to dark with ascending pixel values,<br>\nMONOCHROME2 : ranges from dark to bright with ascending pixel values.\nImage compression is independent of the Photometric Interpretation (0028,0004) attribute. Compression is instead given by the Transfer Syntax UID (0002,0010) attribute <br> <br>\n\n<div class=\"alert alert-warning\" role=\"alert\"> <strong>Hopefull Notebook</strong>:\n    <ul>\n        <li><strong>Preprocessing images to normalize </strong> Notebook link: <a  class=\"anchor\" href=\"https://www.kaggle.com/code/donkeys/preprocessing-images-to-normalize-colors-and-sizes\"  id=\"Per-Processing\"> averagemn  </a> Please Upvote</li>\n        <li><strong>Article Meduim :</strong> Article link : <a  class=\"anchor\" href=\"https://towardsdatascience.com/understanding-dicoms-835cd2e57d0b\"  id=\"Per-Processing\"> Radek Osmulski </a>  read and enjoy </li>\n    </ul>    \n    </div>\n\n\n<strong>Here are some key points on the tag information above:</strong> <br>\n\n**Pixel Data (7fe0 0010)(last entry)** : This is where the raw pixel data is stored. The order of pixels encoded for each image plane is left to right, top to bottom, i.e., the upper left pixel (labeled 1,1) is encoded first <br>\n**Photometric Interpretation (0028, 0004)** : aka color space. In this case it is MONOCHROME2 where pixel data is represented as a single monochrome image plane where the minimum sample value is intended to be displayed as black info <br>\n**Samples per Pixel (0028, 0002)** : This should be 1 as this image is monochrome. This value would be 3 if the color space was RGB for example <br>\n**Bits Stored (0028 0101)** : Number of bits stored for each pixel sample <br>\n**Pixel Represenation (0028 0103)** : can either be unsigned(0) or signed(1) <br>\n**Lossy Image Compression (0028 2110)** : 00 image has not been subjected to lossy compression. 01 image has been subjected to lossy compression. <br>\n**Lossy Image Compression Method (0028 2114)** : states the type of lossy compression used (in this case JPEG Lossy Compression, as denotated by CS:ISO_10918_1)   <br>        \n\n    ","metadata":{}},{"cell_type":"code","source":"balanced_df.head(2)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.071962Z","iopub.status.idle":"2023-02-19T18:40:03.072789Z","shell.execute_reply.started":"2023-02-19T18:40:03.072517Z","shell.execute_reply":"2023-02-19T18:40:03.072542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_image_config_dicom(id_iloc,path_train):\n    data_id_= balanced_df.iloc[id_iloc]\n    patient_path = str(data_id_.patient_id ) +\"/\"+ str(data_id_.image_id)\n    full_path = os.path.join(path_train,patient_path) + \".dcm\"\n    ds = dc.read_file(full_path)\n    config_list = dir(ds)\n### let us get the dicom Configuration Tags\n    for attr in config_list:\n        if attr.startswith(\"_\"):\n            continue\n        if attr == \"PixelData\" or attr == \"pixel_array\":\n            #skip printing the long arrays as they will just spam the output too much with hex code\n            continue\n        var_type = type(getattr(ds,attr))\n        if var_type == types.MethodType:\n            continue\n    return full_path  , getattr(ds,attr)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.074166Z","iopub.status.idle":"2023-02-19T18:40:03.074917Z","shell.execute_reply.started":"2023-02-19T18:40:03.074677Z","shell.execute_reply":"2023-02-19T18:40:03.074699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def show_img(img_path, colormap = None, extra_brightness=0):\n    ds = dc.read_file(img_path)\n    shape = ds.pixel_array.shape\n    target = 255\n    # Convert to float to avoid overflow or underflow losses.\n    image_2d = ds.pixel_array.astype(float)\n    img_data = image_2d\n    print(f\"data min: {img_data.min()}, max: {img_data.max()}\")\n    print(f\"window center: {ds.WindowCenter}, rescale intercept: {ds.RescaleIntercept}\")\n    multival = isinstance(ds.WindowCenter, Iterable)\n    if multival:\n        scale_center = -ds.WindowCenter[0]\n    else:\n        scale_center = -ds.WindowCenter\n    intercept = scale_center+ds.RescaleIntercept+extra_brightness\n    print(f\"final intercept: {intercept}\")\n    #image_2d += intercept \n    print(f\"after applying intercept, min: {image_2d.min()}, max: {image_2d.max()}\")\n\n    # Rescaling grey scale between 0-255\n    image_2d_scaled = (np.maximum(image_2d,0) / image_2d.max()) * 255.0\n    print(f\"after scaling to 0-255, min: {image_2d_scaled.min()}, max: {image_2d_scaled.max()}\")\n\n    # Convert to uint\n    image_2d_scaled = np.uint8(image_2d_scaled)\n\n    plt.figure(figsize=(4,4))\n    plt.imshow(image_2d_scaled, cmap=colormap)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.076279Z","iopub.status.idle":"2023-02-19T18:40:03.077113Z","shell.execute_reply.started":"2023-02-19T18:40:03.076854Z","shell.execute_reply":"2023-02-19T18:40:03.076878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"full_path , config_dicom_file = get_image_config_dicom(2315,Path_train)\nprint(config_dicom_file)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.078384Z","iopub.status.idle":"2023-02-19T18:40:03.079145Z","shell.execute_reply.started":"2023-02-19T18:40:03.07889Z","shell.execute_reply":"2023-02-19T18:40:03.078914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from IPython.display import display_html\ndef restartkernel() :\n    display_html(\"<script>Jupyter.notebook.kernel.restart()</script>\",raw=True)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.080831Z","iopub.status.idle":"2023-02-19T18:40:03.08161Z","shell.execute_reply.started":"2023-02-19T18:40:03.081351Z","shell.execute_reply":"2023-02-19T18:40:03.081376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_img(full_path, colormap=plt.cm.bone) #image 1","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.083039Z","iopub.status.idle":"2023-02-19T18:40:03.083855Z","shell.execute_reply.started":"2023-02-19T18:40:03.083603Z","shell.execute_reply":"2023-02-19T18:40:03.083626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**NOTETTION :** after reseclaing the pixels data between 0-255 that will reduce huge computational cost of finding wieghts Paramaters through BackPropagations <br>\n**Next Step :** use that method of scaling data image to Resampling all data images ","metadata":{"execution":{"iopub.status.busy":"2022-12-03T17:24:55.408253Z","iopub.execute_input":"2022-12-03T17:24:55.408608Z","iopub.status.idle":"2022-12-03T17:24:55.415267Z","shell.execute_reply.started":"2022-12-03T17:24:55.408576Z","shell.execute_reply":"2022-12-03T17:24:55.413924Z"}}},{"cell_type":"code","source":"# https://www.kaggle.com/code/snnclsr/roi-extraction-using-opencv\ndef crop_coords(img):\n    \"\"\"\n    Crop ROI from image.\n    \"\"\"\n    # Otsu's thresholding after Gaussian filtering\n    blur = cv2.GaussianBlur(img, (5, 5), 0)\n    _, breast_mask = cv2.threshold(blur,0,255,cv2.THRESH_BINARY+cv2.THRESH_OTSU)\n    \n    cnts, _ = cv2.findContours(breast_mask.astype(np.uint8), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)\n    cnt = max(cnts, key = cv2.contourArea)\n    x, y, w, h = cv2.boundingRect(cnt)\n    return (x, y, w, h)\n\n\ndef normalization_(img):\n    \"\"\"\n    WindowCenter and normalize pixels in the breast ROI.\n    return: numpy array of the normalized image\n    \"\"\"\n    # Convert to float to avoid overflow or underflow losses.\n    image_2d = img.astype(float)\n    image_2d_scaled = (np.maximum(image_2d,0) / image_2d.max()) * 255.0\n    # Convert to uint\n    normalized = np.uint8(image_2d_scaled)\n    return normalized","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.085259Z","iopub.status.idle":"2023-02-19T18:40:03.086049Z","shell.execute_reply.started":"2023-02-19T18:40:03.085795Z","shell.execute_reply":"2023-02-19T18:40:03.085819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def Processing_img(Path_dir , label_csv , image_id):\n    data_id = label_csv.iloc[image_id]\n    image = str(data_id.patient_id)+ \"/\" + str(data_id.image_id) \n    Image_path = os.path.join(Path_dir,image) + \".dcm\" \n    read_image = dc.read_file(Image_path)\n    Array_pixel = read_image.pixel_array\n#     if read_image.PhotometricInterpretation == \"MONOCHROME1\":\n#         Array_pixel = np.amax(Array_pixel) - Array_pixel\n    (x,y,w,h) = crop_coords(Array_pixel)\n    image_crop = Array_pixel[y:y+h,x:x+w]\n    normalize_img =normalization_(image_crop)\n    # Resize the image to the final shape. \n    img_final = cv2.resize(normalize_img, (120, 120)).astype(np.float32)\n    return img_final , data_id","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.087566Z","iopub.status.idle":"2023-02-19T18:40:03.088324Z","shell.execute_reply.started":"2023-02-19T18:40:03.088075Z","shell.execute_reply":"2023-02-19T18:40:03.088098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# we will need to randome selecte the number rows for not keep it sorted \n# random_label_df = balanced_df.sample(n=2316,random_state =70,replace=True)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.089698Z","iopub.status.idle":"2023-02-19T18:40:03.090445Z","shell.execute_reply.started":"2023-02-19T18:40:03.090188Z","shell.execute_reply":"2023-02-19T18:40:03.090211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## set the clean dataFrame balanced_df\nsums = 0 \nsums_square = 0 ### to computer Normalize means and STD to use later \nNormalizer = 120*120\nProcessed_data = \"./Processed\"\nfor idx_img , Patient_ID in enumerate(tqdm(balanced_df.image_id)):\n    ### Extract label Patient \n    image_processed , label_id = Processing_img(Path_train,balanced_df,image_id=idx_img)\n    Image_id = str(label_id.image_id)\n    label = label_id.cancer\n    Train_or_Val = \"train\" if idx_img < 1800 else \"val\" ## total samples are 2316 we split it by \n    ## creat save folder \n    Save_path = Path(Processed_data + \"/\"+ Train_or_Val +\"/\"+ str(label))\n    Save_path.mkdir(parents=True,exist_ok=True)\n    #         ### the struct dircotry is \n#         ## Train/\n#         #        lable_0/\n#         #              Image_ID\n#         #        labek_1/\n#         #              Image_ID \n#         #####\n    np.save(os.path.join(Save_path,Image_id),image_processed)\n    if Train_or_Val == \"train\":\n        sums += np.sum(image_processed) / Normalizer \n        sums_square += (image_processed ** 2).sum() / Normalizer ","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.091808Z","iopub.status.idle":"2023-02-19T18:40:03.092628Z","shell.execute_reply.started":"2023-02-19T18:40:03.09228Z","shell.execute_reply":"2023-02-19T18:40:03.092303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mean = sums / 2316\nstd = np.sqrt(sums_square / 2316 - mean**2)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.094046Z","iopub.status.idle":"2023-02-19T18:40:03.094836Z","shell.execute_reply.started":"2023-02-19T18:40:03.094587Z","shell.execute_reply":"2023-02-19T18:40:03.09461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mean , std\n### we can see here the STD and Means not range between [0-1] \n### because of the pixle sixe is between 0-255","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.096513Z","iopub.status.idle":"2023-02-19T18:40:03.09727Z","shell.execute_reply.started":"2023-02-19T18:40:03.097023Z","shell.execute_reply":"2023-02-19T18:40:03.097046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Traning Step using Pytorch Lighting** <a class=\"anchor\" id=\"Section-2\"></a>\n* loading the data has been Processed \n  - function to load and save ans float \n     - augementation \n     - Dataloader \n     - model CNN from Renstnet","metadata":{}},{"cell_type":"code","source":"### Loading the numpy format as .npy\ndef Loader_Data(path):\n    return np.load(path).astype(np.float32)\n### Augmenetation pipline from transform Objet trochvision \n\nTransform_aug_Train = transforms.Compose([\n    transforms.ToTensor(),\n    transforms.RandomAffine(degrees=(-5, 5), translate=(0, 0.05), scale=(0.9, 1.1)),\n    ])\nTransform_aug_Val = transforms.Compose([\n    transforms.ToTensor(),\n    ])\nLoading_DataFolder_Train = torchvision.datasets.DatasetFolder(root=\"./Processed/train\",\n                                                              loader=Loader_Data , \n                                                              extensions=\"npy\",\n                                                              transform=Transform_aug_Train)\nLoading_DataFolder_Val= torchvision.datasets.DatasetFolder(root=\"./Processed/val\",\n                                                           loader=Loader_Data , \n                                                           extensions=\"npy\",\n                                                           transform=Transform_aug_Val)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.098814Z","iopub.status.idle":"2023-02-19T18:40:03.099845Z","shell.execute_reply.started":"2023-02-19T18:40:03.099516Z","shell.execute_reply":"2023-02-19T18:40:03.099549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#### Virtualize few samples from Data after augmenration .\nfig , axes = plt.subplots(2,2,figsize=(6,6))\ncounter_Img = 0\nfor i in range(2):\n    for j in range(2):\n        rand_indx = np.random.randint(0,500,size=(1,)).item() ### this will return value betweem [0,24000]\n        Image_ray , label = Loading_DataFolder_Train[rand_indx] # will retrun tuple correspond image X_ray woht label \n        if label == 0 :\n            class_ = \"Negative\"\n        else :\n            class_ = \"Positive\"\n        axes[i][j].imshow(Image_ray[0],cmap=\"gray\")\n        axes[i][j].set_title(f\"label class is : {class_}\")","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.101303Z","iopub.status.idle":"2023-02-19T18:40:03.102131Z","shell.execute_reply.started":"2023-02-19T18:40:03.101872Z","shell.execute_reply":"2023-02-19T18:40:03.101896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"##### check the data if is imblance or not by seeing the distrubtion \nnp.unique(Loading_DataFolder_Train.targets, return_counts=True)\n### here we can see that the data is imblanc but we can go through this even do \n","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.10358Z","iopub.status.idle":"2023-02-19T18:40:03.104636Z","shell.execute_reply.started":"2023-02-19T18:40:03.104285Z","shell.execute_reply":"2023-02-19T18:40:03.104317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Pipline model** Building FourireTransformer <a class=\"anchor\" id=\"Section3\"></a>\n- by calling this API torch \n  * Building Fourier Layer Transformation \n    - **Notation:** the whole Point of this Layer is Transform the Features of Data into Frequnecies domain and Keep Extraction Features be produced on that domain also by **Multiply the Dimenssion of Image with Learnbel Wieghts**\n   \n **Notation** :  <div class=\"alert alert-info\">Explain process <strong>the Main Key idea </strong>\n    in Spetial Domain there's some risk to loss information During Convolution Operation which lead into Unstablized Features map . in my EX i tried to transform the Features into Frequency Domain which will improve some weaknees comes with convolution:\n    1. reduce complexity \n    2. fast traning per Epoch \n    2. Throughput GPU and Latency \n    4. Learning from Different MainField Data Repersentation\n\n</div> <br>\n<div class=\"alert alert-warning\"><strong>\n    that's why we will just run the Model on 10 Epochs and you can raise it to wished value if you would like , i will improve this notebook in the next few days</strong></div>\n\n","metadata":{}},{"cell_type":"code","source":"import torch.nn as nn\nimport torch\nimport torch.nn.functional as F","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.113479Z","iopub.status.idle":"2023-02-19T18:40:03.114541Z","shell.execute_reply.started":"2023-02-19T18:40:03.114164Z","shell.execute_reply":"2023-02-19T18:40:03.114195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Config = {\n  'batch_size': 4 ,\n  'num_workers':2,\n    'lr':3e-4\n    \n}","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.116255Z","iopub.status.idle":"2023-02-19T18:40:03.117015Z","shell.execute_reply.started":"2023-02-19T18:40:03.116766Z","shell.execute_reply":"2023-02-19T18:40:03.116789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"batch_size = 4\nTraining_Loader = torch.utils.data.DataLoader(Loading_DataFolder_Train,batch_size=Config['batch_size'],num_workers=Config['num_workers'], shuffle=True)\nValidation_Loader = torch.utils.data.DataLoader(Loading_DataFolder_Val,batch_size=Config['batch_size'],num_workers=Config['num_workers'],shuffle=False)\nprint(f\"lenght of training data is :{len(Training_Loader)}\\nlenght of validationi s: {len(Validation_Loader)}\")","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.118442Z","iopub.status.idle":"2023-02-19T18:40:03.11923Z","shell.execute_reply.started":"2023-02-19T18:40:03.118958Z","shell.execute_reply":"2023-02-19T18:40:03.118982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class LearnableWeightsFFT(nn.Module):\n    def __init__(self, in_channel: int,out_channle: int,dim: int):\n        super(LearnableWeightsFFT,self).__init__()\n        \"\"\"\n        converting to  the Frequnecies Domain\n        \"\"\"\n        self.W = ( out_channle // 2 ) + 1\n        self.H = in_channel \n        self.pool =nn.MaxPool2d((2,2))\n        self.conv = nn.Conv2d(out_channle, out_channle//2, kernel_size=2, stride=1, padding=1, groups=1, bias=True)\n        self.complex_weight = nn.Parameter(torch.randn(self.H, self.W , dim, 2, dtype=torch.float32) * 0.02)\n        \n    def forward(self, x):\n        B, C, H, W = x.shape\n        x = torch.fft.rfft2(x, dim=(1, 2), norm='ortho')\n        weight = torch.view_as_complex(self.complex_weight)\n        feq_dom = x * weight\n        x = torch.fft.irfft2(feq_dom, s=(H, W), dim=(1, 2), norm='ortho')\n        x = self.conv(x)\n        return self.pool(x)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.121079Z","iopub.status.idle":"2023-02-19T18:40:03.122095Z","shell.execute_reply.started":"2023-02-19T18:40:03.121775Z","shell.execute_reply":"2023-02-19T18:40:03.121806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Classifier(nn.Module):\n    def __init__(self,in_channel , out_channel,n_classes):\n        super(Classifier,self).__init__()\n        self.pool = nn.MaxPool2d((2,2))\n        self.conv_1 = nn.Conv2d(out_channel//2, out_channel//8, kernel_size=(2, 2), stride=1, padding=1, bias=True)\n        self.pool = nn.MaxPool2d((2,2))\n        self.conv_2 = nn.Conv2d(out_channel//8, out_channel*2, kernel_size=(2, 2), stride=1, padding=1, bias=True)\n        self.flatten = nn.Flatten()\n        self.fc = torch.nn.Linear(out_channel*450,n_classes)\n    def forward(self,x):\n        x = self.conv_1(x)\n        x = self.pool(x)\n        x = self.conv_2(x)\n        x = self.pool(x)\n        x = self.flatten(x)\n        x = self.fc(x)\n        return x \n        ","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.123708Z","iopub.status.idle":"2023-02-19T18:40:03.130238Z","shell.execute_reply.started":"2023-02-19T18:40:03.12999Z","shell.execute_reply":"2023-02-19T18:40:03.130014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"###### Creating the model \nclass FourireModel(pl.LightningModule):\n    def __init__(self,in_channel=1, out_channel=120, dim=120,n_classes=1,weight=1):\n        super(FourireModel,self).__init__()\n        self.FFT = LearnableWeightsFFT(in_channel,out_channel,dim)\n        self.Classier = Classifier(out_channel,out_channel,n_classes)\n        ### setup the Optimizer and loss functions \n        self.Optimizer = torch.optim.Adam(self.parameters(),lr = Config['lr'])\n        self.Loss_F = torch.nn.BCEWithLogitsLoss(pos_weight=torch.Tensor([weight]))\n         # simple accuracy computation\n        self.train_acc = torchmetrics.Accuracy(task=\"binary\", num_classes=2)\n        self.val_acc = torchmetrics.Accuracy(task=\"binary\", num_classes=2)\n        self.test_acc = torchmetrics.Accuracy(task=\"binary\", num_classes=2)\n        self.test_prec= torchmetrics.Precision(task=\"binary\", num_classes=2)\n        self.test_recall = torchmetrics.Recall(task=\"binary\", num_classes=2)\n        \n    def forward(self,x):\n        x =  self.FFT(x)\n        x =  self.Classier(x)\n        return x    \n    \n    def training_step(self,batch , batch_idx):\n        Image , label = batch \n        label= label.float()\n        Predicted_label = self(Image)[:,0]\n        loss = self.Loss_F(Predicted_label,label)\n        # Log loss and batch aself.model(input_)ccuracy\n        self.log(\"Train Loss\", loss,sync_dist=True)\n        self.log(\"Step Train Acc\", self.train_acc(torch.sigmoid(Predicted_label), label.int()))\n        return loss\n    \n    def training_epoch_end(self, outs):\n        # After one epoch compute the whole train_data accuracy\n        self.log(\"Train Acc\", self.train_acc.compute(),sync_dist=True)\n        \n    #######\n    ### here we did the same as Traiing PL we changed only the input_data disttro\n    #######\n    def validation_step(self,batch , batch_idx):\n        Image , label = batch \n        label = label.float()\n        Predicted_label = self(Image)[:,0]\n        loss = self.Loss_F(Predicted_label,label)\n        # Log loss and batch accuracy\n        self.log(\"Val Loss\", loss,sync_dist=True)\n        self.log(\"Step Val Acc\", self.val_acc(torch.sigmoid(Predicted_label), label.int()))\n        return loss\n    \n    def validation_epoch_end(self, outs):\n        # After one epoch compute the whole train_data accuracy\n        self.log(\"Val Acc\", self.val_acc.compute(),sync_dist=True)  \n        \n        \n    def test_step(self,batch , batch_idx):\n        Image , label = batch \n        label = label.float()\n        Predicted_label = self(Image)[:,0]\n        loss = self.Loss_F(Predicted_label,label)\n        # Log loss and batch accuracy\n        self.log(\"test Loss\", loss,sync_dist=True)\n        self.log(\"Step test Acc\", self.test_acc(torch.sigmoid(Predicted_label), label.int()))\n        self.log(\"Step test prex\", self.test_prec(torch.sigmoid(Predicted_label), label.int()))\n        self.log(\"Step test Recall\", self.test_recall(torch.sigmoid(Predicted_label), label.int()))\n        \n    def training_epoch_end(self, outs):\n        # After one epoch compute the whole train_data accuracy\n        self.log(\"Test Acc\", self.train_acc.compute(),sync_dist=True) \n        self.log(\"Test Prec\", self.test_prec.compute(),sync_dist=True) \n        self.log(\"Test Recall\",self.test_recall.compute(),sync_dist=True) \n\n\n        \n    def configure_optimizers(self):\n        #Caution! You always need to return a list here (just pack your optimizer into one :))\n        return [self.Optimizer]     \n        ","metadata":{"execution":{"iopub.status.busy":"2023-03-02T18:54:03.074917Z","iopub.execute_input":"2023-03-02T18:54:03.075658Z","iopub.status.idle":"2023-03-02T18:54:03.18966Z","shell.execute_reply.started":"2023-03-02T18:54:03.075561Z","shell.execute_reply":"2023-03-02T18:54:03.18828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### instanace the model from the class \nCheck_Point_Callbacks = ModelCheckpoint(\n    monitor=\"Val Acc\", \n    save_top_k=12,\n    mode=\"max\")","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.134265Z","iopub.status.idle":"2023-02-19T18:40:03.135389Z","shell.execute_reply.started":"2023-02-19T18:40:03.135017Z","shell.execute_reply":"2023-02-19T18:40:03.135056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create the trainer\n# Change the gpus parameter to the number of available gpus on your system. Use 0 for CPU training\n\ngpus = 2 #TODO\nTrainer = pl.Trainer(accelerator='cpu', \n                     logger=TensorBoardLogger(save_dir= \"./processed/logs_weights_ex4\"), log_every_n_steps=1,\n                     callbacks=Check_Point_Callbacks,                    \n                     max_epochs=30,fast_dev_run=False)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.137965Z","iopub.status.idle":"2023-02-19T18:40:03.139066Z","shell.execute_reply.started":"2023-02-19T18:40:03.138779Z","shell.execute_reply":"2023-02-19T18:40:03.13881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Trainer.fit(FourireModel(),Training_Loader,Validation_Loader)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.142363Z","iopub.status.idle":"2023-02-19T18:40:03.143798Z","shell.execute_reply.started":"2023-02-19T18:40:03.143477Z","shell.execute_reply":"2023-02-19T18:40:03.143513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the extension and start TensorBoard\n\n%reload_ext tensorboard\n%tensorboard --logdir ./processed/logs_weights_ex4/lightning_logs","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.148355Z","iopub.status.idle":"2023-02-19T18:40:03.149415Z","shell.execute_reply.started":"2023-02-19T18:40:03.149054Z","shell.execute_reply":"2023-02-19T18:40:03.149095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Evaluation Model** <a class=\"anchor\" id=\"Secteion-Eval\"> </a>\n* here we will load the weights the model \n* preduct few samples from the validation dataset \n* save the reults in tensor data type \n* compute few matrices results ","metadata":{}},{"cell_type":"code","source":"# Create the trainer\n# Change the gpus parameter to the number of available gpus on your system. Use 0 for CPU training\n\ngpus = 2 #TODO\nTrainer_test = pl.Trainer(accelerator='cpu', \n                     logger=TensorBoardLogger(save_dir= \"./processed/logs_weights_ex4\"), log_every_n_steps=1,\n                     callbacks=Check_Point_Callbacks,                    \n                     max_epochs=60,fast_dev_run=False)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.151252Z","iopub.status.idle":"2023-02-19T18:40:03.154152Z","shell.execute_reply.started":"2023-02-19T18:40:03.153863Z","shell.execute_reply":"2023-02-19T18:40:03.15389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import glob\n### Loading the model check point weight \ndevice = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")\npath_model = \"./processed/logs_weights_ex4/lightning_logs/version_0/checkpoints/\"\nlast_epoch = glob.glob(path_model+\"/*.ckpt\")[-1]\nTrainer_test.test(FourireModel(), dataloaders=Validation_Loader, ckpt_path=last_epoch)\n","metadata":{"execution":{"iopub.status.busy":"2023-03-02T17:25:32.783083Z","iopub.execute_input":"2023-03-02T17:25:32.783578Z","iopub.status.idle":"2023-03-02T17:25:32.87254Z","shell.execute_reply.started":"2023-03-02T17:25:32.783537Z","shell.execute_reply":"2023-03-02T17:25:32.870559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_load = FourireModel()\nmodel_load.load_from_checkpoint(last_epoch)\nmodel_load.to(device) \nmodel_load.eval()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.159429Z","iopub.status.idle":"2023-02-19T18:40:03.162694Z","shell.execute_reply.started":"2023-02-19T18:40:03.162441Z","shell.execute_reply":"2023-02-19T18:40:03.162466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predicted_ = []\nlabels_ = []\nwith torch.no_grad():\n    for image , label in tqdm(Loading_DataFolder_Val):\n        ## here we will need to load the data into device by hand \n        image = image.to(device).float().unsqueeze(0)\n        label = label\n        predicted= torch.sigmoid(model_load(image)[0]).cpu()\n        predicted_.append(predicted)\n        labels_.append(label)\n    print(image.shape) \n    print(predicted.shape)\n            \nTensor_Pre = torch.tensor(predicted_)        \nTensor_Lab = torch.tensor(labels_).int() ","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.1641Z","iopub.status.idle":"2023-02-19T18:40:03.164963Z","shell.execute_reply.started":"2023-02-19T18:40:03.164684Z","shell.execute_reply":"2023-02-19T18:40:03.164711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Tensor_Lab.shape , Tensor_Pre.shape","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.168724Z","iopub.status.idle":"2023-02-19T18:40:03.169817Z","shell.execute_reply.started":"2023-02-19T18:40:03.169475Z","shell.execute_reply":"2023-02-19T18:40:03.169509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Submission Files <a class=\"anchor\" id=\"Submission\"></a>\n\n* **here will predict the sampels from vaidations data and save the Patient_ID and predicte value**\n\n","metadata":{}},{"cell_type":"code","source":"acc = torchmetrics.Accuracy(task=\"binary\", num_classes=2)(Tensor_Pre, Tensor_Lab)\nprecision = torchmetrics.Precision(task=\"binary\", num_classes=2)(Tensor_Pre, Tensor_Lab)\nrecall = torchmetrics.Recall(task=\"binary\", num_classes=2)(Tensor_Pre, Tensor_Lab)\ncm = torchmetrics.ConfusionMatrix(task=\"binary\", num_classes=2)(Tensor_Pre, Tensor_Lab)\ncm_threshed = torchmetrics.ConfusionMatrix(task=\"binary\", num_classes=2 ,threshold=0.25)(Tensor_Pre, Tensor_Lab)\n\nprint(f\"Val Accuracy: {acc}\")\nprint(f\"Val Precision: {precision}\")\nprint(f\"Val Recall: {recall}\")\nprint(f\"Confusion Matrix:\\n {cm}\")\nprint(f\"Confusion Matrix 2:\\n {cm_threshed}\")","metadata":{"execution":{"iopub.status.busy":"2023-03-02T20:42:02.476028Z","iopub.execute_input":"2023-03-02T20:42:02.476563Z","iopub.status.idle":"2023-03-02T20:42:02.560272Z","shell.execute_reply.started":"2023-03-02T20:42:02.476524Z","shell.execute_reply":"2023-03-02T20:42:02.558741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n#### Virtualiza some images \nDATA_DIR = \"../input/rsna-breast-cancer-detection/\"\n# Directory to save logs and trained model\nROOT_DIR = '/kaggle/working'\n\ntest_dicom_dir = os.path.join(DATA_DIR, 'test_images/10008')\nprint(test_dicom_dir)\nsub_read = pd.read_csv(\"/kaggle/input/rsna-breast-cancer-detection/sample_submission.csv\")\nsub_read.head(3)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.17779Z","iopub.status.idle":"2023-02-19T18:40:03.178922Z","shell.execute_reply.started":"2023-02-19T18:40:03.178571Z","shell.execute_reply":"2023-02-19T18:40:03.178606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DCM_TEST_IMAGES_PATH = '/kaggle/input/rsna-breast-cancer-detection/test_images/10008'\nRSNA_2022_PATH = '/kaggle/input/rsna-breast-cancer-detection'","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.183444Z","iopub.status.idle":"2023-02-19T18:40:03.18425Z","shell.execute_reply.started":"2023-02-19T18:40:03.183969Z","shell.execute_reply":"2023-02-19T18:40:03.183994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test = pd.read_csv(\"/kaggle/input/rsna-breast-cancer-detection/test.csv\")\ndf_test[\"img_name\"] = df_test[\"patient_id\"].astype(str) + \"_\" + df_test[\"image_id\"].astype(str) + \".png\"\ndf_test[\"dcm_path\"] = df_test[\"patient_id\"].astype(str) + \"_\" + df_test[\"image_id\"].astype(str) + \".dcm\"\ndf_test.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.186144Z","iopub.status.idle":"2023-02-19T18:40:03.192987Z","shell.execute_reply.started":"2023-02-19T18:40:03.192723Z","shell.execute_reply":"2023-02-19T18:40:03.192748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import glob\ndef get_dicom_fps(dicom_dir):\n    dicom_fps = glob.glob(dicom_dir+'/*.dcm')\n    return list(set(dicom_fps))","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.194731Z","iopub.status.idle":"2023-02-19T18:40:03.195754Z","shell.execute_reply.started":"2023-02-19T18:40:03.19544Z","shell.execute_reply":"2023-02-19T18:40:03.195465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get filenames of test dataset DICOM images\ntest_image_fps = get_dicom_fps(DCM_TEST_IMAGES_PATH)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.197884Z","iopub.status.idle":"2023-02-19T18:40:03.198725Z","shell.execute_reply.started":"2023-02-19T18:40:03.198461Z","shell.execute_reply":"2023-02-19T18:40:03.198486Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def models_predict(model, image_fps, min_conf=0.96):\n    with torch.no_grad():\n        predictions = []\n        for idx, image_id in enumerate(tqdm(image_fps)):\n            ds = dc.read_file(image_id)\n            image = ds.pixel_array  / 255 ##normalize \n            image_2d = image.astype(float)\n            image = cv2.resize(image , (120,120)).astype(np.float32)\n            image = np.expand_dims(image,axis=0)\n            image= torch.from_numpy(image)\n            image = image.to(device).float().unsqueeze(0)\n            y_hat= torch.sigmoid(model(image)[0])\n            #pred_proba = torch.argmax(y_hat,dim=1)\n            porbability = y_hat.to(torch.float32).cpu()\n            predictions.append(porbability)\n        return torch.concat(predictions).numpy()","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.200295Z","iopub.status.idle":"2023-02-19T18:40:03.201162Z","shell.execute_reply.started":"2023-02-19T18:40:03.200909Z","shell.execute_reply":"2023-02-19T18:40:03.200933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"models_pred=models_predict(model_load,test_image_fps)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.204807Z","iopub.status.idle":"2023-02-19T18:40:03.205848Z","shell.execute_reply.started":"2023-02-19T18:40:03.205518Z","shell.execute_reply":"2023-02-19T18:40:03.205551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds = np.asarray(models_pred)\npreds","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.211642Z","iopub.status.idle":"2023-02-19T18:40:03.21272Z","shell.execute_reply.started":"2023-02-19T18:40:03.21236Z","shell.execute_reply":"2023-02-19T18:40:03.212394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"prediction_id = df_test['patient_id'].astype(str) + \"_\" + df_test['laterality']","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.214616Z","iopub.status.idle":"2023-02-19T18:40:03.215651Z","shell.execute_reply.started":"2023-02-19T18:40:03.215289Z","shell.execute_reply":"2023-02-19T18:40:03.215332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = {\n    \"prediction_id\": np.array((prediction_id)),\n    \"cancer\": preds.T\n}\n\nsub_df = pd.DataFrame(data=data)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.217056Z","iopub.status.idle":"2023-02-19T18:40:03.21786Z","shell.execute_reply.started":"2023-02-19T18:40:03.21761Z","shell.execute_reply":"2023-02-19T18:40:03.217633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subb = sub_df.drop_duplicates(\"prediction_id\")\nsubb","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.22054Z","iopub.status.idle":"2023-02-19T18:40:03.22489Z","shell.execute_reply.started":"2023-02-19T18:40:03.224518Z","shell.execute_reply":"2023-02-19T18:40:03.224553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"subb.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-02-19T18:40:03.226461Z","iopub.status.idle":"2023-02-19T18:40:03.227328Z","shell.execute_reply.started":"2023-02-19T18:40:03.227045Z","shell.execute_reply":"2023-02-19T18:40:03.227072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-info\" role=\"alert\"> <strong>Next Notebook will be about TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide \nImage Classification Check out in my Profile</strong>\n    </div>\n","metadata":{}},{"cell_type":"markdown","source":"<center>\n<img src=\"https://img.shields.io/badge/Upvote-If%20you%20like%20my%20work-07b3c8?style=for-the-badge&logo=kaggle\">\n</center>","metadata":{}}]}