{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\n#for dirname, _, filenames in os.walk('/kaggle/input'):\n #   for filename in filenames:\n  #      print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-02-20T20:02:53.640319Z","iopub.execute_input":"2023-02-20T20:02:53.640739Z","iopub.status.idle":"2023-02-20T20:02:53.65977Z","shell.execute_reply.started":"2023-02-20T20:02:53.640658Z","shell.execute_reply":"2023-02-20T20:02:53.658733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!lscpu","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:02:53.66516Z","iopub.execute_input":"2023-02-20T20:02:53.665432Z","iopub.status.idle":"2023-02-20T20:02:54.651804Z","shell.execute_reply.started":"2023-02-20T20:02:53.665407Z","shell.execute_reply":"2023-02-20T20:02:54.650603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## To DO list\n\n### Apply Resnet50,101, DenseNet121,vgg16,inception_resnet,efficientnet,etc. with different combinations.Anyother pre_trained model for medical images, with some combinations too if applicable.try to use first, Resnet101, Efficientb7,vgg16,densenet169\n\n### Use AlexNet. create it. One the study shows AlexNet for Breasxt Xrays and for Breast, inception model.\n\n### try to use different filters like gaussian an others in Data Image Gen in model before using pre-trained or designed model. try to remove test from the images but first use gradcam to see whether it's effecting or not the model\n\n### how to use graysacle images as rgb. try to see different techniuqes to improve the visibility.\n\n### Go through different CAM (Class Activation Maps) like GradCam, AblationCap,ScoreCAM etc. Go through Certifications like tensorflow advanced plus AI in medicine.\n\n### Image preprocessing and enhancement : brightness & contrast.\n\n### TensorFlow Filtes check gabor,gaussian,laplacian,prewitt & sobel\n\n# Use shape(224,224) normally if not given the shapes like in Inception-resnet (299,299). normalize test images with train data and also apply other techniques as well like image enhancement. Apply focal loss if you change the data. also use weighted loss as well.\n\n### In Effects on Image Transformation, adam with 0.0005. EfficientnetB0 model applied.SMQT, Wavelet transform, Laplace transform and CLAHE were analyzed.The effect of five image transformations such as CLAHE, Wavelet transform, adaptive gamma correction, and Laplace transform are evaluated on the EfficientNet algorithm.wihtout enhancement, achieved 90.15%.only laplace enhancement give 75%.\n\n### some images has white background and some has black. try to create a similar one for all images.\n\n### look into pos_weight paramter of focal loss. maybe given the pos_weights which can be extracted through some known functions like in AI For Medicine assignment.","metadata":{}},{"cell_type":"markdown","source":"## TO DO LIST\n\n### apply all filters on tensorflow too (also read) & allign with notebook bets filter used in,Image Filtes in Python on Desktop etc. which filters are useful in dealing with medical images(1h)\n\n### Submisison (20m),Deep Dive (20m), Efficientnet (20m), Colab Alexnet plus other pretrained models with filter techniuqes as well, plus further tuing on new dataset with low learnign rate with tensorflow link help. (1.5h)\n\n### apply tpu including removing medical_id if common and also see discussion forum (1h), apply u-net, see information through courses and other stuff as well (1h) \n\n### First read papers which three paper On top including tommorrow downloads (1h).\n### Sort it all stuff for next week including documentation to follow including rsna folders etc(30min), \n### Class Activation Maps & other related stuff(30mins).graysacleimages as rgb for mdeical images(30)\n### google searches(1h), look into scale Data Aug paramter through courses (10min)","metadata":{}},{"cell_type":"markdown","source":"Diving deep into focal loss(just rotate and flip,bacth=16,my_assessment=use 1/.255 for normalization not using training mean, try to make alignment with pre-trained model.)","metadata":{}},{"cell_type":"code","source":"#!dpkg -l libcudnn8*","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:02:54.654686Z","iopub.execute_input":"2023-02-20T20:02:54.655099Z","iopub.status.idle":"2023-02-20T20:02:54.660196Z","shell.execute_reply.started":"2023-02-20T20:02:54.655059Z","shell.execute_reply":"2023-02-20T20:02:54.659032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#!pip install pip install nvidia-cudnn-cu11 --no-index --find-links=file:///kaggle/input/nvidia-cudnn-cu11-880121/storage_dir/","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:02:54.662059Z","iopub.execute_input":"2023-02-20T20:02:54.662563Z","iopub.status.idle":"2023-02-20T20:02:54.670532Z","shell.execute_reply.started":"2023-02-20T20:02:54.662528Z","shell.execute_reply":"2023-02-20T20:02:54.669423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#/kaggle/input/pythongdcm-package\n!pip install python-gdcm --no-index --find-links=file:///kaggle/input/pythongdcm-package/storage_dir/\n!pip install dicomsdl --no-index --find-links=file:///kaggle/input/dicomsdl-package/storage_dir/\n#!pip install tensorflow --no-index --find-links=file:///kaggle/input/tensorflow/storage_dir/\n#!pip install tf_clahe[tensorflow] --no-index --find-links=file:///kaggle/input/tensorflow-clahe/storage_dir/\n#!pip install tensorflow_addons[tensorflow] --no-index --find-links=file:///kaggle/input/tensorflow-addons/storage_dir/\n#!pip install tensorflow-wavelets --no-index --find-links=file:///kaggle/input/tensorflow-wavelets/storage_dir/\n#!pip install cuda-python --no-index --find-links=file:///kaggle/input/cuda-python/storage_dir/\n#!pip install focal-loss --no-index --find-links=file:///kaggle/input/focal-loss-package/storage_dir/","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:02:54.673703Z","iopub.execute_input":"2023-02-20T20:02:54.674055Z","iopub.status.idle":"2023-02-20T20:03:15.763499Z","shell.execute_reply.started":"2023-02-20T20:02:54.674028Z","shell.execute_reply":"2023-02-20T20:03:15.762342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#!dpkg -l libcudnn8*","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:15.765546Z","iopub.execute_input":"2023-02-20T20:03:15.765937Z","iopub.status.idle":"2023-02-20T20:03:15.773554Z","shell.execute_reply.started":"2023-02-20T20:03:15.765898Z","shell.execute_reply":"2023-02-20T20:03:15.77261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#pip install cudatoolkit=11.2 cudnn=8.1.0","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:15.776112Z","iopub.execute_input":"2023-02-20T20:03:15.776392Z","iopub.status.idle":"2023-02-20T20:03:15.781986Z","shell.execute_reply.started":"2023-02-20T20:03:15.776367Z","shell.execute_reply":"2023-02-20T20:03:15.781024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#!pip uninstall -y tensorflow","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:15.783597Z","iopub.execute_input":"2023-02-20T20:03:15.783964Z","iopub.status.idle":"2023-02-20T20:03:15.793461Z","shell.execute_reply.started":"2023-02-20T20:03:15.783929Z","shell.execute_reply":"2023-02-20T20:03:15.792508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#!pip install 'tensorflow==2.6.4' --no-index --find-links=file:///kaggle/input/tensorflow-version-264/storage_dir/","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:15.795085Z","iopub.execute_input":"2023-02-20T20:03:15.795468Z","iopub.status.idle":"2023-02-20T20:03:15.803217Z","shell.execute_reply.started":"2023-02-20T20:03:15.795433Z","shell.execute_reply":"2023-02-20T20:03:15.802327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#!apt install -y --allow-change-held-packages libcudnn8=8.1.0.77-1+cuda11.2","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:15.804678Z","iopub.execute_input":"2023-02-20T20:03:15.805068Z","iopub.status.idle":"2023-02-20T20:03:15.812674Z","shell.execute_reply.started":"2023-02-20T20:03:15.805035Z","shell.execute_reply":"2023-02-20T20:03:15.811761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IS_INTERACTIVE = os.environ['KAGGLE_KERNEL_RUN_TYPE'] == 'Interactive'\nos.environ['CUDA_VISIBLE_DEVICES'] = \"0\"\n#import multiprocessing\n#from multiprocessing import Process\n#multiprocessing.set_start_method('spawn', force=True)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:15.819493Z","iopub.execute_input":"2023-02-20T20:03:15.820209Z","iopub.status.idle":"2023-02-20T20:03:15.825333Z","shell.execute_reply.started":"2023-02-20T20:03:15.820175Z","shell.execute_reply":"2023-02-20T20:03:15.824386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nfrom pydicom import dcmread\nimport dicomsdl as dicom\nfrom skimage.transform import resize\nfrom tensorflow.keras.preprocessing.image import img_to_array, load_img\nimport pydicom \nimport cv2\nfrom joblib import Parallel, delayed\nfrom PIL import Image\nfrom tensorflow.keras.preprocessing.image import img_to_array, load_img\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.applications.densenet import DenseNet121\nfrom tensorflow.keras.applications import resnet\nfrom tensorflow.keras.layers import Dense,Dropout, GlobalAveragePooling2D\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D, Dense, Dropout\n#from focal_loss import BinaryFocalLoss\nfrom tensorflow.keras import backend as K\npd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', None)\nfrom tensorflow.keras.models import load_model\nimport tensorflow_addons as tfa\nfrom sklearn.utils import resample\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.utils import shuffle\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\nimport tensorflow_io as tfio\nfrom tensorflow.keras.models import Sequential \nfrom keras.callbacks import EarlyStopping\n#import util\n#from public_tests import *\n#from test_utils import *\nfrom multiprocessing import cpu_count\ntf.config.threading.set_inter_op_parallelism_threads(num_threads=1)\ncv2.setNumThreads(1)\n#tf.compat.v1.logging.set_verbosity(tf.compat.v1.logging.ERROR)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:15.826765Z","iopub.execute_input":"2023-02-20T20:03:15.827754Z","iopub.status.idle":"2023-02-20T20:03:22.556541Z","shell.execute_reply.started":"2023-02-20T20:03:15.827716Z","shell.execute_reply":"2023-02-20T20:03:22.555573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IS_INTERACTIVE = os.environ['KAGGLE_KERNEL_RUN_TYPE'] == 'Interactive'\n#os.environ['CUDA_VISIBLE_DEVICES'] = \"0\"","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:22.558111Z","iopub.execute_input":"2023-02-20T20:03:22.558813Z","iopub.status.idle":"2023-02-20T20:03:22.564473Z","shell.execute_reply.started":"2023-02-20T20:03:22.558775Z","shell.execute_reply":"2023-02-20T20:03:22.562486Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\ntrain_dir = '/kaggle/input/rsna-breast-cancer-detection/train_images/' \ntrain['images'] = [(train_dir + str(k) + '/' + str(v) + '.dcm') for k,v in train[['patient_id','image_id']].values]\nprint ('train Data Size',train.shape)\ntrain.head() ","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:22.566141Z","iopub.execute_input":"2023-02-20T20:03:22.566485Z","iopub.status.idle":"2023-02-20T20:03:22.811446Z","shell.execute_reply.started":"2023-02-20T20:03:22.56645Z","shell.execute_reply":"2023-02-20T20:03:22.810324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#columns  = [i for i in df_oversampled.columns]\ncolumns  = [i for i in train.columns]\nprint (type(columns))\n(columns.remove('cancer'))\nprint (columns)\n#X = df_oversampled[columns]\n#y = df_oversampled[['cancer']]\nX = train[columns]\ny = train[['cancer']]","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:22.813644Z","iopub.execute_input":"2023-02-20T20:03:22.814666Z","iopub.status.idle":"2023-02-20T20:03:22.830446Z","shell.execute_reply.started":"2023-02-20T20:03:22.814628Z","shell.execute_reply":"2023-02-20T20:03:22.829368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# data split into train,test (later add validation)\nX_ini, X_test, y_ini, y_test = train_test_split(X, y, test_size=0.41,\n                                                    random_state=42,shuffle=True,\n                                                    stratify = y)\n\nX_train,X_val,y_train,y_val = train_test_split(X_ini, y_ini, test_size=0.41,\n                                                    random_state=42,shuffle=True,\n                                                    stratify = y_ini)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:22.831802Z","iopub.execute_input":"2023-02-20T20:03:22.833749Z","iopub.status.idle":"2023-02-20T20:03:23.184476Z","shell.execute_reply.started":"2023-02-20T20:03:22.833699Z","shell.execute_reply":"2023-02-20T20:03:23.183497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.shape,X_test.shape,X_val.shape","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:23.185931Z","iopub.execute_input":"2023-02-20T20:03:23.186294Z","iopub.status.idle":"2023-02-20T20:03:23.193324Z","shell.execute_reply.started":"2023-02-20T20:03:23.186259Z","shell.execute_reply":"2023-02-20T20:03:23.192229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def check_for_leakage(df1, df2, patient_col):\n    df1_patients_unique = set(df1[patient_col].values)\n    df2_patients_unique = set(df2[patient_col].values)\n    patients_in_both_groups = df1_patients_unique.intersection(df2_patients_unique)\n    leakage = len(patients_in_both_groups)\n    return leakage\n    \nprint(\"Leakage between train and test: {}\".format(check_for_leakage(X_train, X_test, 'patient_id')))\nprint(\"Leakage between train and val: {}\".format(check_for_leakage(X_train, X_val, 'patient_id')))\nprint(\"Leakage between test and val: {}\".format(check_for_leakage(X_test, X_val, 'patient_id')))\n\ndef remove_leakage(train_X,test_X,train_y,test_y, patient_col):\n    ids_unique_train = set(train_X[patient_col].values)\n    ids_unique_test  = set(test_X[patient_col].values)\n    patient_overlap = list(ids_unique_train.intersection(ids_unique_test))\n    n_overlap = len(patient_overlap)\n    train_overlap_idxs = []\n    test_overlap_idxs = []\n    \n    for idx in range (n_overlap):\n        train_overlap_idxs.extend(train_X.index[train_X['patient_id'] == patient_overlap[idx]].tolist())\n        test_overlap_idxs.extend(test_X.index[test_X['patient_id'] == patient_overlap[idx]].tolist())\n    \n    #print(f'These are the indices of overlapping patients in the training set: ')\n    #print(f'{train_overlap_idxs}')\n    #print(f'These are the indices of overlapping patients in the test set: ')\n    #print(f'{test_overlap_idxs}')\n    #X_train.drop(patient_overlap,index)\n    #print('Before excluding overlap ids from set',test_X.shape)\n    #print('Before excluding overlap ids from  set labels',test_y.shape)\n    test_X.drop(test_overlap_idxs,inplace=True)\n    test_y.drop(test_overlap_idxs,inplace=True)\n    \n    return test_X,test_y","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:23.194927Z","iopub.execute_input":"2023-02-20T20:03:23.195378Z","iopub.status.idle":"2023-02-20T20:03:23.225287Z","shell.execute_reply.started":"2023-02-20T20:03:23.195343Z","shell.execute_reply":"2023-02-20T20:03:23.224387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test_new,y_test_new = remove_leakage(X_train,X_test,y_train,y_test, 'patient_id')\nX_val_new,y_val_new = remove_leakage (X_train,X_val,y_train,y_val,'patient_id')\nprint ('New Test Data Size',X_test_new.shape,'New Test Label Size',y_test_new.shape)\nprint ('New Val Data Size',X_val_new.shape,'New Val Label Size',y_val_new.shape)    ","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:23.226689Z","iopub.execute_input":"2023-02-20T20:03:23.227302Z","iopub.status.idle":"2023-02-20T20:03:26.978483Z","shell.execute_reply.started":"2023-02-20T20:03:23.227266Z","shell.execute_reply":"2023-02-20T20:03:26.977444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Checking Further Leakage between new test and val: {}\".format(check_for_leakage(X_test_new, X_val_new, 'patient_id')))","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:26.981412Z","iopub.execute_input":"2023-02-20T20:03:26.981688Z","iopub.status.idle":"2023-02-20T20:03:26.990991Z","shell.execute_reply.started":"2023-02-20T20:03:26.981662Z","shell.execute_reply":"2023-02-20T20:03:26.989953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_valid,y_valid = remove_leakage (X_test_new,X_val_new,y_test_new,y_val_new,'patient_id')\nprint ('New Validation Data Size',X_valid.shape,'New Validation Label Size',y_valid.shape)    ","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:26.992791Z","iopub.execute_input":"2023-02-20T20:03:26.993173Z","iopub.status.idle":"2023-02-20T20:03:27.329706Z","shell.execute_reply.started":"2023-02-20T20:03:26.993137Z","shell.execute_reply":"2023-02-20T20:03:27.328625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print  (f\"Test Data Size is {X_test_new.shape[0]}, now almost 25% of train set {X_train.shape[0] * 0.25}\")","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.331109Z","iopub.execute_input":"2023-02-20T20:03:27.331581Z","iopub.status.idle":"2023-02-20T20:03:27.338631Z","shell.execute_reply.started":"2023-02-20T20:03:27.331543Z","shell.execute_reply":"2023-02-20T20:03:27.33737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.shape,X_test_new.shape,X_valid.shape","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.340088Z","iopub.execute_input":"2023-02-20T20:03:27.3409Z","iopub.status.idle":"2023-02-20T20:03:27.351262Z","shell.execute_reply.started":"2023-02-20T20:03:27.340861Z","shell.execute_reply":"2023-02-20T20:03:27.34982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.concat([X_train,y_train],axis=1)\ntest_df = pd.concat([X_test_new,y_test_new],axis=1)\nvalid_df = pd.concat([X_valid,y_valid],axis=1)\ntrain_df.reset_index(inplace=True)\ntest_df.reset_index(inplace=True)\nvalid_df.reset_index(inplace=True)\ntrain_df.drop(['index'],axis=1,inplace=True)\ntest_df.drop(['index'],axis=1,inplace=True)\nvalid_df.drop(['index'],axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.352743Z","iopub.execute_input":"2023-02-20T20:03:27.353211Z","iopub.status.idle":"2023-02-20T20:03:27.378468Z","shell.execute_reply.started":"2023-02-20T20:03:27.353175Z","shell.execute_reply":"2023-02-20T20:03:27.377615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head(1)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.379688Z","iopub.execute_input":"2023-02-20T20:03:27.380056Z","iopub.status.idle":"2023-02-20T20:03:27.396361Z","shell.execute_reply.started":"2023-02-20T20:03:27.380019Z","shell.execute_reply":"2023-02-20T20:03:27.395168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Leakage between train and test on image_id: {}\".format(check_for_leakage(train_df, test_df, 'image_id')))\nprint(\"Leakage between train and val on image_id: {}\".format(check_for_leakage(train_df, valid_df, 'image_id')))\nprint(\"Leakage between test and val on image_id: {}\".format(check_for_leakage(test_df, valid_df, 'image_id')))","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.398005Z","iopub.execute_input":"2023-02-20T20:03:27.398678Z","iopub.status.idle":"2023-02-20T20:03:27.413314Z","shell.execute_reply.started":"2023-02-20T20:03:27.398641Z","shell.execute_reply":"2023-02-20T20:03:27.412448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#create two different dataframe according to majorityity class ('1')\n# train_df oversampling\ntrain_label_one = train_df[train_df['cancer']==1]\ntrain_label_zero = train_df[train_df['cancer']==0]\n# undersample majority classes\ntrain_one_oversampled = resample(train_label_one, \n                                 replace=True,    # sample with replacement\n                                 n_samples= train_label_zero.shape[0], # 820, # to match majority class\n                                 random_state=42)  # reproducible results\n# Combine majority class with upsampled minority class\ntrain_oversampled = pd.concat([train_one_oversampled,train_label_zero])\n\n# index reset and data shuffling\ntrain_oversampled = shuffle (train_oversampled)\ntrain_oversampled.reset_index(inplace=True)\ntrain_oversampled.drop(['index'],axis=1,inplace=True)\nprint ('train',train_oversampled.shape)\ntrain_oversampled.head(1)\n\n#create two different dataframe according to majorityity class ('1')\n# test_df oversampling\ntest_label_one = test_df[test_df['cancer']==1]\ntest_label_zero = test_df[test_df['cancer']==0]\n# undersample majority classes\ntest_one_oversampled = resample(test_label_one, \n                                 replace=True,    # sample with replacement\n                                 n_samples= test_label_zero.shape[0], # 820, # to match majority class\n                                 random_state=42)  # reproducible results\n# Combine majority class with upsampled minority class\ntest_oversampled = pd.concat([test_one_oversampled,test_label_zero])\n\n# index reset and data shuffling\ntest_oversampled = shuffle (test_oversampled)\ntest_oversampled.reset_index(inplace=True)\ntest_oversampled.drop(['index'],axis=1,inplace=True)\nprint ('test',test_oversampled.shape)\ntest_oversampled.head(1)\n\n#create two different dataframe according to majorityity class ('1')\n# valid_df oversampling\nvalid_label_one = valid_df[valid_df['cancer']==1]\nvalid_label_zero = valid_df[valid_df['cancer']==0]\n# undersample majority classes\nvalid_one_oversampled = resample(valid_label_one, \n                                 replace=True,    # sample with replacement\n                                 n_samples= valid_label_zero.shape[0], # 820, # to match majority class\n                                 random_state=42)  # reproducible results\n# Combine majority class with upsampled minority class\nvalid_oversampled = pd.concat([valid_one_oversampled,valid_label_zero])\n\n# index reset and data shuffling\nvalid_oversampled = shuffle (valid_oversampled)\nvalid_oversampled.reset_index(inplace=True)\nvalid_oversampled.drop(['index'],axis=1,inplace=True)\nprint ('valid',valid_oversampled.shape)\nvalid_oversampled.head(1)\n\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.416244Z","iopub.execute_input":"2023-02-20T20:03:27.416513Z","iopub.status.idle":"2023-02-20T20:03:27.489173Z","shell.execute_reply.started":"2023-02-20T20:03:27.416488Z","shell.execute_reply":"2023-02-20T20:03:27.488204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dir = '/kaggle/input/train-rsna-images'\ntest_dir = '/kaggle/input/test-rsna-images'\nval_dir = '/kaggle/input/validation-rsna-images'\ntrain_df['png_images'] = [(str(i) + '.png') for i in train_df['image_id']]\ntest_df['png_images'] = [(str(i) + '.png') for i in test_df['image_id']]\nvalid_df['png_images'] = [(str(i) + '.png') for i in valid_df['image_id']]","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.49054Z","iopub.execute_input":"2023-02-20T20:03:27.49139Z","iopub.status.idle":"2023-02-20T20:03:27.508547Z","shell.execute_reply.started":"2023-02-20T20:03:27.491352Z","shell.execute_reply":"2023-02-20T20:03:27.507492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_oversampled['png_images'] = [(str(i) + '.png') for i in train_oversampled['image_id']]\ntest_oversampled['png_images'] = [(str(i) + '.png') for i in test_oversampled['image_id']]\nvalid_oversampled['png_images'] = [(str(i) + '.png') for i in valid_oversampled['image_id']]","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.517779Z","iopub.execute_input":"2023-02-20T20:03:27.518094Z","iopub.status.idle":"2023-02-20T20:03:27.545635Z","shell.execute_reply.started":"2023-02-20T20:03:27.518068Z","shell.execute_reply":"2023-02-20T20:03:27.544619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print (train_oversampled.shape[0]//8)\nprint (test_oversampled.shape[0]//8)\nprint (valid_oversampled.shape[0]//8)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.54723Z","iopub.execute_input":"2023-02-20T20:03:27.547603Z","iopub.status.idle":"2023-02-20T20:03:27.557143Z","shell.execute_reply.started":"2023-02-20T20:03:27.547567Z","shell.execute_reply":"2023-02-20T20:03:27.555911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_quarter = train_oversampled.loc[:4659,:]\ntest_quarter = test_oversampled.loc[:1193,:]\nvalid_quarter = valid_oversampled.loc[:112,:]\ntrain_quarter = shuffle(train_quarter)\ntest_quarter = shuffle(test_quarter)\nvalid_quarter = shuffle(valid_quarter)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.558369Z","iopub.execute_input":"2023-02-20T20:03:27.55921Z","iopub.status.idle":"2023-02-20T20:03:27.570566Z","shell.execute_reply.started":"2023-02-20T20:03:27.559183Z","shell.execute_reply":"2023-02-20T20:03:27.569679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print (valid_quarter[valid_quarter['cancer']==0].shape[0])\nprint (valid_quarter[valid_quarter['cancer']==1].shape[0])","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.572066Z","iopub.execute_input":"2023-02-20T20:03:27.572716Z","iopub.status.idle":"2023-02-20T20:03:27.582049Z","shell.execute_reply.started":"2023-02-20T20:03:27.572681Z","shell.execute_reply":"2023-02-20T20:03:27.580622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#train_oversampled['full_path'] = [(train_dir + '/' + i) for i in train_oversampled['png_images']]\n#train_oversampled.head(1)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.583612Z","iopub.execute_input":"2023-02-20T20:03:27.584236Z","iopub.status.idle":"2023-02-20T20:03:27.589689Z","shell.execute_reply.started":"2023-02-20T20:03:27.584196Z","shell.execute_reply":"2023-02-20T20:03:27.588509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"source_path = '/kaggle/input/validation-rsna-images'\nimages_no_size = []\nfiles = os.listdir(source_path)\nfor file in (files):\n    #print(file)\n    #break\n    if os.path.getsize(os.path.join(source_path,file)) == 0:\n      print (f'{file} is zero length, so ignoring.')\n      images_no_size.append(file)\nif len(images_no_size) == 0:\n  print ('There are no images with zero size')","metadata":{"execution":{"iopub.status.busy":"2023-02-16T17:36:34.064259Z","iopub.execute_input":"2023-02-16T17:36:34.064575Z","iopub.status.idle":"2023-02-16T17:36:34.122904Z","shell.execute_reply.started":"2023-02-16T17:36:34.064544Z","shell.execute_reply":"2023-02-16T17:36:34.121782Z"}}},{"cell_type":"code","source":"#os.path.getsize(os.path.join(train_oversampled['full_path'][0]))","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.591436Z","iopub.execute_input":"2023-02-20T20:03:27.591961Z","iopub.status.idle":"2023-02-20T20:03:27.598958Z","shell.execute_reply.started":"2023-02-20T20:03:27.591925Z","shell.execute_reply":"2023-02-20T20:03:27.59788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def adjust_gamma(image):\n    return tf.image.adjust_gamma(\n    image, gamma=0.2, gain=1\n    )","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.600555Z","iopub.execute_input":"2023-02-20T20:03:27.601081Z","iopub.status.idle":"2023-02-20T20:03:27.607376Z","shell.execute_reply.started":"2023-02-20T20:03:27.601045Z","shell.execute_reply":"2023-02-20T20:03:27.606115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#@tf.function(experimental_compile=True) \n#import tf_clahe\ndef tfclahe(img):\n    return tf_clahe.clahe(img)#, gpu_optimized=True)#return tf_clahe.clahe(img)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.608759Z","iopub.execute_input":"2023-02-20T20:03:27.60963Z","iopub.status.idle":"2023-02-20T20:03:27.619223Z","shell.execute_reply.started":"2023-02-20T20:03:27.609594Z","shell.execute_reply":"2023-02-20T20:03:27.618475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"%%time\nIMG_PX_SIZE =299\nimage = '/kaggle/input/rsna-breast-cancer-detection/test_images/10008/361203119.dcm'\n#image = '/kaggle/input/samples-test-rsna-images/361203119.png'\nds = pydicom.dcmread(image,force=True) #pydicom.read_file\ndata = ds.pixel_array\nresized_img = resize(data, (IMG_PX_SIZE, IMG_PX_SIZE), anti_aliasing=True)\n#img = Image.open(image)\n#img = img.resize((299,299))\nimg = img_to_array(resized_img)\nimg = img_to_array(img)\n#img = img.convert(\"RGBA\")\nprint (img.shape)\n#img = img.repeat(3,axis=-1)\n#img = img_to_array(img)\n\n#img = cv2.cvtColor(img, cv2.COLOR_RGB2GRAY)\n#img = cv2.cvtColor(img,cv2.COLOR_GRAY2RGB)\n#img = gaussian_filter2d(img)\n#print (img.shape)\n#img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)\nplt.imshow(img,cmap='gray')\n#plt.axis('on')","metadata":{"execution":{"iopub.status.busy":"2023-02-16T17:39:16.000379Z","iopub.execute_input":"2023-02-16T17:39:16.001487Z","iopub.status.idle":"2023-02-16T17:39:16.905509Z","shell.execute_reply.started":"2023-02-16T17:39:16.001419Z","shell.execute_reply":"2023-02-16T17:39:16.904345Z"}}},{"cell_type":"markdown","source":"img = gaussian_filter2d(img)\nprint (img.shape)\n#img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)\nplt.imshow(img,cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2023-02-16T02:33:54.074099Z","iopub.execute_input":"2023-02-16T02:33:54.075209Z","iopub.status.idle":"2023-02-16T02:33:54.239554Z","shell.execute_reply.started":"2023-02-16T02:33:54.075165Z","shell.execute_reply":"2023-02-16T02:33:54.238835Z"}}},{"cell_type":"markdown","source":"img = tfclahe(img)\nplt.imshow(img,cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2023-02-16T02:34:27.766429Z","iopub.execute_input":"2023-02-16T02:34:27.767087Z","iopub.status.idle":"2023-02-16T02:34:27.911397Z","shell.execute_reply.started":"2023-02-16T02:34:27.767043Z","shell.execute_reply":"2023-02-16T02:34:27.910686Z"}}},{"cell_type":"markdown","source":"mg = adjust_gamma(img)\nplt.imshow(img,cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2023-02-16T02:35:23.131705Z","iopub.execute_input":"2023-02-16T02:35:23.132134Z","iopub.status.idle":"2023-02-16T02:35:23.274916Z","shell.execute_reply.started":"2023-02-16T02:35:23.132103Z","shell.execute_reply":"2023-02-16T02:35:23.273648Z"}}},{"cell_type":"markdown","source":"image = train_oversampled['full_path'][0]\n#print(image)\nimg = Image.open(image)\n#img = cv2.imread(image)\nimg = img.resize((299,299))\nimg = img_to_array(img)\nprint (img.shape)\nimg = cv2.cvtColor(img,cv2.COLOR_GRAY2RGB)\nplt.imshow(img)","metadata":{"execution":{"iopub.status.busy":"2023-02-16T01:47:00.331648Z","iopub.status.idle":"2023-02-16T01:47:00.331984Z","shell.execute_reply.started":"2023-02-16T01:47:00.331824Z","shell.execute_reply":"2023-02-16T01:47:00.331839Z"}}},{"cell_type":"code","source":"# implementing Gaussian Filter\nimport tensorflow as tf\nfrom tensorflow_addons.image import utils as img_utils\nfrom tensorflow_addons.utils import keras_utils\nfrom tensorflow_addons.utils.types import TensorLike\n\nfrom typing import Optional, Union, List, Tuple, Iterable\n\n@tf.function\ndef gaussian_filter2d(\n    image: TensorLike,\n    filter_shape: Union[int, Iterable[int]] = (3, 3),\n    sigma: Union[List[float], Tuple[float], float] = 1.0,\n    padding: str = \"REFLECT\",\n    constant_values: TensorLike = 0,\n    name: Optional[str] = None,\n) -> TensorLike:\n    \n    \n    \n    def _pad(\n    image: TensorLike,\n    filter_shape: Union[List[int], Tuple[int]],\n    mode: str = \"CONSTANT\",\n    constant_values: TensorLike = 0,\n        ) -> tf.Tensor:\n        if mode.upper() not in {\"REFLECT\", \"CONSTANT\", \"SYMMETRIC\"}:\n            raise ValueError(\n                'padding should be one of \"REFLECT\", \"CONSTANT\", or \"SYMMETRIC\".'\n                )\n        constant_values = tf.convert_to_tensor(constant_values, image.dtype)\n        filter_height, filter_width = filter_shape\n        pad_top = (filter_height - 1) // 2\n        pad_bottom = filter_height - 1 - pad_top\n        pad_left = (filter_width - 1) // 2\n        pad_right = filter_width - 1 - pad_left\n        paddings = [[0, 0], [pad_top, pad_bottom], [pad_left, pad_right], [0, 0]]\n        return tf.pad(image, paddings, mode=mode, constant_values=constant_values)\n        \n    def _get_gaussian_kernel(sigma, filter_shape):\n        sigma = tf.convert_to_tensor(sigma)\n        x = tf.range(-filter_shape // 2 + 1, filter_shape // 2 + 1)\n        x = tf.cast(x**2, sigma.dtype)\n        x = tf.nn.softmax(-x / (2.0 * (sigma**2)))\n        return x\n    \n    def _get_gaussian_kernel_2d(gaussian_filter_x, gaussian_filter_y):\n        gaussian_kernel = tf.matmul(gaussian_filter_x, gaussian_filter_y)\n        return gaussian_kernel\n    \n   \n    \n    \n    with tf.name_scope(name or \"gaussian_filter2d\"):\n        if isinstance(sigma, (list, tuple)):\n            if len(sigma) != 2:\n                raise ValueError(\"sigma should be a float or a tuple/list of 2 floats\")\n        else:\n            sigma = (sigma,) * 2\n\n        if any(s < 0 for s in sigma):\n            raise ValueError(\"sigma should be greater than or equal to 0.\")\n\n        image = tf.convert_to_tensor(image, name=\"image\")\n        sigma = tf.convert_to_tensor(sigma, name=\"sigma\")\n\n        original_ndims = img_utils.get_ndims(image)\n        image = img_utils.to_4D_image(image)\n\n        # Keep the precision if it's float;\n        # otherwise, convert to float32 for computing.\n        orig_dtype = image.dtype\n        if not image.dtype.is_floating:\n            image = tf.cast(image, tf.float32)\n\n        channels = tf.shape(image)[3]\n        filter_shape = keras_utils.normalize_tuple(filter_shape, 2, \"filter_shape\")\n\n        sigma = tf.cast(sigma, image.dtype)\n        gaussian_kernel_x = _get_gaussian_kernel(sigma[1], filter_shape[1])\n        gaussian_kernel_x = gaussian_kernel_x[tf.newaxis, :]\n\n        gaussian_kernel_y = _get_gaussian_kernel(sigma[0], filter_shape[0])\n        gaussian_kernel_y = gaussian_kernel_y[:, tf.newaxis]\n\n        gaussian_kernel_2d = _get_gaussian_kernel_2d(\n            gaussian_kernel_y, gaussian_kernel_x\n        )\n        gaussian_kernel_2d = gaussian_kernel_2d[:, :, tf.newaxis, tf.newaxis]\n        gaussian_kernel_2d = tf.tile(gaussian_kernel_2d, [1, 1, channels, 1])\n\n        image = _pad(image, filter_shape, mode=padding, constant_values=constant_values)\n\n        output = tf.nn.depthwise_conv2d(\n            input=image,\n            filter=gaussian_kernel_2d,\n            strides=(1, 1, 1, 1),\n            padding=\"VALID\",\n        )\n        output = img_utils.from_4D_image(output, original_ndims)\n        return tf.cast(output, orig_dtype)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.620766Z","iopub.execute_input":"2023-02-20T20:03:27.621443Z","iopub.status.idle":"2023-02-20T20:03:27.824089Z","shell.execute_reply.started":"2023-02-20T20:03:27.621407Z","shell.execute_reply":"2023-02-20T20:03:27.822896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def adjust_gamma(image):\n    return tf.image.adjust_gamma(\n    image, gamma=0.2, gain=1\n    )","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.82589Z","iopub.execute_input":"2023-02-20T20:03:27.826313Z","iopub.status.idle":"2023-02-20T20:03:27.838058Z","shell.execute_reply.started":"2023-02-20T20:03:27.826275Z","shell.execute_reply":"2023-02-20T20:03:27.83707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#import tf_clahe\n#@tf.function(experimental_compile=True) \ndef tfclahe(img):\n    return tf_clahe.clahe(img)#,gpu_optimized=True)#return tf_clahe.clahe(img)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.839774Z","iopub.execute_input":"2023-02-20T20:03:27.840198Z","iopub.status.idle":"2023-02-20T20:03:27.850492Z","shell.execute_reply.started":"2023-02-20T20:03:27.840162Z","shell.execute_reply":"2023-02-20T20:03:27.849604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_train_generator(df, image_dir, x_col, y_cols, shuffle=True, batch_size=128, seed=1,\n                        target_w = 299, target_h = 299):\n    print(\"getting train generator...\") \n    # normalize images\n    image_generator = ImageDataGenerator(\n        preprocessing_function = gaussian_filter2d,#tfclahe,#adjust_gamma,#gaussian_blur,\n        #samplewise_center=True,\n        #samplewise_std_normalization= True,\n        rescale = 1.0/255.0,\n        rotation_range = 5,\n        #width_shift_range = 0.2,\n        #height_shift_range = 0.2,\n        #shear_range = 0.1,\n        #zoom_range = 0.1,\n        #horizontal_flip = True,\n        fill_mode='nearest',\n     #gaussian_filter2d\n    )\n    \n    # flow from directory with specified batch size\n    # and target image size\n    generator = image_generator.flow_from_dataframe(\n            dataframe=df,\n            directory=image_dir,\n            x_col=x_col,\n            y_col=y_cols,\n            class_mode=\"raw\",\n            color_mode='rgb', #\"grayscale\"\n            batch_size=batch_size,\n            shuffle=shuffle,\n            seed=seed,\n            target_size=(target_w,target_h))\n    \n    return generator","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.852223Z","iopub.execute_input":"2023-02-20T20:03:27.852577Z","iopub.status.idle":"2023-02-20T20:03:27.861147Z","shell.execute_reply.started":"2023-02-20T20:03:27.852543Z","shell.execute_reply":"2023-02-20T20:03:27.860221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_test_and_valid_generator(valid_df, test_df, train_df,train_dir,test_dir, val_dir, x_col, y_cols,\n                                 sample_size=300, batch_size=128, seed=1, target_w = 299, target_h =299):\n    print(\"getting train and valid generators...\")\n    # get generator to sample dataset\n    raw_train_generator = ImageDataGenerator().flow_from_dataframe(\n        dataframe=train_df, \n        directory=train_dir, \n        x_col=\"png_images\", \n        y_col='cancer', \n        class_mode=\"raw\", \n        color_mode='rgb',#\"grayscale\",#rgb',\n        batch_size=sample_size, \n        shuffle=True, \n        target_size=(target_w, target_h))\n    \n    # get data sample\n    batch = raw_train_generator.next()\n    data_sample = batch[0]\n\n    # use sample to fit mean and std for test set generator\n    image_generator = ImageDataGenerator(\n        preprocessing_function = gaussian_filter2d,#tfclahe,#adjust_gamma,#,#gaussian_blur,\n        rescale = 1./255.,\n        #featurewise_center=True,\n        #featurewise_std_normalization= True,\n        rotation_range = 5,\n        #width_shift_range = 0.2,\n        #height_shift_range = 0.2,\n        shear_range = 0.1,\n        zoom_range = 0.1,\n        #horizontal_flip = True,\n        fill_mode='nearest',\n        #gaussian_filter2d\n    )\n    \n    # fit generator to sample from training data\n    #image_generator.fit(data_sample)\n\n    # get test generator\n    valid_generator = image_generator.flow_from_dataframe(\n            dataframe=valid_df,\n            directory=val_dir,\n            x_col=x_col,\n            y_col=y_cols,\n            class_mode=\"raw\",\n            color_mode='rgb',#\"grayscale\",#rgb',\n            batch_size=batch_size,\n            shuffle=False,\n            seed=seed,\n            target_size=(target_w,target_h))\n\n    test_generator = image_generator.flow_from_dataframe(\n            dataframe=test_df,\n            directory=test_dir,\n            x_col=x_col,\n            y_col=y_cols,\n            class_mode=\"raw\",\n            color_mode='rgb', #\"grayscale\"\n            batch_size=batch_size,\n            shuffle=False,\n            seed=seed,\n            target_size=(target_w,target_h))\n    return valid_generator, test_generator","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.862586Z","iopub.execute_input":"2023-02-20T20:03:27.863035Z","iopub.status.idle":"2023-02-20T20:03:27.875224Z","shell.execute_reply.started":"2023-02-20T20:03:27.862999Z","shell.execute_reply":"2023-02-20T20:03:27.874219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.shape,test_df.shape,valid_df.shape,train_quarter.shape,test_quarter.shape,valid_quarter.shape","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.876753Z","iopub.execute_input":"2023-02-20T20:03:27.877122Z","iopub.status.idle":"2023-02-20T20:03:27.887771Z","shell.execute_reply.started":"2023-02-20T20:03:27.87708Z","shell.execute_reply":"2023-02-20T20:03:27.886821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_oversampled.shape,test_oversampled.shape,valid_oversampled.shape","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.889174Z","iopub.execute_input":"2023-02-20T20:03:27.889525Z","iopub.status.idle":"2023-02-20T20:03:27.899034Z","shell.execute_reply.started":"2023-02-20T20:03:27.889491Z","shell.execute_reply":"2023-02-20T20:03:27.897906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_generator_oversampled = get_train_generator(train_quarter, train_dir, \"png_images\", 'cancer')\n\nvalid_generator_oversampled, test_generator_oversampled = get_test_and_valid_generator(valid_quarter,\\\n                                                            test_quarter,\n                                                          train_quarter,train_dir,test_dir,val_dir,\n                                                              \"png_images\", 'cancer')","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:27.900476Z","iopub.execute_input":"2023-02-20T20:03:27.901461Z","iopub.status.idle":"2023-02-20T20:03:45.031905Z","shell.execute_reply.started":"2023-02-20T20:03:27.901424Z","shell.execute_reply":"2023-02-20T20:03:45.030205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# For shape (224,224) + Gaussian + Data Aug with 5 rotation.\nx,y = train_generator_oversampled.__getitem__(0)\ni =0\nprint ('Label',y[i])\nplt.imshow(x[i])","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:45.033489Z","iopub.execute_input":"2023-02-20T20:03:45.033861Z","iopub.status.idle":"2023-02-20T20:03:52.028498Z","shell.execute_reply.started":"2023-02-20T20:03:45.033807Z","shell.execute_reply":"2023-02-20T20:03:52.027042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# For shape (224,224) + Gaussian + Data Aug with 5 rotation.\nx,y = train_generator_oversampled.__getitem__(0)\ni =10\nprint ('Label',y[i])\nplt.imshow(x[i])","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:52.029719Z","iopub.execute_input":"2023-02-20T20:03:52.030013Z","iopub.status.idle":"2023-02-20T20:03:55.222481Z","shell.execute_reply.started":"2023-02-20T20:03:52.029987Z","shell.execute_reply":"2023-02-20T20:03:55.221544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.optimizers.schedules import PolynomialDecay\n\n#batch_size = 32\nbatch_size = 128 #16 * tpu_strategy.num_replicas_in_sync\nnum_epochs = 1\n# The number of training steps is the number of samples in the dataset, divided by the batch size then multiplied\n# by the total number of epochs. Note that the tf_train_dataset here is a batched tf.data.Dataset,\n# not the original Hugging Face Dataset, so its len() is already num_samples // batch_size.\nnum_train_steps = (train_oversampled.shape[0]//batch_size) * num_epochs\nlr_scheduler = PolynomialDecay(\n    initial_learning_rate =5e-5, end_learning_rate=0.00, decay_steps=num_train_steps\n)\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.optimizers import SGD\nfrom tensorflow.keras.optimizers import RMSprop\n\n#opt = Adam(learning_rate=lr_scheduler)\n#opt = Adam(learning_rate=lr_scheduler)\nopt = Adam(learning_rate=lr_scheduler)#,momentum = 0.9)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:55.22411Z","iopub.execute_input":"2023-02-20T20:03:55.224463Z","iopub.status.idle":"2023-02-20T20:03:55.235948Z","shell.execute_reply.started":"2023-02-20T20:03:55.224428Z","shell.execute_reply":"2023-02-20T20:03:55.234866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class pFBeta(tf.keras.metrics.Metric):\n    def __init__(self, beta=1, name='pF1', **kwargs):\n        super().__init__(name=name, **kwargs)\n        self.beta = beta\n        self.epsilon = 1e-10\n        self.pos = self.add_weight(name='pos', initializer='zeros')\n        self.ctp = self.add_weight(name='ctp', initializer='zeros')\n        self.cfp = self.add_weight(name='cfp', initializer='zeros')\n\n    def update_state(self, y_true, y_pred, sample_weight=None):\n        y_true = tf.cast(y_true, tf.float32)\n        y_pred = tf.clip_by_value(y_pred, 0, 1)\n        pos = tf.cast(tf.reduce_sum(y_true), tf.float32)\n        ctp = tf.cast(tf.reduce_sum(y_pred[y_true == 1]), tf.float32)\n        cfp = tf.cast(tf.reduce_sum(y_pred[y_true == 0]), tf.float32)\n        self.pos.assign_add(pos)\n        self.ctp.assign_add(ctp)\n        self.cfp.assign_add(cfp)\n\n    def result(self):\n        beta2 = self.beta * self.beta\n        prec = self.ctp / (self.ctp + self.cfp + self.epsilon)\n        reca = self.ctp / (self.pos + self.epsilon)\n        return (1 + beta2) * prec * reca / (beta2 * prec + reca)\n\n    def reset_state(self):\n        self.pos.assign(0.)\n        self.ctp.assign(0.)\n        self.cfp.assign(0.)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:55.237887Z","iopub.execute_input":"2023-02-20T20:03:55.238509Z","iopub.status.idle":"2023-02-20T20:03:55.252901Z","shell.execute_reply.started":"2023-02-20T20:03:55.23847Z","shell.execute_reply":"2023-02-20T20:03:55.251902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data is somewhat balanced as some chunk of data is extracted from oversampled_data. AT moment, using binary_cross_entropy. maybe if model is showing something, then use other loss to imporve loss or accuracy","metadata":{}},{"cell_type":"code","source":"#del model","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:55.254481Z","iopub.execute_input":"2023-02-20T20:03:55.254877Z","iopub.status.idle":"2023-02-20T20:03:55.266064Z","shell.execute_reply.started":"2023-02-20T20:03:55.25482Z","shell.execute_reply":"2023-02-20T20:03:55.265116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create the base pre-trained model\nimport tensorflow_addons as tfa # DenseNet121\ntf.keras.backend.clear_session()\n#enable XLA optmizations\ntf.config.optimizer.set_jit(True)\ndensenet169_weights = '/kaggle/input/densenet169-weights/densenet169_weights_tf_dim_ordering_tf_kernels_notop.h5'\ndensenet121_weights = '/kaggle/input/desnet121-weights-notop/densenet121_weights_tf_dim_ordering_tf_kernels_notop.h5'\nresnet50_weights = '/kaggle/input/resnet50-weights/resnet50_weights_tf_dim_ordering_tf_kernels_notop.h5'\nresnet101_weights = '/kaggle/input/resnet101-weights/resnet101_weights_tf_dim_ordering_tf_kernels_notop.h5'\ninception_resnet_weights = '/kaggle/input/inception-resnet-v2-weights/inception_resnet_v2_weights_tf_dim_ordering_tf_kernels_notop.h5'\nefficientnetb7_weights = '/kaggle/input/efficientnetb7-weights/efficientnetb7_notop.h5'\nefficientnetb0_weights = '/kaggle/input/efficientnetb0-weights/efficientnetb0_notop.h5'\nbase_model = tf.keras.applications.inception_resnet_v2.InceptionResNetV2(input_shape = (299,299,3),\n                              weights=inception_resnet_weights, include_top=False)\n\n# Make all the layers in the pre-trained model non-trainable\nfor layer in base_model.layers:\n    layer.trainable = False\n\n#K.set_learning_phase(1)\nx = base_model.output\n\n# add a global spatial average pooling layer\nx = GlobalAveragePooling2D()(x)\n\n# 1st conv block\n#x = tf.keras.layers.Conv2D(256, 3, padding='same')(x)\n#x = tf.keras.layers.BatchNormalization()(x)\n#x = tf.keras.layers.Activation('relu')(x)\n#x = tf.keras.layers.GlobalAveragePooling2D(keepdims = True)(x)\n\n# 2nd conv block\n#x = tf.keras.layers.Conv2D(128, 3, padding='same')(x)\n#x = tf.keras.layers.BatchNormalization()(x)\n#x = tf.keras.layers.Activation('relu')(x)\n#x = tf.keras.layers.GlobalAveragePooling2D(keepdims = True)(x)\n\n# 1st FC layer\n#x = tf.keras.layers.Flatten()(x) \n#x = tf.keras.layers.Dense(64)(x)\n#x = tf.keras.layers.BatchNormalization()(x)\n#x = tf.keras.layers.Activation('relu')(x)\n\n# 2nd FC layer\n#x = tf.keras.layers.Dense(32, activation = 'relu')(x)\n#x = tf.keras.layers.BatchNormalization()(x)\n#x = tf.keras.layers.Activation('relu')(x)\n#x = tf.keras.layers.Dropout(.2)(x)\n\n#x = tf.keras.layers.Dense(1, 'sigmoid')(x)\n# x = layers.Flatten()(base_model.output)\n# add further layers to tackle overfitting\nx = Dense(1024, activation='relu', name='fc1')(x)\nx = Dropout(0.2,name='drop')(x)\n#x = tf.keras.layers.GaussianNoise(0.2)(x)\n\n# and a logistic layer\npredictions = Dense(1, activation=\"sigmoid\")(x)\n\nmodel = Model(inputs=base_model.input, outputs=predictions)\nmodel.compile(optimizer=opt, \n              loss=[{#'get_weighted_loss' : get_weighted_loss(pos_weights, neg_weights)}],\n                     #'BinaryFocalLoss':focal_loss(),\n              'binary_crossentropy':'binary_crossentropy'}],\n             metrics=['accuracy',tfa.metrics.F1Score(num_classes=1, threshold=0.5),\n                      pFBeta(beta=1, name='pF1')])\n                   \n        #['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:03:55.26936Z","iopub.execute_input":"2023-02-20T20:03:55.269681Z","iopub.status.idle":"2023-02-20T20:04:03.273576Z","shell.execute_reply.started":"2023-02-20T20:03:55.269654Z","shell.execute_reply":"2023-02-20T20:04:03.272618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model Results\n\n## Resnet 50\n### bs=64,opt=sgd. Epoch 10/15\n582/582 [==============================] - 387s 664ms/step - loss: 0.0382 - accuracy: 0.9981 - f1_score: 0.9981 - pF1: 0.9647 - val_loss: 1.7631 - val_accuracy: 0.5145 - val_f1_score: 0.1035 - val_pF1: 0.1870. TP=96,FP=12,TN=108,FN=0\n### bs=64,opt=RMSprop. worst results\n### bs=64,opt=adam. worst results.\n### bs=64,opt = sgd, adding two more layers in model. Worst results.\n### bs=32,opt=sgd,just dropout, failed. \n### bs=32,opt=sgd,gaussian,failed.\n### bs=32,opt=sgd,guassian + layer gaussian. failed.\n### bs=32,opt=sgd,guassian + Data Aug. Epoch 1/3\n2023-02-10 16:14:12.619380: I tensorflow/stream_executor/cuda/cuda_dnn.cc:369] Loaded cuDNN version 8005\n1164/1164 [==============================] - 1352s 1s/step - loss: 0.6373 - accuracy: 0.6370 - f1_score: 0.6362 - pF1: 0.5761 - val_loss: 0.7014 - val_accuracy: 0.6263 - val_f1_score: 0.5090 - val_pF1: 0.5180\nEpoch 2/3\n1164/1164 [==============================] - 1168s 1s/step - loss: 0.4889 - accuracy: 0.7590 - f1_score: 0.7612 - pF1: 0.6860 - val_loss: 0.7731 - val_accuracy: 0.5923 - val_f1_score: 0.5528 - val_pF1: 0.5452\nEpoch 3/3\n1164/1164 [==============================] - 1151s 989ms/step - loss: 0.3912 - accuracy: 0.8234 - f1_score: 0.8245 - pF1: 0.7554 - val_loss: 0.8508 - val_accuracy: 0.6042 - val_f1_score: 0.5348 - val_pF1: 0.5251\nCPU times: user 58min 10s, sys: 1min 35s, total: 59min 45s\n### giving good results onyl for 1 epochs. I have run for 6 epochs, the score become worst. For 2 epochs, the probs are varing but nothing predicting cancer patients. All no cancer patients with false positves too.\n\n## Inception-Resnetv2\n### bs=32,opt=sgd,gaussian + Data AUg, run for 11 epochs, the model was learning very slowly. on 11 peochs, the probabilities was very low. But on first 5-6 epochs, it was giving good results.Freezing layers is showing some progress in overcome overfitting. I didn't add further layers like Dense & dropout. just use global average pooling layer and then fed into sigmoid layer.\n### bs=32,opt=sgd,gamma + Data Aug,freezing all layers plus without adding further layers. On three epochs, the results are worst.\n### bs=32,opt=sgd,gamma + Data Aug, unfreezing layers. overfitting. don't use wihtout freezling.\n\n## DenseNet169\n## bs=32,opt=sgd,image_Den = minor Data Aug with blur function. AFter reached to 3 epochs, no such learning \n### bs=32,opt=sgd,image_Den = minor Data Aug with no custom function. On 3 epcohs, its giving me 47%. \n\n## EfficientNetb7\n### bs=32,opt='adam',image_Den=minor Data Aug with gamma custom function. It was giving me smooth outcomes when after 10 epochs, the model was not learning at all.\n\n## EfficientNetb0\n### bs=32,opt=adam,image_den= minior Data Aug with clahe custom function. not working.\n\n## Inception_resnetv2:\n### bs=64,opt=adam,data_img = Minro Data Aug + Clahe. Its working to some extent.","metadata":{}},{"cell_type":"markdown","source":"checkpoint_path = \"training_1/cp.ckpt\"\ncheckpoint_dir = os.path.dirname(checkpoint_path)\n\n# Create a callback that saves the model's weights\ncp_callback = tf.keras.callbacks.ModelCheckpoint(filepath=checkpoint_path,\n                                                 save_weights_only=True,\n                                                 verbose=1)\n#os.listdir(checkpoint_dir)","metadata":{}},{"cell_type":"markdown","source":"# Include the epoch in the file name (uses `str.format`)\ncheckpoint_path = \"training_2/cp-{epoch:04d}.ckpt\"\ncheckpoint_dir = os.path.dirname(checkpoint_path)\n\nbatch_size = 32\n\n# Create a callback that saves the model's weights every 5 epochs\ncp_callback = tf.keras.callbacks.ModelCheckpoint(\n    filepath=checkpoint_path, \n    verbose=1, \n    save_weights_only=True,\n    save_freq=5*batch_size)\n\n# Create a new model instance\nmodel = create_model()\n\n# Save the weights using the `checkpoint_path` format\nmodel.save_weights(checkpoint_path.format(epoch=0))\n\nlatest = tf.train.latest_checkpoint(checkpoint_dir)\nlatest\n\n# Create a new model instance\nmodel = create_model() # that would not be necessary! check it.\n\n# Load the previously saved weights\nmodel.load_weights(latest)","metadata":{}},{"cell_type":"code","source":"#del model","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:04:03.275094Z","iopub.execute_input":"2023-02-20T20:04:03.275435Z","iopub.status.idle":"2023-02-20T20:04:03.280273Z","shell.execute_reply.started":"2023-02-20T20:04:03.275394Z","shell.execute_reply":"2023-02-20T20:04:03.279052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#import tensorflow as tf\n#physical_devices = tf.config.list_physical_devices('GPU')\n#tf.config.experimental.set_memory_growth(physical_devices[0], enable=True)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:04:03.281985Z","iopub.execute_input":"2023-02-20T20:04:03.282563Z","iopub.status.idle":"2023-02-20T20:04:03.291498Z","shell.execute_reply.started":"2023-02-20T20:04:03.282526Z","shell.execute_reply":"2023-02-20T20:04:03.290567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# inception resnet with gaussian2d\nfrom keras.callbacks import ModelCheckpoint, Callback, EarlyStopping\nes = EarlyStopping(monitor='val_loss', mode='min', verbose=1,patience=5)\n#es = EarlyStopping(patience=5, restore_best_weights=True)\n# checkpoint to save model\n#chkpt = ModelCheckpoint(filepath=\"model1\", save_best_only=True)\nSTEPS_PER_EPOCH = train_oversampled.shape[0] // batch_size\n#VAL_SUBSPLITS = 5\nVALIDATION_STEPS = test_oversampled.shape[0]//batch_size #//VAL_SUBSPLITS\n# First 5\nhistory = model.fit(train_generator_oversampled,\n                    validation_data=test_generator_oversampled,\n                 #steps_per_epoch=STEPS_PER_EPOCH,validation_steps=VALIDATION_STEPS,\n                 callbacks=[es],#[es,chkpt],class_weight={0:1.0, 1:0.33},\n                 epochs = 1,verbose=1)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:04:03.293151Z","iopub.execute_input":"2023-02-20T20:04:03.293521Z","iopub.status.idle":"2023-02-20T20:08:00.855116Z","shell.execute_reply.started":"2023-02-20T20:04:03.293487Z","shell.execute_reply":"2023-02-20T20:08:00.853929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.evaluate(valid_generator_oversampled)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:08:00.856958Z","iopub.execute_input":"2023-02-20T20:08:00.857355Z","iopub.status.idle":"2023-02-20T20:08:08.895644Z","shell.execute_reply.started":"2023-02-20T20:08:00.857314Z","shell.execute_reply":"2023-02-20T20:08:08.894743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#!mkdir -p saved_model\n#model.save('efficientnetb7.h5')#('save_model/inception_resnetv2_model')","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:08:08.898175Z","iopub.execute_input":"2023-02-20T20:08:08.898539Z","iopub.status.idle":"2023-02-20T20:08:08.904119Z","shell.execute_reply.started":"2023-02-20T20:08:08.898506Z","shell.execute_reply":"2023-02-20T20:08:08.90296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"history = model.fit(train_generator_oversampled,\n                    validation_data=test_generator_oversampled,\n                 steps_per_epoch=STEPS_PER_EPOCH,validation_steps=VALIDATION_STEPS,\n                 callbacks=[es],#[es,chkpt],class_weight={0:1.0, 1:0.33},\n                 epochs = 2)","metadata":{}},{"cell_type":"markdown","source":"history = model.fit(train_generator_oversampled,\n                    validation_data=test_generator_oversampled,\n                 steps_per_epoch=STEPS_PER_EPOCH,validation_steps=VALIDATION_STEPS,\n                 callbacks=[es],#[es,chkpt],class_weight={0:1.0, 1:0.33},\n                 epochs = 1)","metadata":{}},{"cell_type":"markdown","source":"history = model.fit(train_generator_oversampled,\n                    validation_data=test_generator_oversampled,\n                 steps_per_epoch=STEPS_PER_EPOCH,validation_steps=VALIDATION_STEPS,\n                 callbacks=[es],#[es,chkpt],class_weight={0:1.0, 1:0.33},\n                 epochs = 1)","metadata":{}},{"cell_type":"markdown","source":"%%time\n# 11 epochs\nhistory = model.fit(train_generator_oversampled,\n                    validation_data=test_generator_oversampled,\n                 steps_per_epoch=STEPS_PER_EPOCH,validation_steps=VALIDATION_STEPS,\n                 callbacks=[es],#[es,chkpt],class_weight={0:1.0, 1:0.33},\n                 epochs = 5)","metadata":{}},{"cell_type":"markdown","source":"%%time\n# with bs=64,opt=RMSProp\nfrom keras.callbacks import ModelCheckpoint, Callback, EarlyStopping\nes = EarlyStopping(monitor='val_loss', mode='min', verbose=1,patience=90)\n#es = EarlyStopping(patience=5, restore_best_weights=True)\n# checkpoint to save model\n#chkpt = ModelCheckpoint(filepath=\"model1\", save_best_only=True)\nSTEPS_PER_EPOCH = train_oversampled.shape[0] // batch_size\n#VAL_SUBSPLITS = 5\nVALIDATION_STEPS = test_oversampled.shape[0]//batch_size #//VAL_SUBSPLITS\n# First 5\nhistory = model.fit(train_generator_oversampled,\n                    validation_data=test_generator_oversampled,\n                 steps_per_epoch=STEPS_PER_EPOCH,validation_steps=VALIDATION_STEPS,\n                 callbacks=[es],#[es,chkpt],class_weight={0:1.0, 1:0.33},\n                 epochs = 3)","metadata":{}},{"cell_type":"markdown","source":"### tf.image.adjust_contrast\n### tf.image.adjust_hue\n### tf.image.adjust_gamma","metadata":{}},{"cell_type":"code","source":"# Graphs For \"Loss\" and \"AUC\"\ndef plot_graphs(history, metric):\n    plt.plot(history.history[metric])\n    plt.plot(history.history[f'val_{metric}'])\n    plt.xlabel(\"Epochs\")\n    plt.ylabel(metric)\n    plt.legend([metric, f'val_{metric}'])\n    plt.show()\n#plot_graphs(history, \"accuracy\")\n#plot_graphs(history,pFBeta)\nplot_graphs(history, \"loss\")","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:08:08.905836Z","iopub.execute_input":"2023-02-20T20:08:08.906218Z","iopub.status.idle":"2023-02-20T20:08:09.113776Z","shell.execute_reply.started":"2023-02-20T20:08:08.906163Z","shell.execute_reply":"2023-02-20T20:08:09.112801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/test.csv')\n#sample_dir = '/kaggle/input/samples-test-rsna-images'\nsample_test_dir = '/kaggle/input/rsna-breast-cancer-detection/test_images'\ntest['images'] = [(sample_test_dir + '/' + str(k) + '/' + str(v) + '.dcm') \\\n                  for k,v in test[['patient_id','image_id']].values]\nprint ('train Data Size',test.shape)\ntest.head() ","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:02:01.749534Z","iopub.execute_input":"2023-02-20T21:02:01.749933Z","iopub.status.idle":"2023-02-20T21:02:01.772382Z","shell.execute_reply.started":"2023-02-20T21:02:01.749899Z","shell.execute_reply":"2023-02-20T21:02:01.771383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"%%time\nimport dicomsdl as dicom, time\nIMG_PX_SIZE = 299\nimage  =test['images'][0]\nds = pydicom.dcmread(image)#,force=True) # dicom.open(image)#p\ndata = ds.pixel_array #pixel_array#\nresized_img = resize(data, (IMG_PX_SIZE, IMG_PX_SIZE), anti_aliasing=True)\n#img = img_to_array(resized_img)\n#img = img.resize((299,299,1))\n#img = img_to_array(resized_img)\n#print (data.shape)\n#print (type(data))\nimg = resized_img.repeat(3,axis=-1)\nimg = img_to_array(img)\nprint (img.shape)\n#img = cv2.cvtColor(img,cv2.COLOR_GRAY2RGB)\n#print (data.shape)\n#img = gaussian_filter2d(img)\n\nplt.imshow(img,cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2023-02-18T22:22:27.844345Z","iopub.status.idle":"2023-02-18T22:22:27.845774Z","shell.execute_reply.started":"2023-02-18T22:22:27.845403Z","shell.execute_reply":"2023-02-18T22:22:27.845439Z"}}},{"cell_type":"markdown","source":"def predictions(image):\n    import dicomsdl as dicom\n    IMG_PX_SIZE = 299\n    try:\n        ds = dicom.open(image)\n        data = ds.pixelData()\n        resized_img = resize(data, (IMG_PX_SIZE, IMG_PX_SIZE), anti_aliasing=True)\n        img = img_to_array(resized_img)\n        img = cv2.cvtColor(img,cv2.COLOR_GRAY2RGB)\n        x = img/127.0\n        x=np.expand_dims(x, axis=0)\n        images = np.vstack([x])\n        classes = model.predict(images)\n        return classes[0][0]\n    except:\n        print ('Not Working')","metadata":{"execution":{"iopub.status.busy":"2023-02-20T18:07:48.340651Z","iopub.execute_input":"2023-02-20T18:07:48.341357Z","iopub.status.idle":"2023-02-20T18:07:48.348529Z","shell.execute_reply.started":"2023-02-20T18:07:48.341319Z","shell.execute_reply":"2023-02-20T18:07:48.347322Z"}}},{"cell_type":"markdown","source":"%%time\njobs  = [delayed(predictions)(i) for i in test['images']]\ntest['probabilities'] = Parallel(\n    n_jobs=cpu_count(),\n    verbose=IS_INTERACTIVE,\n    #backend='multiprocessing',\n    prefer='threads',\n)(jobs)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T18:07:53.655802Z","iopub.execute_input":"2023-02-20T18:07:53.656366Z","iopub.status.idle":"2023-02-20T18:07:57.021745Z","shell.execute_reply.started":"2023-02-20T18:07:53.656321Z","shell.execute_reply":"2023-02-20T18:07:57.020724Z"}}},{"cell_type":"code","source":"def predictions(image):\n    IMG_PX_SIZE = 299\n    ds = dicom.open(image)\n    data = ds.pixelData()\n    resized_img = resize(data, (IMG_PX_SIZE, IMG_PX_SIZE), anti_aliasing=True)\n    img = img_to_array(resized_img)\n    img = cv2.cvtColor(img,cv2.COLOR_GRAY2RGB)\n    x = img/127.0\n    x=np.expand_dims(x, axis=0)\n    images = np.vstack([x])\n    #classes = model.predict(images)\n    images = tf.convert_to_tensor(images)\n    return model.predict(images)[0][0]#classes[0][0] # predict_on_batch","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:02:03.268601Z","iopub.execute_input":"2023-02-20T21:02:03.268986Z","iopub.status.idle":"2023-02-20T21:02:03.276048Z","shell.execute_reply.started":"2023-02-20T21:02:03.268955Z","shell.execute_reply":"2023-02-20T21:02:03.274707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\njobs  = [delayed(predictions)(i) for i in test['images']]\ntest['probabilities'] = Parallel(\n    n_jobs=cpu_count(),\n    verbose=IS_INTERACTIVE,\n    #backend='multiprocessing',\n    prefer='threads',\n)(jobs)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:02:04.385519Z","iopub.execute_input":"2023-02-20T21:02:04.385909Z","iopub.status.idle":"2023-02-20T21:02:07.81351Z","shell.execute_reply.started":"2023-02-20T21:02:04.385874Z","shell.execute_reply":"2023-02-20T21:02:07.812437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:17:24.745445Z","iopub.execute_input":"2023-02-20T20:17:24.745794Z","iopub.status.idle":"2023-02-20T20:17:24.761129Z","shell.execute_reply.started":"2023-02-20T20:17:24.745758Z","shell.execute_reply":"2023-02-20T20:17:24.759966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n#extract_images = Parallel(n_jobs=4,prefer=\"threads\")(delayed(predictions)(i) for i in test['images'])\n#test['probabilities'] = Parallel(n_jobs=2,prefer=\"threads\")(delayed(model.predict)(i) for i in test['probabilities'])\n#test.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:17:24.762477Z","iopub.execute_input":"2023-02-20T20:17:24.763422Z","iopub.status.idle":"2023-02-20T20:17:24.769793Z","shell.execute_reply.started":"2023-02-20T20:17:24.763385Z","shell.execute_reply":"2023-02-20T20:17:24.768727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#%%time\n#test['predictions'] = [(predictions(i)) for i in test['images']]\n#test.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:17:24.772153Z","iopub.execute_input":"2023-02-20T20:17:24.773153Z","iopub.status.idle":"2023-02-20T20:17:24.781814Z","shell.execute_reply.started":"2023-02-20T20:17:24.773117Z","shell.execute_reply":"2023-02-20T20:17:24.780719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"%%time\nIMG_PX_SIZE = 299\n#ds = dicom.open(test['images'][0])\ndata = dicom.open(test['images'][0]).pixelData() \nresized_img = resize(data, (IMG_PX_SIZE, IMG_PX_SIZE), anti_aliasing=True)\n#resized_img = resize(data(299, 299, 3),anti_aliasing=True)\nimg = img_to_array(resized_img)\nprint (img.shape)\n#img = img.repeat(resized_img,axis=-1)\nimg = cv2.cvtColor(img,cv2.COLOR_GRAY2RGB)\nx = img/255.0\nx=np.expand_dims(x, axis=0)\nimages = np.vstack([x])\nclasses = model.predict(images)\nprint (classes[0][0])","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:35:08.011891Z","iopub.execute_input":"2023-02-20T02:35:08.01228Z","iopub.status.idle":"2023-02-20T02:35:08.941691Z","shell.execute_reply.started":"2023-02-20T02:35:08.012249Z","shell.execute_reply":"2023-02-20T02:35:08.940651Z"}}},{"cell_type":"markdown","source":"%%time\n#extract_images = Parallel(n_jobs=4,prefer=\"threads\")(delayed(predictions)(i) for i in test['images'])\ntest['probabilities'] = Parallel(n_jobs=4,prefer=\"threads\")(delayed(predictions)(i) for i in test['images'])\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:08:32.124572Z","iopub.execute_input":"2023-02-20T20:08:32.124978Z","iopub.status.idle":"2023-02-20T20:08:35.363719Z","shell.execute_reply.started":"2023-02-20T20:08:32.124942Z","shell.execute_reply":"2023-02-20T20:08:35.362574Z"}}},{"cell_type":"markdown","source":"def get_image_train(X_train):\n    image_bytes = tf.io.read_file(X_train)\n    image = tfio.image.decode_dicom_image(image_bytes,color_dim=True, scale='preserve', dtype=tf.uint16)\n    image= tf.image.resize(image, (299, 299))\n    #image = image.numpy()\n    image = tf.image.grayscale_to_rgb(image)\n    image = image/127.0\n    prob = model.predict(image)#,batch_size=1)\n    #image = img_to_array(image)\n    #image = cv2.cvtColor(image[:1,:,],cv2.COLOR_GRAY2RGB)\n    #image = tf.reshape(image,[299,299,3])\n    #image= tf.image.per_image_standardization(image)\n    return prob[0][0]","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:27:57.77697Z","iopub.execute_input":"2023-02-20T20:27:57.777343Z","iopub.status.idle":"2023-02-20T20:27:57.784109Z","shell.execute_reply.started":"2023-02-20T20:27:57.777314Z","shell.execute_reply":"2023-02-20T20:27:57.782906Z"}}},{"cell_type":"markdown","source":"%%time\nIMG_PX_SIZE = 299\n#ds = dicom.open(test['images'][0])\ndata = dicom.open(test['images'][0]).pixelData() \nresized_img = resize(data, (IMG_PX_SIZE, IMG_PX_SIZE), anti_aliasing=True)\nimg = img_to_array(resized_img)\nprint (img.shape)\n#img = img.repeat(resized_img,axis=-1)\nimg = cv2.cvtColor(img,cv2.COLOR_GRAY2RGB)\nx = img/255.0\nx=np.expand_dims(x, axis=0)\nimages = np.vstack([x])\nprint (model.predict(images)[0][0])\n#print (classes[0][0])","metadata":{"execution":{"iopub.status.busy":"2023-02-20T16:40:00.52601Z","iopub.execute_input":"2023-02-20T16:40:00.526486Z","iopub.status.idle":"2023-02-20T16:40:01.478147Z","shell.execute_reply.started":"2023-02-20T16:40:00.526449Z","shell.execute_reply":"2023-02-20T16:40:01.477138Z"}}},{"cell_type":"markdown","source":"def get_image_train(X_train):\n    image_bytes = tf.io.read_file(X_train)\n    image = tfio.image.decode_dicom_image(image_bytes,color_dim=True, scale='preserve', dtype=tf.uint16)\n    image= tf.image.resize(image, (299, 299))\n    image = tf.image.grayscale_to_rgb(image)\n    image = image/127.0\n    prob = model.predict_on_batch(image)\n    return prob[0][0]","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:28:07.833036Z","iopub.execute_input":"2023-02-20T20:28:07.83343Z","iopub.status.idle":"2023-02-20T20:28:07.840554Z","shell.execute_reply.started":"2023-02-20T20:28:07.833397Z","shell.execute_reply":"2023-02-20T20:28:07.839396Z"}}},{"cell_type":"markdown","source":"%%time\nimport tensorflow_io as tfio\njobs  = [delayed(get_image_train)(i) for i in test['images']]\ntest['probabilities'] = Parallel(\n    n_jobs=cpu_count(),\n    verbose=IS_INTERACTIVE,\n    #backend='multiprocessing',\n    prefer='threads',\n)(jobs)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:29:09.81873Z","iopub.execute_input":"2023-02-20T20:29:09.81924Z","iopub.status.idle":"2023-02-20T20:29:09.976814Z","shell.execute_reply.started":"2023-02-20T20:29:09.819201Z","shell.execute_reply":"2023-02-20T20:29:09.975836Z"}}},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:17:27.363962Z","iopub.execute_input":"2023-02-20T20:17:27.364365Z","iopub.status.idle":"2023-02-20T20:17:27.388342Z","shell.execute_reply.started":"2023-02-20T20:17:27.36433Z","shell.execute_reply":"2023-02-20T20:17:27.387221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#get_image_train(test['images'][1])","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:17:27.764207Z","iopub.execute_input":"2023-02-20T20:17:27.764556Z","iopub.status.idle":"2023-02-20T20:17:27.769085Z","shell.execute_reply.started":"2023-02-20T20:17:27.764528Z","shell.execute_reply":"2023-02-20T20:17:27.768032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#%%time\n#test['probabilities']  = [get_image_train(i) for i in test['images']]","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:17:28.0859Z","iopub.execute_input":"2023-02-20T20:17:28.086272Z","iopub.status.idle":"2023-02-20T20:17:28.092518Z","shell.execute_reply.started":"2023-02-20T20:17:28.08624Z","shell.execute_reply":"2023-02-20T20:17:28.090898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"%%time\ntest['probabilities']  = Parallel(n_jobs=cpu_count(),prefer=\"threads\",)(delayed(get_image_train)(i) for i in test['images'])","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:52:29.262116Z","iopub.execute_input":"2023-02-20T02:52:29.262814Z","iopub.status.idle":"2023-02-20T02:52:31.674768Z","shell.execute_reply.started":"2023-02-20T02:52:29.262781Z","shell.execute_reply":"2023-02-20T02:52:31.673695Z"}}},{"cell_type":"markdown","source":"%%time\ndef predictions(image):\n    IMG_PX_SIZE = 299\n    try:\n        ds = pydicom.dcmread(image,force=True)\n        data = ds.pixel_array \n        resized_img = resize(data, (IMG_PX_SIZE, IMG_PX_SIZE), anti_aliasing=True)\n        img = img_to_array(resized_img)\n        #img = img.repeat(img,axis=-1)\n        img = cv2.cvtColor(img,cv2.COLOR_GRAY2RGB)\n        x = img/255.0\n        x=np.expand_dims(x, axis=0)\n        images = np.vstack([x])\n        classes = model.predict(images)\n        return classes[0][0]\n    except:\n        pass\n#predictions(test['images'][0])\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T02:24:15.527417Z","iopub.execute_input":"2023-02-20T02:24:15.527814Z","iopub.status.idle":"2023-02-20T02:24:15.535833Z","shell.execute_reply.started":"2023-02-20T02:24:15.527783Z","shell.execute_reply":"2023-02-20T02:24:15.53448Z"}}},{"cell_type":"markdown","source":"%%time\nIMG_PX_SIZE = 299\n#ds = dicom.open(test['images'][0])\ndata = dicom.open(test['images'][0]).pixelData() \nresized_img = resize(data, (IMG_PX_SIZE, IMG_PX_SIZE), anti_aliasing=True)\nimg = img_to_array(resized_img)\nprint (img.shape)\n#img = img.repeat(resized_img,axis=-1)\nimg = cv2.cvtColor(img,cv2.COLOR_GRAY2RGB)\nx = img/255.0\nx=np.expand_dims(x, axis=0)\nimages = np.vstack([x])\nprint (model.predict(images)[0][0])\n#print (classes[0][0])","metadata":{"execution":{"iopub.status.busy":"2023-02-20T04:09:03.238104Z","iopub.execute_input":"2023-02-20T04:09:03.238481Z","iopub.status.idle":"2023-02-20T04:09:04.218017Z","shell.execute_reply.started":"2023-02-20T04:09:03.238451Z","shell.execute_reply":"2023-02-20T04:09:04.216954Z"}}},{"cell_type":"code","source":"((2.62/4) * 32000)/3600","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:18:50.086449Z","iopub.execute_input":"2023-02-20T20:18:50.086894Z","iopub.status.idle":"2023-02-20T20:18:50.094573Z","shell.execute_reply.started":"2023-02-20T20:18:50.086834Z","shell.execute_reply":"2023-02-20T20:18:50.093282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.drop_duplicates(subset=['prediction_id'],inplace=True)\ntest.reset_index(inplace=True)\ntest.drop(['index'],axis=1,inplace=True)\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:17:30.232915Z","iopub.execute_input":"2023-02-20T20:17:30.23356Z","iopub.status.idle":"2023-02-20T20:17:30.253732Z","shell.execute_reply.started":"2023-02-20T20:17:30.233526Z","shell.execute_reply":"2023-02-20T20:17:30.252869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/sample_submission.csv')\nsample.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:17:30.598874Z","iopub.execute_input":"2023-02-20T20:17:30.599981Z","iopub.status.idle":"2023-02-20T20:17:30.613857Z","shell.execute_reply.started":"2023-02-20T20:17:30.599939Z","shell.execute_reply":"2023-02-20T20:17:30.612778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample.prediction_id = test.prediction_id.copy()\nsample.cancer = test.probabilities.copy()\nsample.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:17:31.152697Z","iopub.execute_input":"2023-02-20T20:17:31.153362Z","iopub.status.idle":"2023-02-20T20:17:31.164412Z","shell.execute_reply.started":"2023-02-20T20:17:31.153327Z","shell.execute_reply":"2023-02-20T20:17:31.163362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample.to_csv('submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:08:35.679882Z","iopub.status.idle":"2023-02-20T20:08:35.680392Z","shell.execute_reply.started":"2023-02-20T20:08:35.680126Z","shell.execute_reply":"2023-02-20T20:08:35.680149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"val_dir = '/kaggle/input/validation-rsna-images'\nvalid_oversampled['png_images1'] = [(val_dir + '/' + str(i) + '.png') for i in valid_oversampled['image_id']]","metadata":{"execution":{"iopub.status.busy":"2023-02-14T23:46:27.286834Z","iopub.execute_input":"2023-02-14T23:46:27.287885Z","iopub.status.idle":"2023-02-14T23:46:27.294987Z","shell.execute_reply.started":"2023-02-14T23:46:27.287836Z","shell.execute_reply":"2023-02-14T23:46:27.293617Z"}}},{"cell_type":"markdown","source":"valid_oversampled['predictions'] = [(predictions(i)[1][0]) for i in valid_oversampled['png_images1']]\nvalid_oversampled['probabilities'] = [(predictions(i)[0][0]) for i in valid_oversampled['png_images1']]\nvalid_oversampled.head(1)","metadata":{"execution":{"iopub.status.busy":"2023-02-14T23:46:28.473005Z","iopub.execute_input":"2023-02-14T23:46:28.473425Z","iopub.status.idle":"2023-02-14T23:49:34.585015Z","shell.execute_reply.started":"2023-02-14T23:46:28.473393Z","shell.execute_reply":"2023-02-14T23:49:34.583922Z"}}},{"cell_type":"markdown","source":"valid_oversampled[['cancer','predictions','probabilities']].head(110)","metadata":{"execution":{"iopub.status.busy":"2023-02-14T23:49:34.588968Z","iopub.execute_input":"2023-02-14T23:49:34.589385Z","iopub.status.idle":"2023-02-14T23:49:34.613387Z","shell.execute_reply.started":"2023-02-14T23:49:34.58935Z","shell.execute_reply":"2023-02-14T23:49:34.612207Z"},"jupyter":{"outputs_hidden":true}}},{"cell_type":"markdown","source":"i=0\nfor k,v in valid_oversampled[['cancer','predictions']].values:\n    if k==v:\n        i+=1\nprint (i/len(valid_oversampled)*100)","metadata":{"execution":{"iopub.status.busy":"2023-02-14T23:50:10.671697Z","iopub.execute_input":"2023-02-14T23:50:10.672387Z","iopub.status.idle":"2023-02-14T23:50:10.681055Z","shell.execute_reply.started":"2023-02-14T23:50:10.672349Z","shell.execute_reply":"2023-02-14T23:50:10.679543Z"}}},{"cell_type":"markdown","source":"# Get the confusion matrix\nfrom sklearn.metrics import confusion_matrix\nfrom mlxtend.plotting import plot_confusion_matrix\ncm  = confusion_matrix(valid_oversampled['cancer'], valid_oversampled['predictions'])\nplt.figure()\nplot_confusion_matrix(cm,figsize=(10,5), hide_ticks=True, cmap=plt.cm.Blues,)\nplt.xticks(range(2), ['No Cancer', 'Cancer'], fontsize=16)\nplt.yticks(range(2), ['No Cancer', 'Cancer'], fontsize=16)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-14T23:50:13.853689Z","iopub.execute_input":"2023-02-14T23:50:13.854094Z","iopub.status.idle":"2023-02-14T23:50:14.130791Z","shell.execute_reply.started":"2023-02-14T23:50:13.854054Z","shell.execute_reply":"2023-02-14T23:50:14.129072Z"}}},{"cell_type":"markdown","source":"from sklearn.metrics import classification_report\n \nprint(classification_report(valid_oversampled['cancer'], valid_oversampled['predictions']))\n","metadata":{"execution":{"iopub.status.busy":"2023-02-14T23:50:23.008419Z","iopub.execute_input":"2023-02-14T23:50:23.010974Z","iopub.status.idle":"2023-02-14T23:50:23.021823Z","shell.execute_reply.started":"2023-02-14T23:50:23.010928Z","shell.execute_reply":"2023-02-14T23:50:23.02101Z"}}},{"cell_type":"code","source":"#for i in range (2):\n #   sample.loc[len(sample.index)] = [i,i] ","metadata":{"execution":{"iopub.status.busy":"2023-02-20T20:08:35.682102Z","iopub.status.idle":"2023-02-20T20:08:35.68259Z","shell.execute_reply.started":"2023-02-20T20:08:35.68234Z","shell.execute_reply":"2023-02-20T20:08:35.682364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}