{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Import relevant libraries <br>\nWe will explore how each library serves its purpose later on. ","metadata":{}},{"cell_type":"code","source":"# General Data Libraries\nimport pandas as pd \nimport numpy as np \nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.134924Z","iopub.execute_input":"2022-08-30T11:27:49.13554Z","iopub.status.idle":"2022-08-30T11:27:49.142019Z","shell.execute_reply.started":"2022-08-30T11:27:49.135475Z","shell.execute_reply":"2022-08-30T11:27:49.14057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Domain Knowledge\n\n### What are Whole Slide Images? \n\nWhole Slide Images are created from a process called Whole Slide Imaging. This process involves scanning a microscopic slide and creating a single high resolution digital file. This digital file can come in many dimensions depending on the research. These dimensions can vary according to microscope zoom levels.","metadata":{}},{"cell_type":"markdown","source":"### Exploratory Data Analysis\n\nA recommended notebook for detailed EDA <br>\nhttps://www.kaggle.com/code/manikanthgoud/mayo-clinic-strip-ai-exploratory-data-analysis","metadata":{}},{"cell_type":"code","source":"seed = 24\nnp.random_seed = 24","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.144475Z","iopub.execute_input":"2022-08-30T11:27:49.145779Z","iopub.status.idle":"2022-08-30T11:27:49.158483Z","shell.execute_reply.started":"2022-08-30T11:27:49.145723Z","shell.execute_reply":"2022-08-30T11:27:49.15753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv(\"../input/mayo-clinic-strip-ai/train.csv\")\ntest_df = pd.read_csv(\"../input/mayo-clinic-strip-ai/test.csv\")\nother_df = pd.read_csv(\"../input/mayo-clinic-strip-ai/other.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.160071Z","iopub.execute_input":"2022-08-30T11:27:49.160425Z","iopub.status.idle":"2022-08-30T11:27:49.186178Z","shell.execute_reply.started":"2022-08-30T11:27:49.160393Z","shell.execute_reply":"2022-08-30T11:27:49.185254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.18851Z","iopub.execute_input":"2022-08-30T11:27:49.189209Z","iopub.status.idle":"2022-08-30T11:27:49.204807Z","shell.execute_reply.started":"2022-08-30T11:27:49.189171Z","shell.execute_reply":"2022-08-30T11:27:49.203121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Memory saving function credit to https://www.kaggle.com/gemartin/load-data-reduce-memory-usage\n# Avoid memory intensive tasks which could cause the kernel to restart\n\ndef reduce_mem_usage(df): \n    for col in df.columns:\n        col_type = df[col].dtype\n\n        if col_type != object:\n            c_min = df[col].min()\n            c_max = df[col].max()\n            if str(col_type)[:3] == 'int':\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)  \n            \n            else:\n                if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n\n    return df.head(3)","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.20607Z","iopub.execute_input":"2022-08-30T11:27:49.207209Z","iopub.status.idle":"2022-08-30T11:27:49.2224Z","shell.execute_reply.started":"2022-08-30T11:27:49.207157Z","shell.execute_reply":"2022-08-30T11:27:49.220801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"BOLD = '\\033[1m'\n\nprint(BOLD + 'Train_df' + BOLD)\nprint(reduce_mem_usage(train_df))\nprint('\\n')\nprint(BOLD + 'Test_df' + BOLD)\nprint(reduce_mem_usage(test_df))\nprint('\\n')\nprint(BOLD + 'Other_df' + BOLD)\nprint(reduce_mem_usage(other_df))\nprint('\\n')\ntrain_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.223936Z","iopub.execute_input":"2022-08-30T11:27:49.224518Z","iopub.status.idle":"2022-08-30T11:27:49.263534Z","shell.execute_reply.started":"2022-08-30T11:27:49.224481Z","shell.execute_reply":"2022-08-30T11:27:49.262147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_df.shape)\nprint(test_df.shape)\nprint(other_df.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.265882Z","iopub.execute_input":"2022-08-30T11:27:49.26637Z","iopub.status.idle":"2022-08-30T11:27:49.274242Z","shell.execute_reply.started":"2022-08-30T11:27:49.266322Z","shell.execute_reply":"2022-08-30T11:27:49.272839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['image_path'] = train_df['image_id'].apply(lambda val: \"../input/mayo-clinic-strip-ai/train/\" + val + \".tif\")\ntrain_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.276011Z","iopub.execute_input":"2022-08-30T11:27:49.276411Z","iopub.status.idle":"2022-08-30T11:27:49.296585Z","shell.execute_reply.started":"2022-08-30T11:27:49.276364Z","shell.execute_reply":"2022-08-30T11:27:49.29531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.columns","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.302501Z","iopub.execute_input":"2022-08-30T11:27:49.302902Z","iopub.status.idle":"2022-08-30T11:27:49.312966Z","shell.execute_reply.started":"2022-08-30T11:27:49.30287Z","shell.execute_reply":"2022-08-30T11:27:49.311228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = train_df[['image_id', 'image_path', 'image_num', 'center_id', 'patient_id', 'label']]\ntrain_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.31448Z","iopub.execute_input":"2022-08-30T11:27:49.315328Z","iopub.status.idle":"2022-08-30T11:27:49.33287Z","shell.execute_reply.started":"2022-08-30T11:27:49.315292Z","shell.execute_reply":"2022-08-30T11:27:49.331656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.isnull().sum() ","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.334722Z","iopub.execute_input":"2022-08-30T11:27:49.335665Z","iopub.status.idle":"2022-08-30T11:27:49.349902Z","shell.execute_reply.started":"2022-08-30T11:27:49.335537Z","shell.execute_reply":"2022-08-30T11:27:49.349008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['image_id'].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.351064Z","iopub.execute_input":"2022-08-30T11:27:49.351409Z","iopub.status.idle":"2022-08-30T11:27:49.364991Z","shell.execute_reply.started":"2022-08-30T11:27:49.351377Z","shell.execute_reply":"2022-08-30T11:27:49.363636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Unique Values for: ')\ntrain_df.nunique()\n\nprint('\\n')\n\nnu = train_df.nunique().reset_index()\nnu.columns = ['feature','nunique']\nplt.figure(figsize=(12,4))\nax = sns.barplot(x='feature', y='nunique', data=nu)\nax.bar_label(ax.containers[0])","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.366454Z","iopub.execute_input":"2022-08-30T11:27:49.366903Z","iopub.status.idle":"2022-08-30T11:27:49.639299Z","shell.execute_reply.started":"2022-08-30T11:27:49.366868Z","shell.execute_reply":"2022-08-30T11:27:49.637976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info() ","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.640999Z","iopub.execute_input":"2022-08-30T11:27:49.641457Z","iopub.status.idle":"2022-08-30T11:27:49.656672Z","shell.execute_reply.started":"2022-08-30T11:27:49.641413Z","shell.execute_reply":"2022-08-30T11:27:49.655597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['center_id'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.658466Z","iopub.execute_input":"2022-08-30T11:27:49.658926Z","iopub.status.idle":"2022-08-30T11:27:49.671353Z","shell.execute_reply.started":"2022-08-30T11:27:49.658892Z","shell.execute_reply":"2022-08-30T11:27:49.670288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12,6))\nsns.countplot(data=train_df, x='center_id', palette='flare', order=train_df['center_id'].value_counts(ascending=True).index)","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.672765Z","iopub.execute_input":"2022-08-30T11:27:49.67427Z","iopub.status.idle":"2022-08-30T11:27:49.928006Z","shell.execute_reply.started":"2022-08-30T11:27:49.67422Z","shell.execute_reply":"2022-08-30T11:27:49.926778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Take note of the min and max values for number of slides per patient","metadata":{}},{"cell_type":"code","source":"train_df['patient_id'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.930634Z","iopub.execute_input":"2022-08-30T11:27:49.931029Z","iopub.status.idle":"2022-08-30T11:27:49.942039Z","shell.execute_reply.started":"2022-08-30T11:27:49.930995Z","shell.execute_reply":"2022-08-30T11:27:49.940415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['patient_id'].value_counts().describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.943718Z","iopub.execute_input":"2022-08-30T11:27:49.944207Z","iopub.status.idle":"2022-08-30T11:27:49.958459Z","shell.execute_reply.started":"2022-08-30T11:27:49.944163Z","shell.execute_reply":"2022-08-30T11:27:49.957499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[train_df['patient_id']=='91b9d3']","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.960048Z","iopub.execute_input":"2022-08-30T11:27:49.960796Z","iopub.status.idle":"2022-08-30T11:27:49.975747Z","shell.execute_reply.started":"2022-08-30T11:27:49.960751Z","shell.execute_reply":"2022-08-30T11:27:49.974339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12,6))\nsns.countplot(data=train_df, x='image_num', order=train_df['image_num'].value_counts().index, palette='mako')","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:49.977548Z","iopub.execute_input":"2022-08-30T11:27:49.978183Z","iopub.status.idle":"2022-08-30T11:27:50.131527Z","shell.execute_reply.started":"2022-08-30T11:27:49.978142Z","shell.execute_reply":"2022-08-30T11:27:50.130626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Most patients have only 1 image ","metadata":{}},{"cell_type":"markdown","source":"### Check whether dataset is balanced","metadata":{}},{"cell_type":"code","source":"train_df['label'].value_counts()\nprint('\\n')\nsns.countplot(data=train_df, x='label')","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:50.132906Z","iopub.execute_input":"2022-08-30T11:27:50.133942Z","iopub.status.idle":"2022-08-30T11:27:50.24646Z","shell.execute_reply.started":"2022-08-30T11:27:50.133892Z","shell.execute_reply":"2022-08-30T11:27:50.2455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Unbalanced Dataset, CE makes up ~73% of the entire dataset <br>\nWe can use data augmentation on the image dataset for deep learning.","metadata":{}},{"cell_type":"markdown","source":"### Reading a sample TIF image <br> \ncredits https://www.kaggle.com/code/junjitakeshima/mayo-simple-cnn-starter-eng","metadata":{}},{"cell_type":"markdown","source":"As we can see, there are many blank areas that could affect the performance of our model. <br> \nHence, we can perform a technique called seam carving which removes unwanted areas and retains important information.","metadata":{}},{"cell_type":"markdown","source":"### Let's examine a few slides","metadata":{}},{"cell_type":"code","source":"from PIL import Image","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:50.247793Z","iopub.execute_input":"2022-08-30T11:27:50.248839Z","iopub.status.idle":"2022-08-30T11:27:50.254492Z","shell.execute_reply.started":"2022-08-30T11:27:50.248793Z","shell.execute_reply":"2022-08-30T11:27:50.253182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nImage.MAX_IMAGE_PIXELS = None\nsize = (512,512)\n\n\nsample_slides = train_df.sample(n=3, random_state=1) \nfig, axs = plt.subplots(1, len(sample_slides.index), figsize = (12, 12))\n\nfor i in range(len(sample_slides.index)): \n    \n    slide = Image.open(sample_slides.iloc[i]['image_path'])\n    slide.thumbnail(size)\n    axs[i].set_title(f'Image Size: {slide.size}')\n    axs[i].imshow(slide)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-30T11:27:50.256537Z","iopub.execute_input":"2022-08-30T11:27:50.257245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Seam Carving\nNote all the empty spaces. We should consider removing the background to speed up the classification process. <br>\nWhat is Seam Carving? Simply put, seam carving is an algorithm that allows for images to be resized without losing important features. It makes use of a seam, which is a line of 8-connected pixels that run either vertically or horizontally. The importance of a seam (a line of pixels) is calculated via its neighbouring pixels found either on the sides of the edges of a pixel. If the energy level (gradient function that measures importance) is low, the seam is deemed as less important. Less important pixels are usually make up the background of an image. <br>\n<br>\ncredits: https://www.kaggle.com/code/ren4yu/mayo-clinic-removing-background-via-seam-carving ","metadata":{}},{"cell_type":"code","source":"! pip install seam-carving","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seam_carving","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfig, axs = plt.subplots(1, len(sample_slides.index), figsize = (12, 12))\n\nfor i in range(len(sample_slides.index)): \n    \n    slide = Image.open(sample_slides.iloc[i]['image_path'])\n    size = (512,512)\n    slide = slide.resize(size)\n    slide = seam_carving.resize(slide, (128,128))\n    axs[i].set_title(f'Image Size: {slide.size}')\n    axs[i].imshow(slide)\n\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Check the shape of 1 slide, esp. the color channel on the 3rd index","metadata":{}},{"cell_type":"code","source":"slide1 = Image.open(sample_slides.iloc[1]['image_path'])\nslide1.thumbnail(size) \nslide1 = np.array(slide1) ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"slide1.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}