{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <B> Mayo Clinic - STRIP AI</B>\n\n## <u>EDA - Image Processing</u>\n\n**Organiser:** The Mayo Clinic is a nonprofit American academic medical center focused on integrated health care, education, and research. It employs over 4,500 physicians and scientists, along with another 58,400 administrative and allied health staff, across three major campuses:\n* Rochester, Minnesota\n* Jacksonville, Florida\n* Phoenix/Scottsdale, Arizona.\nThe practise specializes in treating difficult cases through tertiary care and destination medicine. It spends over 660 million dollars a year on research and has more than 3,000 full-time research personnel\nCompetition's sponsor webpage: https://www.mayoclinic.org/\n\n**Competition overview:**\n* Type of problem: Classification\n* Target: binary\n* Subject: blood clots \n\nWhat are the blook clots you can read [here](https://www.hematology.org/education/patients/blood-clots).\n\nIn this notebook I'll show you how to open .tiff files in a couple of ways using following libraries:\n* Matplotlib\n* CV2\n* Scikit-image","metadata":{}},{"cell_type":"markdown","source":"[](http://)","metadata":{}},{"cell_type":"markdown","source":"## Importing libraries","metadata":{}},{"cell_type":"code","source":"import os\nimport sys\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport cv2\nfrom skimage import io,img_as_float,img_as_ubyte,color\n\nfrom glob import glob\nfrom pprint import pprint\nfrom collections import defaultdict\nimport gc\n\nfrom warnings import filterwarnings\nfilterwarnings('ignore')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-10T06:09:30.967425Z","iopub.execute_input":"2022-08-10T06:09:30.967892Z","iopub.status.idle":"2022-08-10T06:09:30.976706Z","shell.execute_reply.started":"2022-08-10T06:09:30.967859Z","shell.execute_reply":"2022-08-10T06:09:30.97489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Images metadata  \n\nIn this section we'll explore tabular data related to the provided images.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('../input/mayo-clinic-strip-ai/train.csv')\ntest_df = pd.read_csv('../input/mayo-clinic-strip-ai/test.csv')\nother_df = pd.read_csv('../input/mayo-clinic-strip-ai/other.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:53.926641Z","iopub.execute_input":"2022-08-10T05:56:53.927524Z","iopub.status.idle":"2022-08-10T05:56:53.956232Z","shell.execute_reply.started":"2022-08-10T05:56:53.927468Z","shell.execute_reply":"2022-08-10T05:56:53.955213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:54.067703Z","iopub.execute_input":"2022-08-10T05:56:54.068504Z","iopub.status.idle":"2022-08-10T05:56:54.083145Z","shell.execute_reply.started":"2022-08-10T05:56:54.068453Z","shell.execute_reply":"2022-08-10T05:56:54.08116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:54.19818Z","iopub.execute_input":"2022-08-10T05:56:54.198705Z","iopub.status.idle":"2022-08-10T05:56:54.209608Z","shell.execute_reply.started":"2022-08-10T05:56:54.198667Z","shell.execute_reply":"2022-08-10T05:56:54.207821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:54.32767Z","iopub.execute_input":"2022-08-10T05:56:54.3283Z","iopub.status.idle":"2022-08-10T05:56:54.340537Z","shell.execute_reply.started":"2022-08-10T05:56:54.328267Z","shell.execute_reply":"2022-08-10T05:56:54.339509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:54.457252Z","iopub.execute_input":"2022-08-10T05:56:54.458158Z","iopub.status.idle":"2022-08-10T05:56:54.466264Z","shell.execute_reply.started":"2022-08-10T05:56:54.458115Z","shell.execute_reply":"2022-08-10T05:56:54.464624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"other_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:54.576749Z","iopub.execute_input":"2022-08-10T05:56:54.577497Z","iopub.status.idle":"2022-08-10T05:56:54.599813Z","shell.execute_reply.started":"2022-08-10T05:56:54.577455Z","shell.execute_reply":"2022-08-10T05:56:54.598006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"other_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:54.69673Z","iopub.execute_input":"2022-08-10T05:56:54.697142Z","iopub.status.idle":"2022-08-10T05:56:54.705424Z","shell.execute_reply.started":"2022-08-10T05:56:54.697109Z","shell.execute_reply":"2022-08-10T05:56:54.703918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Get a concise summary of a DataFrame.\ntrain_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:54.816544Z","iopub.execute_input":"2022-08-10T05:56:54.81702Z","iopub.status.idle":"2022-08-10T05:56:54.835698Z","shell.execute_reply.started":"2022-08-10T05:56:54.816983Z","shell.execute_reply":"2022-08-10T05:56:54.834488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Generate descriptive statistics.\ntrain_df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:54.946947Z","iopub.execute_input":"2022-08-10T05:56:54.947649Z","iopub.status.idle":"2022-08-10T05:56:54.969893Z","shell.execute_reply.started":"2022-08-10T05:56:54.947612Z","shell.execute_reply":"2022-08-10T05:56:54.968536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Generate descriptive statistics for categorical data.\ntrain_df.describe(include='object')","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:55.068913Z","iopub.execute_input":"2022-08-10T05:56:55.069668Z","iopub.status.idle":"2022-08-10T05:56:55.090577Z","shell.execute_reply.started":"2022-08-10T05:56:55.069618Z","shell.execute_reply":"2022-08-10T05:56:55.089301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Check for null values\ntrain_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:55.197314Z","iopub.execute_input":"2022-08-10T05:56:55.198049Z","iopub.status.idle":"2022-08-10T05:56:55.2092Z","shell.execute_reply.started":"2022-08-10T05:56:55.197994Z","shell.execute_reply":"2022-08-10T05:56:55.207887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's check number of unique patients in each dataset.","metadata":{}},{"cell_type":"code","source":"patients_train = train_df['patient_id'].nunique()\npatients_test = test_df['patient_id'].nunique()\npatients_other = other_df['patient_id'].nunique()\n\nprint(f\"Number of unique patients in train set: {patients_train}\")\nprint(f\"Number of unique patients in test set: {patients_test}\")\nprint(f\"Number of unique patients in the 'other' set: {patients_other}\")","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-08-10T05:56:55.307387Z","iopub.execute_input":"2022-08-10T05:56:55.308589Z","iopub.status.idle":"2022-08-10T05:56:55.319428Z","shell.execute_reply.started":"2022-08-10T05:56:55.308518Z","shell.execute_reply":"2022-08-10T05:56:55.318107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get percentage of each unique value present in target column\nlabels=train_df[\"label\"].value_counts(normalize=True)*100\nlabels","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:55.427558Z","iopub.execute_input":"2022-08-10T05:56:55.428316Z","iopub.status.idle":"2022-08-10T05:56:55.442813Z","shell.execute_reply.started":"2022-08-10T05:56:55.428281Z","shell.execute_reply":"2022-08-10T05:56:55.441036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels.index","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:55.548264Z","iopub.execute_input":"2022-08-10T05:56:55.548723Z","iopub.status.idle":"2022-08-10T05:56:55.557267Z","shell.execute_reply.started":"2022-08-10T05:56:55.548689Z","shell.execute_reply":"2022-08-10T05:56:55.555986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels.values","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:55.667024Z","iopub.execute_input":"2022-08-10T05:56:55.6683Z","iopub.status.idle":"2022-08-10T05:56:55.676352Z","shell.execute_reply.started":"2022-08-10T05:56:55.668249Z","shell.execute_reply":"2022-08-10T05:56:55.67545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Check whether data is balanced or not\nplt.figure(figsize=(6,6))\nplt.pie(x=labels.values, labels=labels.index, autopct='%1.1f%%')\nplt.title('Target Proportion', fontsize=15)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:55.726491Z","iopub.execute_input":"2022-08-10T05:56:55.727673Z","iopub.status.idle":"2022-08-10T05:56:55.854406Z","shell.execute_reply.started":"2022-08-10T05:56:55.727616Z","shell.execute_reply":"2022-08-10T05:56:55.852789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We got unbalanced dataset.","metadata":{}},{"cell_type":"code","source":"# Get percentage of each unique value present in center_id column\ncenters=train_df[\"center_id\"].value_counts(normalize=True).sort_index(ascending=True)*100\ncenters","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:55.857544Z","iopub.execute_input":"2022-08-10T05:56:55.859522Z","iopub.status.idle":"2022-08-10T05:56:55.879371Z","shell.execute_reply.started":"2022-08-10T05:56:55.859441Z","shell.execute_reply":"2022-08-10T05:56:55.877621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,6))\n\nplt.subplot(1,2,1)\nsns.barplot(x=labels.index, y=labels.values)\nplt.title(\"Distribution of a target variable\")\n\nplt.subplot(1,2,2)\nsns.barplot(x=centers.index, y=centers.values)\nplt.title(\"Images per clinic center\")\n\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:55.917834Z","iopub.execute_input":"2022-08-10T05:56:55.919025Z","iopub.status.idle":"2022-08-10T05:56:56.37966Z","shell.execute_reply.started":"2022-08-10T05:56:55.918973Z","shell.execute_reply":"2022-08-10T05:56:56.378042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Other types of blood clots not a part of this competition:\")\nprint(list(other_df['other_specified'].unique()))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T05:56:56.382464Z","iopub.execute_input":"2022-08-10T05:56:56.383387Z","iopub.status.idle":"2022-08-10T05:56:56.392418Z","shell.execute_reply.started":"2022-08-10T05:56:56.383337Z","shell.execute_reply":"2022-08-10T05:56:56.390865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Images - Exploration and Processing  \n\nFirst let's see how many images do we have to deal with.","metadata":{}},{"cell_type":"code","source":"train_images = glob(\"/kaggle/input/mayo-clinic-strip-ai/train/*\")\ntest_images = glob(\"/kaggle/input/mayo-clinic-strip-ai/test/*\")\nother_images = glob(\"/kaggle/input/mayo-clinic-strip-ai/other/*\")\nprint(f\"Number of images in a train set: {len(train_images)}\")\nprint(f\"Number of images in a test set: {len(test_images)}\")\nprint(f\"Number of other images: {len(other_images)}\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T05:56:56.395838Z","iopub.execute_input":"2022-08-10T05:56:56.396527Z","iopub.status.idle":"2022-08-10T05:56:56.468377Z","shell.execute_reply.started":"2022-08-10T05:56:56.396455Z","shell.execute_reply":"2022-08-10T05:56:56.467489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.1 Getting images' statistics using OpenSlide package\n\nOpenSlide is a library created for reading high-resolution pictures used in digital pathology. It has a couple of practical features used when dealing with these specific heavy images.  \nDocs: https://openslide.org/api/python/","metadata":{}},{"cell_type":"code","source":"from openslide import OpenSlide","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T05:56:56.470673Z","iopub.execute_input":"2022-08-10T05:56:56.471332Z","iopub.status.idle":"2022-08-10T05:56:56.476469Z","shell.execute_reply.started":"2022-08-10T05:56:56.471294Z","shell.execute_reply":"2022-08-10T05:56:56.474817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_prop = defaultdict(list)\n\nfor i, path in enumerate(train_images):\n    img_path = train_images[i]\n    slide = OpenSlide(img_path)    \n    img_prop['image_id'].append(img_path[-12:-4])\n    img_prop['width'].append(slide.dimensions[0])\n    img_prop['height'].append(slide.dimensions[1])\n    img_prop['size'].append(round(os.path.getsize(img_path) / 1e6, 2))\n    img_prop['path'].append(img_path)\n\n# Creating a dataframe of image details\nimage_data = pd.DataFrame(img_prop)\nimage_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:56:56.477986Z","iopub.execute_input":"2022-08-10T05:56:56.478865Z","iopub.status.idle":"2022-08-10T05:57:05.97398Z","shell.execute_reply.started":"2022-08-10T05:56:56.478828Z","shell.execute_reply":"2022-08-10T05:57:05.972387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add new aspect_ration column to datafarme\nimage_data['img_aspect_ratio'] = image_data['width']/image_data['height']","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:57:05.976781Z","iopub.execute_input":"2022-08-10T05:57:05.977332Z","iopub.status.idle":"2022-08-10T05:57:05.999685Z","shell.execute_reply.started":"2022-08-10T05:57:05.977283Z","shell.execute_reply":"2022-08-10T05:57:05.998341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Sort dataframe & reset index\nimage_data.sort_values(by='image_id', inplace=True)\nimage_data.reset_index(inplace=True, drop=True)\nimage_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:58:57.877973Z","iopub.execute_input":"2022-08-10T05:58:57.87935Z","iopub.status.idle":"2022-08-10T05:58:57.898285Z","shell.execute_reply.started":"2022-08-10T05:58:57.879259Z","shell.execute_reply":"2022-08-10T05:58:57.896424Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:59:09.516052Z","iopub.execute_input":"2022-08-10T05:59:09.517444Z","iopub.status.idle":"2022-08-10T05:59:09.532758Z","shell.execute_reply.started":"2022-08-10T05:59:09.517366Z","shell.execute_reply":"2022-08-10T05:59:09.530926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merge with train_df based on common column \"image_id\"\nimage_data = pd.merge(image_data,train_df, on='image_id')\nimage_data","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:01:11.587558Z","iopub.execute_input":"2022-08-10T06:01:11.588029Z","iopub.status.idle":"2022-08-10T06:01:11.620696Z","shell.execute_reply.started":"2022-08-10T06:01:11.587993Z","shell.execute_reply":"2022-08-10T06:01:11.618504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_data.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:04:36.828625Z","iopub.execute_input":"2022-08-10T06:04:36.829086Z","iopub.status.idle":"2022-08-10T06:04:36.868297Z","shell.execute_reply.started":"2022-08-10T06:04:36.829051Z","shell.execute_reply":"2022-08-10T06:04:36.866855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_num=image_data.select_dtypes([\"int64\",\"float64\"])\nplt.figure(figsize=(15,8))\n\nfor feature,i in zip(df_num.columns,range(1,7)):\n    plt.subplot(2,3,i)\n    sns.distplot(df_num[feature],hist=True,kde=True)\n    plt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T06:11:59.995273Z","iopub.execute_input":"2022-08-10T06:11:59.995717Z","iopub.status.idle":"2022-08-10T06:12:02.446079Z","shell.execute_reply.started":"2022-08-10T06:11:59.995682Z","shell.execute_reply":"2022-08-10T06:12:02.444945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.2 Displaying images\n","metadata":{}},{"cell_type":"code","source":"CE_imgs = image_data.loc[image_data['label']=='CE','path']\nLAA_imgs = image_data.loc[image_data['label']=='LAA','path']","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:09:43.258229Z","iopub.execute_input":"2022-08-10T05:09:43.258631Z","iopub.status.idle":"2022-08-10T05:09:43.267776Z","shell.execute_reply.started":"2022-08-10T05:09:43.258595Z","shell.execute_reply":"2022-08-10T05:09:43.266338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**First impressions:**\n1. Images sizes are from small ones to a high-resolution ones\n2. Images have different aspect ratios\n3. A significant amount of images is a background\n4. Backgrounds have different colours\n5. Cloths are usually in the form of multiple small pieces\n6. Blood cloths have different colours","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:43:16.533119Z","iopub.execute_input":"2022-07-09T10:43:16.533551Z","iopub.status.idle":"2022-07-09T10:43:16.721438Z","shell.execute_reply.started":"2022-07-09T10:43:16.533516Z","shell.execute_reply":"2022-07-09T10:43:16.72019Z"}}},{"cell_type":"markdown","source":"## 2.2.1 Displaying, resizing and manipulation using CV2 and skimage\n\n- Here we're going to open the full image using *Computer Vision* library CV2. We can use this library also for image manipulations. You can also detect edges by using the Canny edge detector. \n\n\nNOTE: CV2 library has an image size limit and will give you an error with some of the images from our dataset. There's a workaround by changing the environmental variable (see section 0) but even then some images will not fit into memory. That's why we are going also to use an other library later.","metadata":{}},{"cell_type":"code","source":"img_path=train_images[261]\nimg_path","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:31:11.208235Z","iopub.execute_input":"2022-08-10T05:31:11.209053Z","iopub.status.idle":"2022-08-10T05:31:11.219745Z","shell.execute_reply.started":"2022-08-10T05:31:11.208963Z","shell.execute_reply":"2022-08-10T05:31:11.218024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_name=img_path[-12:-4]\nimg_name","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:31:40.107799Z","iopub.execute_input":"2022-08-10T05:31:40.108297Z","iopub.status.idle":"2022-08-10T05:31:40.117844Z","shell.execute_reply.started":"2022-08-10T05:31:40.108262Z","shell.execute_reply":"2022-08-10T05:31:40.116161Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sel_img = image_data[image_data['image_id']==img_name]\nsel_img","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:31:47.457832Z","iopub.execute_input":"2022-08-10T05:31:47.458246Z","iopub.status.idle":"2022-08-10T05:31:47.478387Z","shell.execute_reply.started":"2022-08-10T05:31:47.458211Z","shell.execute_reply":"2022-08-10T05:31:47.477117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_BGR = cv2.imread(img_path, cv2.IMREAD_COLOR) ## Opencv read image as BGR not RBG\nimg_RGB = cv2.cvtColor(img_BGR, cv2.COLOR_BGR2RGB)\nresize_img = cv2.resize(img_RGB, (0,0), fx=0.05, fy=0.05) ## fx anf fy are scale factors\n\nplt.figure(figsize=(12,12))\nplt.subplot(1,2,1)\nplt.imshow(img_RGB)\nplt.title(\"Original Image\")\n\nplt.subplot(1,2,2)\nplt.imshow(resize_img)\nplt.title(\"Resized Image\")\n\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:09:43.301062Z","iopub.execute_input":"2022-08-10T05:09:43.301599Z","iopub.status.idle":"2022-08-10T05:10:19.922341Z","shell.execute_reply.started":"2022-08-10T05:09:43.301562Z","shell.execute_reply":"2022-08-10T05:10:19.919922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"Type of CV2 image: {type(img_RGB)}\")\nprint(f\"Memory occupied by CV2 image: {sys.getsizeof(img_RGB)/1e6:.2f}MB\")\nprint(f\"Memory occupied by resized CV2 image: {sys.getsizeof(resize_img)/1e6:.2f}MB\")\nprint(f\"Shape of a original image:{img_RGB.shape}\")\nprint(f\"Shape of a resized image:{resize_img.shape}\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T05:10:19.925137Z","iopub.execute_input":"2022-08-10T05:10:19.925817Z","iopub.status.idle":"2022-08-10T05:10:19.938905Z","shell.execute_reply.started":"2022-08-10T05:10:19.925765Z","shell.execute_reply":"2022-08-10T05:10:19.937087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"CV2 stores image as a numpy array. \n\nBelow are 3 exemplary transformations using CV2:\n* Increasing contrast. For this we have to use `addWeighted` function\n* Detecting edges. For this we use Canny method.\n* Reading in a grayscale. Here we use `cv2.COLOR_RGB2GRAY` ","metadata":{}},{"cell_type":"code","source":"contrast_img = cv2.addWeighted(resize_img, 2.5, np.zeros(resize_img.shape,resize_img.dtype), 0, 0)\nimg_gray = cv2.cvtColor(img_RGB, cv2.COLOR_BGR2GRAY)\n\nmin_intensity_grad, max_intensity_grad = 100, 200\nimg_edge = cv2.Canny(resize_img, min_intensity_grad, max_intensity_grad)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:10:19.94064Z","iopub.execute_input":"2022-08-10T05:10:19.941203Z","iopub.status.idle":"2022-08-10T05:10:20.06886Z","shell.execute_reply.started":"2022-08-10T05:10:19.941154Z","shell.execute_reply":"2022-08-10T05:10:20.067792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,15))\n\nplt.subplot(1,3,1)\nplt.imshow(contrast_img)\nplt.title(\"Increased Constrast Image\")\n\nplt.subplot(1,3,2)\nplt.imshow(img_gray,cmap='gray',vmin = 0, vmax = 255)\nplt.title(\"Gray Image\")\n\nplt.subplot(1,3,3)\nplt.imshow(img_edge)\nplt.title(\"Canny Egde Detection\")\n\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:10:20.070178Z","iopub.execute_input":"2022-08-10T05:10:20.071007Z","iopub.status.idle":"2022-08-10T05:10:30.504789Z","shell.execute_reply.started":"2022-08-10T05:10:20.070968Z","shell.execute_reply":"2022-08-10T05:10:30.503608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#gc.collect()\n#del contrast_img, edge_img, img_gray, image_resized","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:10:30.506326Z","iopub.execute_input":"2022-08-10T05:10:30.507713Z","iopub.status.idle":"2022-08-10T05:10:30.51334Z","shell.execute_reply.started":"2022-08-10T05:10:30.507664Z","shell.execute_reply":"2022-08-10T05:10:30.511646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.2.2 Transformations using scikit-image\n\nNow the same transformations will be performed using Scikit-Image. This is an another popular python package containing set image processing funcions. Official page: https://scikit-image.org/ but the most interesting section is this: https://scikit-image.org/docs/stable/user_guide/transforming_image_data.html. Skimage stores images as numpy arrays.\n","metadata":{}},{"cell_type":"code","source":"from skimage import io, exposure, feature\nfrom skimage import transform \nfrom skimage import color","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:10:40.76824Z","iopub.execute_input":"2022-08-10T05:10:40.768709Z","iopub.status.idle":"2022-08-10T05:10:40.776129Z","shell.execute_reply.started":"2022-08-10T05:10:40.768674Z","shell.execute_reply":"2022-08-10T05:10:40.774958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_sk = io.imread(img_path)\nresize_img_sk = transform.rescale(img_sk, 0.2, multichannel=True, anti_aliasing=True) # creating a scaled image\ncontrast_img_sk = exposure.adjust_gamma(resize_img_sk, 2)\nimg_gray = color.rgb2gray(resize_img_sk)\nimg_edge_sk = feature.canny(img_gray, sigma=2) # canny works only with grayscale images","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:26:27.102874Z","iopub.execute_input":"2022-08-10T05:26:27.103435Z","iopub.status.idle":"2022-08-10T05:27:34.73598Z","shell.execute_reply.started":"2022-08-10T05:26:27.103379Z","shell.execute_reply":"2022-08-10T05:27:34.734084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,15))\nplt.subplot(1,4,1)\nplt.imshow(resize_img_sk)\nplt.title(\"Resized Image\")\n\nplt.subplot(1,4,2)\nplt.imshow(contrast_img_sk)\nplt.title(\"Increased Contrast Image\")\n\nplt.subplot(1,4,3)\nplt.imshow(img_gray,cmap='gray')\nplt.title(\"Grayscale Image\")\n\nplt.subplot(1,4,4)\nplt.imshow(img_edge_sk)\nplt.title(\"Canny Edge Detection\")\n\nplt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T05:27:34.739964Z","iopub.execute_input":"2022-08-10T05:27:34.740998Z","iopub.status.idle":"2022-08-10T05:27:42.348899Z","shell.execute_reply.started":"2022-08-10T05:27:34.740941Z","shell.execute_reply":"2022-08-10T05:27:42.347436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Analysis of color channels\nEach image has 3 color channels (RGB): Red, Green, Blue. Below I'll display these channels. As each channel represents only intensity of a given single color I'll display these images in a grayscale to avoid any confusion (by displayed colors otherwise).\n\nNOTE: Values of each colour channel are in range from 0 to 255.","metadata":{}}]}