{"cells":[{"metadata":{},"cell_type":"markdown","source":"Pipeline Step 1 - Pre-Processing:\nStep 1 Part B - Exploring Masks to gain labels","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\nimport os\nimport matplotlib\nimport matplotlib.pyplot as plt # plotting figures\nfrom PIL import Image # open and display images\nimport cv2 #computer vision library\nfrom tqdm.notebook import tqdm # progress bar, tqdm shorthand for progress in Arabic \nimport skimage.io #image processing","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# setting the main directory and loading the train CSV file\nMAIN_DIR = '../input/prostate-cancer-grade-assessment'\ntrain = pd.read_csv(os.path.join(MAIN_DIR, 'train.csv')).set_index('image_id')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#dropping the mislabelled case we identified in EDA, this will return an error if ran more than once\nsusp = train[(train.gleason_score == '4+3') & (train.isup_grade != 3)]\ntrain = train.drop([susp.index[0]])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Check we removed our suspicious case\nprint(train.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Set our directories for our resized images and ensure we have all 10616 images and 10516 masks\nresize_dir = '../input/panda-resized-train-data-512x512/'\nimg_dir = resize_dir + 'train_images/train_images/'\nmask_dir = img_dir.replace('images', 'label_masks')\nprint(len(os.listdir(img_dir)))\nprint(len(os.listdir(mask_dir)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\n# there are 100 images with no masks we'll just create an image to display for the case where there are no masks\nno_mask_array = np.zeros((512,512,3), dtype = 'uint8')\nno_mask_array[:,:,2] = np.identity(512, dtype = 'uint8')*2\n\n# creating a batch of ids to test\nid_batch = train.index[0:10]\n\n#creating a function to take an image id and display image or mask (if exists) as required\ndef id2array(id, type):\n    if type == 'mask':\n        if os.path.isfile(os.path.join(mask_dir + id + '_mask.png')) == True:\n            array = skimage.io.imread(os.path.join(mask_dir + id + '_mask.png'))\n        else:\n            array = no_mask_array\n    else:\n        array = skimage.io.imread(os.path.join(img_dir + id + '.png'))\n    return array","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"img_array_batch = [id2array(item, 'image') for item in id_batch]\nfig, axs = plt.subplots(5, 2, figsize=(25,25))\nfor i in range(0,10):\n    axs[(i//2), (i%2)].imshow(img_array_batch[i])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#let us now test visualise one of our masks\nmask_test = id2array(train.index[8032], 'mask')\n\nplt.figure()\nplt.title(\"Mask with default cmap\")\nplt.imshow(mask_test[:,:,2], interpolation='nearest')\nplt.show()\n\nplt.figure()\nplt.title(\"Mask with custom cmap\")\n# Optional: create a custom color map\ncmap = matplotlib.colors.ListedColormap(['black', 'gray', 'green', 'yellow', 'orange', 'red'])\nplt.imshow(mask_test[:,:,2], cmap=cmap, interpolation='nearest', vmin=0, vmax=5)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The masks are labeled differently according to provider so we shall write a function that displays masks in a similar way regardless of data provider. \n\nThe label masks of Radboudumc were semi-automatically generated by several deep learning algorithms, contain noise, and can be considered as weakly-supervised labels. The label masks of Karolinska were semi-autotomatically generated based on annotations by a pathologist.\n\n\nRadboudumc: Prostate glands are individually labelled. Valid values are:\n0: background (non tissue) or unknown\n1: stroma (connective tissue, non-epithelium tissue)\n2: healthy (benign) epithelium\n3: cancerous epithelium (Gleason 3)\n4: cancerous epithelium (Gleason 4)\n5: cancerous epithelium (Gleason 5)\n\nKarolinska: Regions are labelled. Valid values:\n0: background (non tissue) or unknown\n1: benign tissue (stroma and epithelium combined)\n2: cancerous tissue (stroma and epithelium combined)\n\nWe will label apply Karolinska's method of not distinguishing between stroma and epithelium. We will use a graded color scheme to denote gleason 3,4 or 5 in radboud. While all cancerous cells from Karolinska will have a different block colour.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# we set up two colour maps as described above\ncmap_rad = matplotlib.colors.ListedColormap(['black', 'gray', 'gray', 'yellow', 'orange', 'red'])\ncmap_kar = matplotlib.colors.ListedColormap(['black', 'gray', 'purple'])\n\n\n# this function will take 5 image ids and display an image and the related mask if there is one if not display the no mask image defined above\ndef plot5(ids):\n    img_arrays = [id2array(item, 'image') for item in ids]\n    mask_arrays = [id2array(item, 'mask') for item in ids]\n    fig, axs = plt.subplots(5, 2, figsize=(15,25))\n    for i in range(0,5):\n        image_id = ids[i]\n        data_provider = train.loc[image_id, 'data_provider']\n        gleason_score = train.loc[image_id, 'gleason_score']\n        axs[i, 0].imshow(img_arrays[i])\n        mask_array = mask_arrays[i]\n        if data_provider == 'karolinska':\n            axs[i, 1].imshow(mask_array[:,:,2], cmap=cmap_kar, interpolation='nearest', vmin=0, vmax=2)\n        else:\n            axs[i, 1].imshow(mask_array[:,:,2], cmap=cmap_rad, interpolation='nearest', vmin=0, vmax=5)\n        for j in range(0,2):\n            axs[i,j].set_title(f\"ID: {image_id}\\nSource: {data_provider} Gleason: {gleason_score}\")\n    plt.show()\n    ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot5(train.index[1103:1108])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We now have a better understanding of our images and masks and can explore images and masks side by side. We also can access our images and masks as numerical numpy arrays. We are now ready to begin obtaining labeled 128x128 tiles, which we can use to train our Neural Network.","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}