{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Good Quality Tiles that respect the CPU limits","metadata":{}},{"cell_type":"markdown","source":"<hr>\n\n**TL/DR;** We will resize the images to a standard horizontal size, keeping the aspect ratio intact so as not to distort, and then chop up this image into multiple tiles. The tiles will be scored and sequenced, and bad tiles will be discarded and good tiles will be retained, and all be well with the world.\n<hr>\n\nOnce images are loaded into the working space, the next step is tiling them. This is one approach to care of the large patches of blank canvas that will otherwise most certainly impact the efficiency and efficacy of the model.\n\nSo the task at hand is:\n\n1. Load the image but resize it to a \"convenient\" size.\n2. Send it off to a function that'll dutifully chop it up and return a packet of tiles, nicely wrapped up.\n3. Send off this packet of tiles to another function that will help you choose the \"best\" tiles.\n\nBefore we dive in, let's clarify what is a **\"convenient\"** size and what are the **\"best tiles\"**<br><hr><br>\n\n**`CONVENIENT SIZE`**\n\nThe images are of wildly different sizes, as anyone who has more than a passing interest in the Mayo Clinic STRIP AI competition is well aware by now. So to resize them, one has to think, resize them to what?\n\nOne way to look at it is the fact that these are slides, viewed by pathologists under a microscope, so in real life, they would probably be of a similar size, at least in terms of orders of magnitude. The resolutions are, of course, all over the place, but I don't really think one tile would be 10 cm wide and another would be 2 meters across - you get the drift.\n\nSo if I can fix up a horizontal number of pixels, say 1000, and resize to that, keeping the aspect ratio in mind, I would end up with.<br>\n    **a** All slides that are exactly 1000 pixels across.<br>\n    **b** No slide is forcefully distorted to a square shape since we will maintain the aspect ratio.<br>\n    **c** The images will obviously have different heights. \n    \n*Point to consider: A little over a hundred images are oriented in the landscape format, which kind of messes up the philosophy a little. However, that is a battle for another day.*\n    \nThe actual horizontal size we are going to pick will be a multiple of 224 pixels, so that our final tiles can be of that size. This lets us get away with just one resizing, at the time of loading the image. That is, assuming you want to feed 224x224 sized tiles to the model.\n<br><hr><br>\n**`BEST TILES`**\n\nWhat is a good tile?\n\nWe need tiles with a fair bit going on. No point having a sea of pink or a cloudless blue sky, right?\n\nSince the background colour is not consistent across images, it is not possible to say, \"get rid of the pixels where the RGB is 210,200,180\" or something along those lines.\n\nBut what we do know is, similar colours have similar RGB values.\n\nWhich means we will be able to use good old coefficient of variation to figure out which are the happening hotspots.\n\nThe rationale is, if a tile has a high coefficient of variation, there must be a diversity of pixels. On the other hand, a tile with low CV is probably just as bland as... Well, let's refrain from useless comparisons.\n\nSo we shall score each tile, then sort them in order of the said score, and if a certain threshold is met, this tile can be considered good enough to be retained. Having done all this, we shall end up with **\"good quality tiles\"**. Fingers crossed.\n\n<hr>\n\n**`Also, may I please solicit upvotes if you like this notebook. Thank you.`**\n<hr>","metadata":{}},{"cell_type":"markdown","source":"# Libraries","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\nimport skimage.io as io\nimport cv2\nimport matplotlib.pyplot as plt\nfrom PIL import Image\nImage.MAX_IMAGE_PIXELS = None # just so that I can read the image size of those huge images\nprint('libraries imported')","metadata":{"execution":{"iopub.status.busy":"2022-09-23T04:57:17.528232Z","iopub.execute_input":"2022-09-23T04:57:17.528765Z","iopub.status.idle":"2022-09-23T04:57:18.497644Z","shell.execute_reply.started":"2022-09-23T04:57:17.528661Z","shell.execute_reply":"2022-09-23T04:57:18.495064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read in the data","metadata":{}},{"cell_type":"code","source":"# Reading in the train data\ntrain_csv = pd.read_csv('../input/mayo-clinic-strip-ai/train.csv')\n# Cutoff size\ncutoff_size=3500000000\nprint(f'train_data read in, BIG image cutoff size set at {cutoff_size} pixels')","metadata":{"execution":{"iopub.status.busy":"2022-09-23T04:57:18.500911Z","iopub.execute_input":"2022-09-23T04:57:18.502536Z","iopub.status.idle":"2022-09-23T04:57:18.532816Z","shell.execute_reply.started":"2022-09-23T04:57:18.502474Z","shell.execute_reply":"2022-09-23T04:57:18.531532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Get in some image metadata","metadata":{}},{"cell_type":"code","source":"def enhance_df(df):\n    df['image_path']     = df.image_id.apply(lambda x: os.path.join(\"../input/mayo-clinic-strip-ai/train/\", x+\".tif\"))    \n    df[\"image_size\"]     = df.image_path.apply(lambda x: Image.open(x).size)\n    df[\"image_pixels\"]   = df[\"image_size\"].apply(lambda x: int(x[0]*int(x[1])))    \n    df[\"image_width\"]    = df[\"image_size\"].apply(lambda x: int(x[0]))\n    df[\"image_height\"]   = df[\"image_size\"].apply(lambda x: int(x[1]))\n    df[\"aspect_ratio\"]   = df[\"image_width\"]/df[\"image_height\"]\n    df[\"too_big\"]        = df[\"image_pixels\"].apply(lambda x: 1 if x>cutoff_size else 0)\n    df[\"majority_class\"] = df[\"label\"].apply(lambda x: 1 if x==\"CE\" else 0)  # Booleany to indicate if the row belongs to the majority class \n    return df\n\n# Getting image paths and other metadata into the df\ntrain_csv=enhance_df(train_csv)\n\n# Outcomes\ndisplay(train_csv.head())\nprint('enhanced train dataframe ready')","metadata":{"execution":{"iopub.status.busy":"2022-09-23T04:57:18.534891Z","iopub.execute_input":"2022-09-23T04:57:18.535813Z","iopub.status.idle":"2022-09-23T04:57:38.071717Z","shell.execute_reply.started":"2022-09-23T04:57:18.535761Z","shell.execute_reply":"2022-09-23T04:57:38.070495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Adjustables","metadata":{}},{"cell_type":"markdown","source":"Please note that the parameters set here determine both the tile size, as well as the number of such tiles that will be created horizontally. This means that if an tile size is set of 224 and 4 such tiles are required horizontally, the image will also loaded and resized so that the width of the loaded image is resized to `224*4=896`<br>\nThis is exactly in line with the philosophy of equivalence of images in the real world, no matter their resolution.","metadata":{}},{"cell_type":"code","source":"# Adjustables -------------------------------------------------------------------------\nout_tile_size=224             # the size of the square tiles I want to produce\nn_tiles_horizontally=5        # The number of tiles I would like to create horizontally. \nhorizontal_size=out_tile_size*n_tiles_horizontally\n\nprint('Adjustables set')","metadata":{"execution":{"iopub.status.busy":"2022-09-23T04:57:38.074988Z","iopub.execute_input":"2022-09-23T04:57:38.075984Z","iopub.status.idle":"2022-09-23T04:57:38.084632Z","shell.execute_reply.started":"2022-09-23T04:57:38.075942Z","shell.execute_reply":"2022-09-23T04:57:38.082244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Util function to show an image and its size","metadata":{}},{"cell_type":"code","source":"def show(img):\n    '''Just a convenience'''\n    plt.imshow(img)\n    plt.show()\n    print(f'Image size')\n    print(f'     Width: {img.shape[1]} pixels')\n    print(f'    Height: {img.shape[0]} pixels')","metadata":{"execution":{"iopub.status.busy":"2022-09-23T04:57:38.08843Z","iopub.execute_input":"2022-09-23T04:57:38.088933Z","iopub.status.idle":"2022-09-23T04:57:38.133128Z","shell.execute_reply.started":"2022-09-23T04:57:38.088899Z","shell.execute_reply":"2022-09-23T04:57:38.131713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preprocessing functions - tile_resize_stitch","metadata":{}},{"cell_type":"code","source":"def tile_resize_stich():\n    '''\n    Break up the large image, resize individual tiles, put them back together.\n    The function as written here in this notebook is a dummy implementation. For the full implementation,\n    please refer to the notebook https://www.kaggle.com/code/saurabhsawhney/handling-large-images-tile-resize-and-stitch\n    '''\n    print('Dummy implementation to keep this notebook clutter free.')\n    print('For a full implementation of tile-resize-stitch,\\nplease refer to the notebook https://www.kaggle.com/code/saurabhsawhney/handling-large-images-tile-resize-and-stitch')\n    return np.zeros(shape=(horizontal_size,horizontal_size,3))","metadata":{"execution":{"iopub.status.busy":"2022-09-23T04:57:38.135488Z","iopub.execute_input":"2022-09-23T04:57:38.136329Z","iopub.status.idle":"2022-09-23T04:57:38.149547Z","shell.execute_reply.started":"2022-09-23T04:57:38.136292Z","shell.execute_reply":"2022-09-23T04:57:38.147737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preprocessing functions - load_image","metadata":{}},{"cell_type":"code","source":"def load_image(row, horizontal_size=2240, verbose=False):\n    '''\n    row = dataframe row\n    Pass the row of the enhanced dataframe. The function will pickup the image path, load, resize and return it.\n    '''\n    if verbose: print(f'Original image width: {row.image_width}, height: {row.image_height}')\n    # Get the image path -----------------------------------------------------------------\n    image_path=row.image_path\n    final_h_size=horizontal_size\n    final_v_size=int(horizontal_size/row.aspect_ratio)\n    dsize=(final_h_size, final_v_size)\n    # Handling a SMALL image -------------------------------------------------------------    \n    if not row.too_big:  # small image, we can handle it here\n        img=io.imread(row.image_path)\n        img = cv2.resize(img, dsize=dsize, interpolation=cv2.INTER_NEAREST)\n        if verbose: \n            print('Image loading and resizing completed')\n            show(img)\n        return img            \n    else:  # handling a big image by outsourcing\n        img= tile_resize_stich()\n        if verbose: \n            print('Image loading and resizing completed')\n            show(img)        \n        return img","metadata":{"execution":{"iopub.status.busy":"2022-09-23T04:57:38.151462Z","iopub.execute_input":"2022-09-23T04:57:38.152585Z","iopub.status.idle":"2022-09-23T04:57:38.165639Z","shell.execute_reply.started":"2022-09-23T04:57:38.152522Z","shell.execute_reply":"2022-09-23T04:57:38.164627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preprocessing functions - make_tiles","metadata":{}},{"cell_type":"code","source":"def make_tiles(img, tile_size=out_tile_size, show_img=True, debug=False):\n    '''\n    We are going to make tiles !!!\n    '''\n    if debug:\n        print('>'*80)\n        print(f'INSIDE THE make_tiles FUNCTION')\n        print('This is the image received.')\n        show(img)\n        print('>'*80)        \n        \n    how_many_tiles_horizontally=img.shape[1]//tile_size \n    how_many_tiles_vertically  =img.shape[0]//tile_size \n    \n    if debug:\n        print('how_many_tiles_horizontally:',how_many_tiles_horizontally)\n        print('how_many_tiles_vertically  :',how_many_tiles_vertically)\n    \n    # The tiling process\n    tiles=[]\n    counter=0\n    for gridrow in range(how_many_tiles_vertically):\n        for gridcol in range(how_many_tiles_horizontally):\n            row_start=gridrow*tile_size\n            row_end=gridrow*tile_size+tile_size\n            col_start=gridcol*tile_size\n            col_end=gridcol*tile_size+tile_size\n            counter=counter+1\n            tile=img[row_start:row_end,\n                     col_start:col_end,\n                     :]\n            tiles.append(tile)\n    if show_img:\n        nrows = how_many_tiles_vertically\n        ncols = how_many_tiles_horizontally\n        fig, ax = plt.subplots(nrows=nrows, ncols=ncols, figsize = (3*ncols,3*nrows))        \n        for i,t in enumerate(tiles):\n            row_pos=int(i/ncols)\n            col_pos=i%ncols            \n            ax[row_pos,col_pos].imshow(t)\n        fig.tight_layout()\n        plt.show()            \n    return np.array(tiles)","metadata":{"execution":{"iopub.status.busy":"2022-09-23T04:57:38.167424Z","iopub.execute_input":"2022-09-23T04:57:38.167784Z","iopub.status.idle":"2022-09-23T04:57:38.181124Z","shell.execute_reply.started":"2022-09-23T04:57:38.16775Z","shell.execute_reply":"2022-09-23T04:57:38.180221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Now, we need to look at quality issues. How does one pick good quality tiles?","metadata":{}},{"cell_type":"markdown","source":"The idea is, having generated the tiles, we only need to work woth those tiles that carry useful information. A sea of white or pink is not much use, right?\n\nSo the plan is this.\n\n1. We take up all the tiles generated.\n1. Each tile is converted to grayscale and the standard deviation and Coefficient of Variation is calculated.\n1. All CoV values are averaged to determine an image-level-CoV.\n1. The image-level Cov includes both useful pixels (with lots of variations) as well as those we want to discard. Therefore, this is a kind of middle-level value.\n1. Each tile gets a score which can simply be the ratio of tile-level-CoV to image-level-CoV.\n1. Tiles with lots of action going on would likely have CoV higher than the average, thus a score more than 1.\n1. Lacklustre tiles will probably have lower than image-level Cov, thus a score less than 1.\n1. The tiles can be sorted by score.\n1. What would be ideal is to have a threshold score which is based on the image-level CoV, against which individual tiles can be compared. If a tile cannot beat the threshold, then we don't accept it. This would happen if the entire image has little to offer in terms of variety.\n1. Finally, a user can choose to define how many tiles are required off an image. (This can potentially address class imbalance as well). So if say 25 tiles have been sent to the function, and the user needs only 10, then the top 10 scorers can be returned, as long as each of them also manages to beat the threshold.\n\nRight, enough talk. Let's write down some code.\n","metadata":{}},{"cell_type":"code","source":"def grayConversion(image):\n    '''\n    We'll use greyscale images to determine a color-independent image standard deviation, for use as a tile selection criterion.\n    '''\n    grayValue = 0.07 * image[:,:,2] + 0.72 * image[:,:,1] + 0.21 * image[:,:,0]\n    gray_img = grayValue.astype(np.uint8)\n    return gray_img\n\ndef pick_n_tiles(tiles, n=2, show_img=False, debug=False):\n    ''' Pick CoV based best n tiles from the tiles generated.\n        Also returning scores for later analysis, if required.'''\n    if debug:\n        print(f\"Received {len(tiles)} tiles for selection\")\n        \n    # Calculating CoV for all tiles sent our way -----------------------------------------        \n    coeffs_of_variation=[]    \n    for idx in range(tiles.shape[0]):\n        grey_img=grayConversion(tiles[idx])\n        sd = np.std(grey_img)\n        if debug: print(f'For this tile, calculated sd {sd}')\n        if sd==0: cov=0 # sd=0 will generate a NaN value for cov, so we best not go there\n        else:     cov = sd/np.mean(grey_img)\n        coeffs_of_variation.append(cov)\n        \n    # What if all tiles are plain white (or some other colour) sheets? -------------------\n    if max(coeffs_of_variation)==0: # If all covs are zero, that's not good\n        return [tiles[0]],[1e-8]  # =========================================== RETURN ===\n        # returning just one tile so that downstram code doesn't break.\n        # This is open to improvement, of course.\n        \n    # Finding out image level covariance in the presence of multiple zero cov values------\n    non_zero_cov_values=[x for x in coeffs_of_variation if x>0]\n    image_level_cov= np.mean(non_zero_cov_values) # not the best idea, I know...\n    reasonable_cov = 0.5\n    threshold_score = 0.9 + reasonable_cov/(image_level_cov+reasonable_cov) # as the image level sd falls, the background is more, so we need stricter criterion\n    kept_tiles = []\n    scores=[]\n    \n    # Time to choose wisely --------------------------------------------------------------\n    for tile, cov in zip(tiles, coeffs_of_variation):\n        score=cov/(image_level_cov+1e-8)\n        if debug: print(f'Tile selection process score: {score}, threshold score {threshold_score}, image_level cov {image_level_cov}')\n        tile_status=1 if score >= threshold_score else 0\n        if tile_status:\n            kept_tiles.append(tile)\n            scores.append(score)  \n            \n    # What if all tiles are bad? --------------------------------------------------------- \n    if len(kept_tiles)==0:\n        return [tiles[0]],[1e-8]  # =========================================== RETURN ===\n    \n    # Sorting -----( psst...I want Gryffindor)--------------------------------------------        \n    score_dict={k:v for v,k in enumerate(scores) }\n    sorted_scores=list(score_dict.keys())\n    sorted_scores.sort(reverse=True)\n    sorted_tiles=[]\n    for k in sorted_scores[:n]:\n        sorted_tiles.append(kept_tiles[score_dict[k]])\n        if debug: print(score_dict[k], k)\n            \n    if show_img: # show and tell ------------------------------------------------------------\n        for tile in sorted_tiles:\n            show(tile)\n            \n    return sorted_tiles,sorted_scores[:n]  # ================================= RETURN ===","metadata":{"execution":{"iopub.status.busy":"2022-09-23T04:57:38.182752Z","iopub.execute_input":"2022-09-23T04:57:38.183146Z","iopub.status.idle":"2022-09-23T04:57:38.20538Z","shell.execute_reply.started":"2022-09-23T04:57:38.183093Z","shell.execute_reply":"2022-09-23T04:57:38.204293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Testing our functions - moment of truth !!","metadata":{}},{"cell_type":"markdown","source":"## Step 1 - we load an image","metadata":{}},{"cell_type":"code","source":"index_position=109   # Choose your battle here. Try 109,110,111 or 522 for smaller images\nrow=train_csv.loc[index_position]\nimg=load_image(row, horizontal_size, verbose=True)","metadata":{"execution":{"iopub.status.busy":"2022-09-23T04:57:38.209035Z","iopub.execute_input":"2022-09-23T04:57:38.209503Z","iopub.status.idle":"2022-09-23T04:57:39.524229Z","shell.execute_reply.started":"2022-09-23T04:57:38.209461Z","shell.execute_reply":"2022-09-23T04:57:39.522839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 2 - we make tiles","metadata":{}},{"cell_type":"code","source":"tiles=make_tiles(img, tile_size=out_tile_size, show_img=True, debug=False)","metadata":{"execution":{"iopub.status.busy":"2022-09-23T04:57:39.526485Z","iopub.execute_input":"2022-09-23T04:57:39.527483Z","iopub.status.idle":"2022-09-23T04:57:41.633966Z","shell.execute_reply.started":"2022-09-23T04:57:39.527431Z","shell.execute_reply":"2022-09-23T04:57:41.632626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 3 - we pick the good ones","metadata":{}},{"cell_type":"code","source":"# Let's say I want 6 good tiles from the above display - well, only if we have 6 good ones\n\ntiles, scores=pick_n_tiles(tiles, n=6, show_img=False, debug=False)","metadata":{"execution":{"iopub.status.busy":"2022-09-23T04:57:41.636016Z","iopub.execute_input":"2022-09-23T04:57:41.636645Z","iopub.status.idle":"2022-09-23T04:57:42.796086Z","shell.execute_reply.started":"2022-09-23T04:57:41.636574Z","shell.execute_reply":"2022-09-23T04:57:42.794753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Hope you enjoyed the notebook. Please do upvote, and all the best for the competition !!**\n  \n<p><div style=\"text-align: left; font-family:mistral; font-size:2em\"><b> - Saurabh Sawhney </b></div></p><hr>","metadata":{}}]}