{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# The tile_resize_stitch function","metadata":{}},{"cell_type":"markdown","source":"The **Mayo Clinic STRIP AI challenge** presented a difficult situation. Some of the images were simply too large to be read into memory. Many others were too large to handle effectively once loaded.\n\nThe largest one is truly impressive with 4.89 billion pixels, not to mention 4 channels.\n\nHere are my thoughts and solution.\n\nIt may not be the most elegant implementation, but it lets me resize the big images before I can start tiling them.\n\nThis is important because I want the tiles to be uniform, and for that to happen, the input images that go into the tiles must have dimensions that are the same order of magnitude. Otherwise, the area captured by a 100x100 tile from a billion-pixel image would contain far less information than a 100x100 tile from a million-pixel image would contain.\n\nThe plan is to use openslide to read in portions of the images. These portions are then individually resized and finally stitched together.\n\nThe function contains a lot of 'unnecessary' instructions, but that reflects the thinking process as the code developed. By setting the 'show' and 'debug' flags to False, all commentary can be suppressed and time taken to execute reduced considerably. By all means, feel free to experiment with the code and share your thoughts.\n\nIf you set show=True, you can see how the tiles are put together step-by-step.\n**The idea of this notebook is primarily to walk you through the steps of piecemeal loading a huge image, resizing those bits, and then assembling the image back together.**\n\nI have kept the default value of tiles_per_side as 5. This means a large image will be broken down into 25 parts, reducing the in-memory requirements to only 4% (approx) of the value otherwise.\n\n\n**UPDATE:** Do take a look at this companion notebook:\nhttps://www.kaggle.com/code/saurabhsawhney/good-quality-tiles-that-respect-the-cpu-limits","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\n# import skimage.io as io\nfrom openslide import OpenSlide\nimport cv2\nimport matplotlib.pyplot as plt\nfrom tqdm.notebook import tqdm\nfrom PIL import Image\nImage.MAX_IMAGE_PIXELS = None\nprint('libraries imported')","metadata":{"execution":{"iopub.status.busy":"2022-09-08T02:33:18.325254Z","iopub.execute_input":"2022-09-08T02:33:18.325729Z","iopub.status.idle":"2022-09-08T02:33:18.334031Z","shell.execute_reply.started":"2022-09-08T02:33:18.325691Z","shell.execute_reply":"2022-09-08T02:33:18.332532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reading in the train and test data\ntrain_csv = pd.read_csv('../input/mayo-clinic-strip-ai/train.csv')\ncols = train_csv.columns\n# Getting image paths into the df\ntrain_csv['image_path'] = train_csv.image_id.apply(lambda x: os.path.join(\"../input/mayo-clinic-strip-ai/train/\", x+\".tif\"))\ndef enhance_df(df):\n    df[\"image_size\"]   = df.image_path.apply(lambda x: Image.open(x).size)\n    df[\"image_pixels\"] = df[\"image_size\"].apply(lambda x: int(x[0]*int(x[1])))    \n    df[\"image_width\"]  = df[\"image_size\"].apply(lambda x: int(x[0]))\n    df[\"image_height\"] = df[\"image_size\"].apply(lambda x: int(x[1]))\n    df[\"aspect_ratio\"] = df[\"image_width\"]/df[\"image_height\"]\n    return df\ntrain_csv=enhance_df(train_csv)\nprint('enhanced train dataframe ready')\n\n# The largest and smallest images\ndisplay(train_csv[train_csv.image_pixels==train_csv.image_pixels.max()])\ndisplay(train_csv[train_csv.image_pixels==train_csv.image_pixels.min()])","metadata":{"execution":{"iopub.status.busy":"2022-09-08T02:33:18.3367Z","iopub.execute_input":"2022-09-08T02:33:18.338597Z","iopub.status.idle":"2022-09-08T02:33:19.899179Z","shell.execute_reply.started":"2022-09-08T02:33:18.338519Z","shell.execute_reply":"2022-09-08T02:33:19.897909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# The tile_resize_stitch function","metadata":{}},{"cell_type":"code","source":"%%time\ndef tile_resize_stich(row, horizontal_size=4000, cutoff_size=3500000000, tiles_per_side=5, show=False, debug=False):\n    '''\n    Break up the large image, resize individual tiles, put them back together\n    Keep horizontal size at the same size specified for general resizing.\n    '''\n    def demarc(): print('='*100)\n    \n    # get the metadata -----------------------------------------------------------        \n    image_path=row.image_path\n    image_width=row.image_width\n    image_height=row.image_height\n    \n    # Show the original image, if req'd ------------------------------------------\n    if show and row.image_pixels<cutoff_size:\n        orig_image = np.array(cv2.imread(row.image_path))\n        orig_image = cv2.cvtColor(orig_image, cv2.COLOR_RGB2BGR)        \n        print('Original Image')\n        plt.imshow(orig_image); plt.show()\n        demarc()\n        \n    # What should be the tile size? ----------------------------------------------    \n    # Before  resizing -----------\n    tile_size=(int(image_width/tiles_per_side),int(image_height/tiles_per_side))\n    # After resizing -------------\n    h_size=int(horizontal_size/tiles_per_side) # target horizontal size of each tile after resizing\n    v_size=int(h_size/row.aspect_ratio)        # target vertical size of each tile after resizing, to maintain aspect ratio\n\n    # Let's make tiles from this -------------------------------------------------\n    slide=OpenSlide(image_path) # using OpenSlide to get access to the image    \n    if debug:\n        print(f'Original image_width {image_width}   image_height {image_height}')\n        print(f'individual tile_size before resizing: {tile_size}')\n    tiles=[]\n    big_tile_number=1\n    for v in tqdm(range(0,image_height-tile_size[1]+1,tile_size[1])): # The +1 is just to manage the last step in the range\n        for h in range(0,image_width-tile_size[0]+1,tile_size[0]):\n            if debug: print('processing big tile', big_tile_number)\n            image = slide.read_region((h,v),0, tile_size)  # reading a tile_size area of the image, starting at h,v,position.\n            image = np.array(image)\n            image = cv2.resize(image, dsize=(h_size,v_size), interpolation=cv2.INTER_NEAREST) # INTER_CUBIC takes far longer\n            tiles.append(image)\n            big_tile_number+=1\n            if debug: print('Tile shape:', image.shape)\n            \n    # Showing off the tiles in a grid structure ----------------------------------        \n    if show:\n        print('showing tiles')\n        fig, ax = plt.subplots(nrows = tiles_per_side,ncols = tiles_per_side, figsize = (6,6/row.aspect_ratio))\n        for i,t in enumerate(tiles):\n            x_grid=int(i/tiles_per_side)            \n            y_grid=i%tiles_per_side\n            ax[x_grid,y_grid].imshow(t)\n            ax[x_grid,y_grid].axis('off')\n        fig.tight_layout()\n        plt.show()\n        \n    # Stitching it all up --------------------------------------------------------\n    stitched = np.array(Image.new('RGBA', (h_size*tiles_per_side, v_size*tiles_per_side)))\n    if debug:\n        print('Beginning the stitching process...')\n        print('First, a placeholder image of shape', stitched.shape)\n    for pos, individual_tile in enumerate(tiles):\n            x_grid=pos%tiles_per_side\n            y_grid=int(pos/tiles_per_side)\n            if debug: print(f'GRID POSITIONS: {x_grid}, {y_grid}')\n            stitched[\n                     y_grid*v_size:y_grid*v_size+v_size,\n                     x_grid*h_size:x_grid*h_size+h_size,                \n                    :] = individual_tile\n    # Showing off the stitching process -----------------------------------------                \n            if show: \n                plt.imshow(stitched)\n                plt.show()\n                demarc()\n    # All done ------------------------------------------------------------------              \n    return stitched\n\n################################################\n# The largest image is 329 - tackling the beast now\nrow=train_csv.loc[329]\nprint('='*100)\nstitched = tile_resize_stich(row, tiles_per_side=5, show=True, debug=True)\nprint('='*100)\nplt.imshow(stitched)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-08T02:33:19.901171Z","iopub.execute_input":"2022-09-08T02:33:19.901946Z","iopub.status.idle":"2022-09-08T02:43:57.108759Z","shell.execute_reply.started":"2022-09-08T02:33:19.901906Z","shell.execute_reply":"2022-09-08T02:43:57.106203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Results of some experiments changing the number of tiles per side\n\n\n1. Largest image wall times with 5 tiles per side\n   - With show and debug: 10min 37s\n<br><br><br>\n1. Largest image wall times with 10 tiles per side\n   - Without show and debug: 7min 29s\n   - With show and debug:    11min 17s\n<br><br><br>\n1. Largest image wall times with 200 tiles per side\n   - Without show and debug: 5min 44s\n\n\n**A word of caution:** At some point of time, increasing the tiles per side will break the code. This will happen sooner with smaller images.\n\n\n**UPDATE:** Do take a look at this companion notebook: \nhttps://www.kaggle.com/code/saurabhsawhney/good-quality-tiles-that-respect-the-cpu-limits\n\n**PLEASE UPVOTE IF THIS HELPED YOU OR MAYBE IF YOU JUST LIKED THIS.**\n\n**THANKS. :)**","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}