{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**What are you trying to do in this notebook?**\n\nIn this competition, I'm working on to classify the blood clot origins in ischemic stroke. Using whole slide digital pathology images, I'll build a model that differentiates between the two major acute ischemic stroke (AIS) etiology subtypes: cardiac and large artery atherosclerosis.\n\nMy work will enable healthcare providers to better identify the origins of blood clots in deadly strokes, making it easier for physicians to prescribe the best post-stroke therapeutic management and reducing the likelihood of a second stroke.\n\n**File and Data Field Descriptions** :-\n\n**train/** - A folder containing images in the TIFF format to be used as training data.\n\n**test/** - A folder containing images to be used as test data. The actual test data comprises about 280 images.\n\n**other/** - A supplemental set of images with a either an unknown etiology or an etiology other than CE or LAA.\ntrain.csv Contains annotations for images in the train/ folder.\n\n**image_id** - A unique identifier for this instance having the form {patient_id}_{image_num}. Corresponds to the image {image_id}.tif.\n\n**center_id** - Identifies the medical center where the slide was obtained.\n\n**patient_id** - Identifies the patient from whom the slide was obtained.\n\n**image_num** - Enumerates images of clots obtained from the same patient.\n\n**label** - The etiology of the clot, either CE or LAA. This field is the classification target.\n\n**test.csv** - Annotations for images in the test/ folder. Has the same fields as train.csv excluding label.\n\n**other.csv** - Annotations for images in the other/ folder. Has the same fields as train.csv. The center_id is unavailable for these images however.\n\n**label** - The etiology of the clot, either Unknown or Other.\n\n**other_specified** - The specific etiology, when known, in case the etiology is labeled as Other.\n\n**sample_submission.csv** - A sample submission file in the correct format. See the Evaluation page for more details. Note in particular that you should make one prediction per patient_id, not per image_id.\n\n\n**Why are you trying it?**\n\nTo decrease the chances of subsequent strokes, With the help of Mayo Clinic Neurovascular Research Laboratory it encourages me to improve artificial intelligence-based etiology classification so that physicians can be better equipped to prescribe the correct treatment. New computational and artificial intelligence approaches could help save the lives of stroke survivors and help me better understand the world's second-leading cause of death.\n\nMy task is to classify the etiology (CE or LAA) of the slides in the test set for each patient. The slides comprising the training and test sets depict clots with an etiology (that is, origin) known to be either CE (Cardioembolic) or LAA (Large Artery Atherosclerosis). I include a set of supplemental slides with a either an unknown etiology or an etiology other than CE or LAA.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2022-09-27T05:49:36.967538Z","iopub.execute_input":"2022-09-27T05:49:36.967916Z","iopub.status.idle":"2022-09-27T05:49:37.194593Z","shell.execute_reply.started":"2022-09-27T05:49:36.967844Z","shell.execute_reply":"2022-09-27T05:49:37.193498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os","metadata":{"execution":{"iopub.status.busy":"2022-09-27T05:49:37.196394Z","iopub.execute_input":"2022-09-27T05:49:37.196689Z","iopub.status.idle":"2022-09-27T05:49:37.201792Z","shell.execute_reply.started":"2022-09-27T05:49:37.196663Z","shell.execute_reply":"2022-09-27T05:49:37.200714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from openslide import OpenSlide\nimport cv2\nimport matplotlib.pyplot as plt\nfrom tqdm.notebook import tqdm\nfrom PIL import Image\nImage.MAX_IMAGE_PIXELS = None\nprint('libraries imported')","metadata":{"execution":{"iopub.status.busy":"2022-09-27T05:50:23.851702Z","iopub.execute_input":"2022-09-27T05:50:23.852086Z","iopub.status.idle":"2022-09-27T05:50:24.286306Z","shell.execute_reply.started":"2022-09-27T05:50:23.85206Z","shell.execute_reply":"2022-09-27T05:50:24.284903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_csv = pd.read_csv('../input/mayo-clinic-strip-ai/train.csv')\ncols = train_csv.columns","metadata":{"execution":{"iopub.status.busy":"2022-09-27T05:50:59.561284Z","iopub.execute_input":"2022-09-27T05:50:59.561637Z","iopub.status.idle":"2022-09-27T05:50:59.578628Z","shell.execute_reply.started":"2022-09-27T05:50:59.561612Z","shell.execute_reply":"2022-09-27T05:50:59.577332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_csv['image_path'] = train_csv.image_id.apply(lambda x: os.path.join(\"../input/mayo-clinic-strip-ai/train/\", x+\".tif\"))\ndef enhance_df(df):\n    df[\"image_size\"]   = df.image_path.apply(lambda x: Image.open(x).size)\n    df[\"image_pixels\"] = df[\"image_size\"].apply(lambda x: int(x[0]*int(x[1])))    \n    df[\"image_width\"]  = df[\"image_size\"].apply(lambda x: int(x[0]))\n    df[\"image_height\"] = df[\"image_size\"].apply(lambda x: int(x[1]))\n    df[\"aspect_ratio\"] = df[\"image_width\"]/df[\"image_height\"]\n    return df\ntrain_csv=enhance_df(train_csv)\nprint('enhanced train dataframe ready')","metadata":{"execution":{"iopub.status.busy":"2022-09-27T05:51:16.193221Z","iopub.execute_input":"2022-09-27T05:51:16.19355Z","iopub.status.idle":"2022-09-27T05:51:34.990857Z","shell.execute_reply.started":"2022-09-27T05:51:16.193525Z","shell.execute_reply":"2022-09-27T05:51:34.989104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_csv[train_csv.image_pixels==train_csv.image_pixels.max()])\ndisplay(train_csv[train_csv.image_pixels==train_csv.image_pixels.min()])","metadata":{"execution":{"iopub.status.busy":"2022-09-27T05:51:55.507671Z","iopub.execute_input":"2022-09-27T05:51:55.508984Z","iopub.status.idle":"2022-09-27T05:51:55.544506Z","shell.execute_reply.started":"2022-09-27T05:51:55.50893Z","shell.execute_reply":"2022-09-27T05:51:55.5428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndef tile_resize_stich(row, horizontal_size=4000, cutoff_size=3500000000, tiles_per_side=5, show=False, debug=False):\n    '''\n    Break up the large image, resize individual tiles, put them back together\n    Keep horizontal size at the same size specified for general resizing.\n    '''\n    def demarc(): print('='*100)\n           \n    image_path=row.image_path\n    image_width=row.image_width\n    image_height=row.image_height\n    \n    if show and row.image_pixels<cutoff_size:\n        orig_image = np.array(cv2.imread(row.image_path))\n        orig_image = cv2.cvtColor(orig_image, cv2.COLOR_RGB2BGR)        \n        print('Original Image')\n        plt.imshow(orig_image); plt.show()\n        demarc()\n        \n    tile_size=(int(image_width/tiles_per_side),int(image_height/tiles_per_side))\n    h_size=int(horizontal_size/tiles_per_side) \n    v_size=int(h_size/row.aspect_ratio)        \n\n    slide=OpenSlide(image_path)     \n    if debug:\n        print(f'Original image_width {image_width}   image_height {image_height}')\n        print(f'individual tile_size before resizing: {tile_size}')\n    tiles=[]\n    big_tile_number=1\n    for v in tqdm(range(0,image_height-tile_size[1]+1,tile_size[1])): \n        for h in range(0,image_width-tile_size[0]+1,tile_size[0]):\n            if debug: print('processing big tile', big_tile_number)\n            image = slide.read_region((h,v),0, tile_size)  \n            image = np.array(image)\n            image = cv2.resize(image, dsize=(h_size,v_size), interpolation=cv2.INTER_NEAREST) \n            tiles.append(image)\n            big_tile_number+=1\n            if debug: print('Tile shape:', image.shape)\n                    \n    if show:\n        print('showing tiles')\n        fig, ax = plt.subplots(nrows = tiles_per_side,ncols = tiles_per_side, figsize = (6,6/row.aspect_ratio))\n        for i,t in enumerate(tiles):\n            x_grid=int(i/tiles_per_side)            \n            y_grid=i%tiles_per_side\n            ax[x_grid,y_grid].imshow(t)\n            ax[x_grid,y_grid].axis('off')\n        fig.tight_layout()\n        plt.show()\n        \n    stitched = np.array(Image.new('RGBA', (h_size*tiles_per_side, v_size*tiles_per_side)))\n    if debug:\n        print('Beginning the stitching process...')\n        print('First, a placeholder image of shape', stitched.shape)\n    for pos, individual_tile in enumerate(tiles):\n            x_grid=pos%tiles_per_side\n            y_grid=int(pos/tiles_per_side)\n            if debug: print(f'GRID POSITIONS: {x_grid}, {y_grid}')\n            stitched[\n                     y_grid*v_size:y_grid*v_size+v_size,\n                     x_grid*h_size:x_grid*h_size+h_size,                \n                    :] = individual_tile\n            \n            if show: \n                plt.imshow(stitched)\n                plt.show()\n                demarc()\n                  \n    return stitched\n\nrow=train_csv.loc[329]\nprint('='*100)\nstitched = tile_resize_stich(row, tiles_per_side=5, show=True, debug=True)\nprint('='*100)\nplt.imshow(stitched)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-27T05:55:29.337899Z","iopub.execute_input":"2022-09-27T05:55:29.338258Z","iopub.status.idle":"2022-09-27T06:07:09.393787Z","shell.execute_reply.started":"2022-09-27T05:55:29.338233Z","shell.execute_reply":"2022-09-27T06:07:09.392622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Did it work?**\n\nBy Using whole slide digital pathology images, I'll build a model that differentiates between the two major acute ischemic stroke (AIS) etiology subtypes: cardiac and large artery atherosclerosis and to classify the blood clot origins in ischemic stroke.\n\nMy work will enable healthcare providers to better identify the origins of blood clots in deadly strokes, making it easier for physicians to prescribe the best post-stroke therapeutic management and reducing the likelihood of a second stroke.\n\n**What did you not understand about it?**\n\nWell, everything provides in the competition data page. I've no problem while working on it. The dataset for this competition comprises over a thousand high-resolution whole-slide digital pathology images. Each slide depicts a blood clot from a patient that had experienced an acute ischemic stroke.\n\n**I hope you find this notebook useful , Good Luck!**","metadata":{}}]}