{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# MAYO Clinic Strip AI competition: Tile generation and inferencing notebook (under specified limits of the competition).","metadata":{}},{"cell_type":"markdown","source":"The notebook is used for making the inferencing on the tiled dataset prepared using DeepZoomGenerator, which is a part of MAYO Clinic - STRIP AI competition.\n\n\nThe preprocessed tiled training set can be obtained from following link -> [Dataset Link](https://www.kaggle.com/datasets/tr1gg3rtrash/mayo-clinic)\n\nExample Images from the dataset.\n\n<div> \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4636937%2Ff8fff1fa62dc75cdb3b38981b7a147fd%2F0b7871_0_2_33_15.png?generation=1663412986808182&alt=media\" width=\"300\" height=\"300\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4636937%2Fd6ed6a9702983da36ee9225a9c19e5fc%2F0b25f8_0_13_8_16.png?generation=1663413194154320&alt=media\" width=\"300\" height=\"300\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4636937%2F913f27973e1bd16b956a4551a2e7850d%2F0ba49d_0_3_6_17.png?generation=1663413413713409&alt=media\" width=\"300\" height=\"300\">\n</div>\n\n<br>\n\nThe normalized tiled training set can be obtained from following link -> [Dataset Link](https://www.kaggle.com/datasets/tr1gg3rtrash/mayo-clinic-strip-ai-normalized-dataset)\n\nExample Images from the dataset.\n\n<div> \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4636937%2F107c7f1149a0677f6125e66e44a77fdb%2FUntitled%20drawing.png?generation=1664430814926922&alt=media\">\n</div>\n","metadata":{}},{"cell_type":"markdown","source":"# Importing Libraries","metadata":{}},{"cell_type":"code","source":"import openslide\nimport os\nimport pandas as pd\nimport shutil\nimport matplotlib.pyplot as plt\nimport numpy as np\nfrom PIL import Image\nimport cv2\nimport warnings\nfrom PIL import Image\nimport tensorflow as tf\nfrom openslide import open_slide\nfrom openslide.deepzoom import DeepZoomGenerator\nfrom tqdm import tqdm\nimport skimage.transform as st\nImage.MAX_IMAGE_PIXELS = None\nwarnings.filterwarnings(\"ignore\")","metadata":{"papermill":{"duration":6.243453,"end_time":"2022-10-01T11:23:37.609681","exception":false,"start_time":"2022-10-01T11:23:31.366228","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T13:09:13.312492Z","iopub.execute_input":"2022-10-02T13:09:13.312956Z","iopub.status.idle":"2022-10-02T13:09:13.329864Z","shell.execute_reply.started":"2022-10-02T13:09:13.312912Z","shell.execute_reply":"2022-10-02T13:09:13.324873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Function that checks if image is blank or contains some cell contents in it or not. ","metadata":{}},{"cell_type":"code","source":"def norm_HnE(img, Io=240, alpha=1, beta=0.15):\n    HERef = np.array([[0.5626, 0.2159],\n                      [0.7201, 0.8012],\n                      [0.4062, 0.5581]])\n    maxCRef = np.array([1.9705, 1.0308])\n    h, w, c = img.shape\n    img = img.reshape((-1,3))\n    OD = -np.log10((img.astype(np.float)+1)/Io)\n    ODhat = OD[~np.any(OD < beta, axis=1)]\n    try:\n        eigvals, eigvecs = np.linalg.eigh(np.cov(ODhat.T))\n        return True\n    except:\n        return False","metadata":{"papermill":{"duration":0.022885,"end_time":"2022-10-01T11:23:37.635459","exception":false,"start_time":"2022-10-01T11:23:37.612574","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T13:09:13.331846Z","iopub.execute_input":"2022-10-02T13:09:13.332456Z","iopub.status.idle":"2022-10-02T13:09:13.368486Z","shell.execute_reply.started":"2022-10-02T13:09:13.332421Z","shell.execute_reply":"2022-10-02T13:09:13.36753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Defining model\n\nIn this case, do replace this with your own model","metadata":{}},{"cell_type":"code","source":"model = tf.keras.models.load_model(\"../input/first-model/content/firstExperiment\")","metadata":{"papermill":{"duration":16.632331,"end_time":"2022-10-01T11:23:54.269739","exception":false,"start_time":"2022-10-01T11:23:37.637408","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T13:09:13.369972Z","iopub.execute_input":"2022-10-02T13:09:13.370578Z","iopub.status.idle":"2022-10-02T13:09:29.96362Z","shell.execute_reply.started":"2022-10-02T13:09:13.370539Z","shell.execute_reply":"2022-10-02T13:09:29.962418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"IMG_SIZE = 224","metadata":{"papermill":{"duration":0.026545,"end_time":"2022-10-01T11:23:54.299927","exception":false,"start_time":"2022-10-01T11:23:54.273382","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T13:09:29.966168Z","iopub.execute_input":"2022-10-02T13:09:29.966521Z","iopub.status.idle":"2022-10-02T13:09:29.971623Z","shell.execute_reply.started":"2022-10-02T13:09:29.966488Z","shell.execute_reply":"2022-10-02T13:09:29.970564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Main Inferencing Cell\n\nThe cell uses the test set from the competition data and uses DeepZoomGenerator to create tiles from same. It checks if the tiles is blank or not. If the tiles is blank then the norm_HE function will return false as eigen vectors can't be obtained in that case. Hence, that tile will be discarded. Area of the tile is also checked as a threshold to make sure that areas that have high content, which will be actually useful for predicting are only considered for the tasl. In the end, mean from each tile is calculated in order to obtain prediction for a single image. All notebook compute requirements are taken care of in the notebook. ","metadata":{}},{"cell_type":"code","source":"test_dir = \"../input/mayo-clinic-strip-ai/test\"\ntest_df = pd.read_csv(\"../input/mayo-clinic-strip-ai/test.csv\")\nresult = []\nfor image in test_df[\"image_id\"].values:\n    slide = open_slide(f\"{test_dir}/{image}.tif\")\n    tiles = DeepZoomGenerator(slide, tile_size=600, overlap=0, limit_bounds=False)\n    level = tiles.level_count - 1\n    col, row = tiles.level_tiles[level]\n    inter_result = []\n    for row in tqdm(range(row)):\n        X_test = []\n        for col in range(col):\n            temp_tile = tiles.get_tile(level, (col, row))\n            rgb = temp_tile.convert(\"RGB\")\n            temp_tile_np = np.array(rgb)\n            img = temp_tile_np\n            mean = np.mean(img)\n            std = np.std(img)\n            area = []\n            try:\n                if not norm_HnE(img):\n                    continue\n                for channel in range(3):\n                    portrait = img[:, :, channel]\n                    th, threshed = cv2.threshold(portrait, 240, 255, cv2.THRESH_BINARY_INV)\n                    kernel = cv2.getStructuringElement(cv2.MORPH_ELLIPSE, (1,1))\n                    morphed = cv2.morphologyEx(threshed, cv2.MORPH_CLOSE, kernel)\n                    cnts = cv2.findContours(morphed, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)[-2]\n                    cnt = sorted(cnts, key=cv2.contourArea)[-1]\n                    x,y,w,h = cv2.boundingRect(cnt)\n                    dst = portrait[y:y+h, x:x+w]\n                    area_inter = (dst.shape[0] * dst.shape[1]) / (3600)\n                    area.append(area_inter)\n                area = np.array(area)\n                if mean < 220 and std > 8 and np.mean(area) > 10:\n                    # add image and predict\n                    img = st.resize(img, (IMG_SIZE, IMG_SIZE))\n                    X_test.append(img / 255.)\n                else:\n                    # discard image\n                    continue\n            except:\n                # discard image\n                continue\n                \n        # predicting for one row\n        if len(X_test) != 0:\n            X_test = np.array(X_test)\n            pred = model.predict(X_test)\n            pred = np.reshape(pred, (pred.shape[0],)).tolist()\n            for x in pred:\n                inter_result.append(x)\n        del X_test\n    # predict for single image\n    if len(inter_result) != 0:\n        result.append(np.mean(inter_result))\n    else:\n        result.append(0.5)\npatient_id = test_df[\"patient_id\"].values\nce_pred = np.array([1 - x for x in result])\nlaa_pred = np.array(result)\ndf = pd.DataFrame({ \"patient_id\" : patient_id,\"CE\" : ce_pred, \"LAA\" : laa_pred}).groupby(\"patient_id\").mean().reset_index()\ndf.to_csv('submission.csv',index=False)","metadata":{"papermill":{"duration":376.589704,"end_time":"2022-10-01T11:30:10.891825","exception":false,"start_time":"2022-10-01T11:23:54.302121","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-10-02T13:09:29.97364Z","iopub.execute_input":"2022-10-02T13:09:29.974489Z","iopub.status.idle":"2022-10-02T13:15:40.45405Z","shell.execute_reply.started":"2022-10-02T13:09:29.974449Z","shell.execute_reply":"2022-10-02T13:15:40.452947Z"},"trusted":true},"execution_count":null,"outputs":[]}]}