{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# MAYO Clinic Strip AI competition: Tile generation and thresholding preprocessing notebook.","metadata":{}},{"cell_type":"markdown","source":"The notebook is used for making the preprocessed dataset, which is a part of MAYO Clinic - STRIP AI competition.\nThe preprocessed training set from the following notebook is available at the following link -> [Dataset Link](https://www.kaggle.com/datasets/tr1gg3rtrash/mayo-clinic)\n\nExample Images from the dataset.\n\n<div> \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4636937%2Ff8fff1fa62dc75cdb3b38981b7a147fd%2F0b7871_0_2_33_15.png?generation=1663412986808182&alt=media\" width=\"300\" height=\"300\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4636937%2Fd6ed6a9702983da36ee9225a9c19e5fc%2F0b25f8_0_13_8_16.png?generation=1663413194154320&alt=media\" width=\"300\" height=\"300\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4636937%2F913f27973e1bd16b956a4551a2e7850d%2F0ba49d_0_3_6_17.png?generation=1663413413713409&alt=media\" width=\"300\" height=\"300\">\n</div>","metadata":{}},{"cell_type":"markdown","source":"# Strategy Used   \n\nThe dataset was downloaded from the Kaggle website. Using Openslide, 600 x 600 slides were extracted. From the last level of DeepZoomGenerator, the slides were saved in png format. Then using the thresholding technique of calculating the area of slide contents in the image as well as the total content in the image. If the slide area consists of more than 30% of the total slide area, the image was kept otherwise it was discarded. Totally white images from tiles were deleted manually from the dataset.","metadata":{}},{"cell_type":"markdown","source":"# Library Install and imports","metadata":{}},{"cell_type":"code","source":"!pip install openslide-python","metadata":{"execution":{"iopub.status.busy":"2023-03-17T12:34:36.361447Z","iopub.execute_input":"2023-03-17T12:34:36.36187Z","iopub.status.idle":"2023-03-17T12:34:47.867014Z","shell.execute_reply.started":"2023-03-17T12:34:36.361823Z","shell.execute_reply":"2023-03-17T12:34:47.865395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from PIL import Image\nimport os\nfrom openslide import OpenSlide\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport openslide\nfrom openslide import open_slide\nfrom tqdm import tqdm\nfrom openslide.deepzoom import DeepZoomGenerator","metadata":{"execution":{"iopub.status.busy":"2023-03-19T15:49:05.091802Z","iopub.execute_input":"2023-03-19T15:49:05.092279Z","iopub.status.idle":"2023-03-19T15:49:05.207968Z","shell.execute_reply.started":"2023-03-19T15:49:05.092195Z","shell.execute_reply":"2023-03-19T15:49:05.206939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%rm -rf /kaggle/working/test-tile_dataset","metadata":{"execution":{"iopub.status.busy":"2023-03-17T12:55:52.440012Z","iopub.execute_input":"2023-03-17T12:55:52.440471Z","iopub.status.idle":"2023-03-17T12:55:53.977017Z","shell.execute_reply.started":"2023-03-17T12:55:52.440431Z","shell.execute_reply":"2023-03-17T12:55:53.975445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Tile Generation Section\n\n#### In for loop, slicing of df has been performed as the notebook takes a lot of time to preprocess such large files present in train and other directory. Make sure to remove that slicing when running the notebook. The notebook takes a lot of time to preprocess the dataset but output is worth the wait as it increases the dataset to over 100k+ images. ","metadata":{}},{"cell_type":"code","source":"main_dir = \"test-tile_dataset\"\nos.mkdir(main_dir)\nmain_folder = \"../input/mayo-clinic-strip-ai\"\nfor folder in [\"train,test\"]:\n  os.mkdir(f\"{main_dir}/{folder}\")\n  df = pd.read_csv(f\"{main_folder}/{folder}.csv\")\n  df_new = []\n  for file_values in tqdm(df.values[:]):\n    file_name = file_values[0]\n    slide = open_slide(f\"{main_folder}/{folder}/{file_name}.tif\")\n    tiles = DeepZoomGenerator(slide, tile_size=600, overlap=0, limit_bounds=False)\n    level = tiles.level_count - 1\n    col, row = tiles.level_tiles[level]\n    for row in range(row):\n      for col in range(col):\n        tile_name = os.path.join(main_dir, folder, f\"{file_name}_{row}_{col}_{level}\")\n        temp_tile = tiles.get_tile(level, (col, row))\n        rgb = temp_tile.convert(\"RGB\")\n        temp_tile_np = np.array(rgb)\n        plt.imsave(tile_name + \".png\", temp_tile_np)\n        updated_ds = [x for x in file_values]\n        updated_ds[0] = tile_name\n        df_new.append(updated_ds)\n  dataframe = pd.DataFrame(df_new, columns=df.columns)\n  dataframe.to_csv(f\"{main_dir}/{folder}.csv\", index=True)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-03-19T15:49:12.45327Z","iopub.execute_input":"2023-03-19T15:49:12.453646Z","iopub.status.idle":"2023-03-19T15:58:11.861977Z","shell.execute_reply.started":"2023-03-19T15:49:12.453592Z","shell.execute_reply":"2023-03-19T15:58:11.860798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\ntrain_filenames = tf.io.gfile.glob(str('/kaggle/input/mayo-clinic/First Experiment/train/CE/*'))\nCOUNT_CE = len([filename for filename in train_filenames])\nprint(\"CE images count in training set: \" + str(COUNT_CE))","metadata":{"execution":{"iopub.status.busy":"2023-03-17T12:56:48.740709Z","iopub.execute_input":"2023-03-17T12:56:48.741185Z","iopub.status.idle":"2023-03-17T12:57:07.903752Z","shell.execute_reply.started":"2023-03-17T12:56:48.741143Z","shell.execute_reply":"2023-03-17T12:57:07.902433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from skimage.filters import threshold_otsu\nimport numpy as np\nimport cv2\ndef ostu(img):\n    img = cv2.imread(img)\n    gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)\n    thre = threshold_otsu(gray)\n#     ret1, th =  cv2.threshold(gray,thre,255,cv2.THRESH_BINARY_INV)\n    ret, th = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)\n    kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (5, 5))\n    th = cv2.morphologyEx(th.astype(np.uint8), cv2.MORPH_CLOSE, kernel)\n    th[th>0] = 1\n    th[gray==0] = 0\n    return th","metadata":{"execution":{"iopub.status.busy":"2023-03-17T12:57:07.906443Z","iopub.execute_input":"2023-03-17T12:57:07.906816Z","iopub.status.idle":"2023-03-17T12:57:07.91556Z","shell.execute_reply.started":"2023-03-17T12:57:07.90678Z","shell.execute_reply":"2023-03-17T12:57:07.914181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def check_msk(msk, threshold=0.8):\n    if msk.sum()/(msk.size) > threshold:\n        return True\n    return False","metadata":{"execution":{"iopub.status.busy":"2023-03-17T13:43:31.670768Z","iopub.status.idle":"2023-03-17T13:43:31.671224Z","shell.execute_reply.started":"2023-03-17T13:43:31.671007Z","shell.execute_reply":"2023-03-17T13:43:31.671026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"c = []\nfor i in train_filenames:\n    th = ostu(i)\n    if check_msk(th):\n        c.append(i)","metadata":{"execution":{"iopub.status.busy":"2023-03-17T12:57:51.249637Z","iopub.execute_input":"2023-03-17T12:57:51.250153Z","iopub.status.idle":"2023-03-17T13:43:31.58652Z","shell.execute_reply.started":"2023-03-17T12:57:51.250112Z","shell.execute_reply":"2023-03-17T13:43:31.584906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(c)","metadata":{"execution":{"iopub.status.busy":"2023-03-17T13:43:31.589334Z","iopub.execute_input":"2023-03-17T13:43:31.59028Z","iopub.status.idle":"2023-03-17T13:43:31.600307Z","shell.execute_reply.started":"2023-03-17T13:43:31.590229Z","shell.execute_reply":"2023-03-17T13:43:31.598968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!mkdir train-ce\nimport shutil\nfor i in c:\n    shutil.move(i, '/kaggle/working/train-ce')","metadata":{"execution":{"iopub.status.busy":"2023-03-17T13:58:44.338169Z","iopub.execute_input":"2023-03-17T13:58:44.338658Z","iopub.status.idle":"2023-03-17T13:58:45.517678Z","shell.execute_reply.started":"2023-03-17T13:58:44.338623Z","shell.execute_reply":"2023-03-17T13:58:45.51556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!zip -r file.zip /kaggle/working/test-tile_dataset","metadata":{"execution":{"iopub.status.busy":"2023-03-17T12:50:48.213212Z","iopub.execute_input":"2023-03-17T12:50:48.213685Z","iopub.status.idle":"2023-03-17T12:51:34.094594Z","shell.execute_reply.started":"2023-03-17T12:50:48.213651Z","shell.execute_reply":"2023-03-17T12:51:34.093028Z"},"trusted":true},"execution_count":null,"outputs":[]}]}