{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Memory-saving image loading**\n\nIn this competition, a Notebook Exceeded Allowed Compute occurred due to insufficient memory at the stage of loading images. I share this notebook as a countermeasure.\nThe purpose of this notebook is to share how to load image data with low memory usage.\nNote that the code as is results in a Notebook Timeout for hidden test data. There is room for improvement, such as adding more processing time.\n","metadata":{}},{"cell_type":"markdown","source":"<b>1. Load module and Dataset</b>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport warnings\nwarnings.filterwarnings('ignore')\nimport numpy as np\nimport gc\nimport tensorflow as tf\nimport openslide\nfrom openslide import OpenSlide \nimport math\nfrom tensorflow.python.keras.models import load_model\nfrom tqdm import tqdm\nimport matplotlib.pyplot as plt\nfrom PIL import Image","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-09-19T14:50:31.667209Z","iopub.execute_input":"2022-09-19T14:50:31.668146Z","iopub.status.idle":"2022-09-19T14:50:31.675061Z","shell.execute_reply.started":"2022-09-19T14:50:31.6681Z","shell.execute_reply":"2022-09-19T14:50:31.673808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df  = pd.read_csv('../input/mayo-clinic-strip-ai/test.csv')\ntest_df[\"file_path\"]  = test_df[\"image_id\"].apply(lambda x: \"../input/mayo-clinic-strip-ai/test/\" + x + \".tif\")","metadata":{"execution":{"iopub.status.busy":"2022-09-19T14:38:40.35263Z","iopub.execute_input":"2022-09-19T14:38:40.353072Z","iopub.status.idle":"2022-09-19T14:38:40.380279Z","shell.execute_reply.started":"2022-09-19T14:38:40.353036Z","shell.execute_reply":"2022-09-19T14:38:40.37916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<b>2. low memory preprocessing</b>\n\nThe point of this method is that large images are read in separately. In this case, each image read in is resized to the desired size, enabling pre-processing while minimizing the amount of memory used.\nSince memory shortage seems to occur around the order of 10^8, I can escape memory shortage by keeping it below 10^4 in each of the vertical and horizontal directions.","metadata":{}},{"cell_type":"code","source":"def preprocess_lowMem(image_path, opt_size):\n    slide=OpenSlide(image_path)\n    h = int(slide.properties.get(\"openslide.level[0].height\"))\n    w = int(slide.properties.get(\"openslide.level[0].width\"))\n    ih = max(int(10**(np.log10(h) - 4)), 1)\n    iw = max(int(10**(np.log10(w) - 4)), 1)\n    print(f'h = {h}, w = {w}')\n    print(f'ih = {ih}, w = {iw}')\n    image_all = []\n    for i in range(ih):\n        for j in range(iw):        \n            dw = w//iw\n            dh = h//ih\n            region = (dw*j, dh *i)\n            if i == ih - 1:\n                dh = h - dh*i\n            if j == iw - 1:\n                dw = w - dw*j        \n            size = (dw, dh)\n            image = slide.read_region(region, 0, size)\n            image = tf.image.resize(np.array(image), (opt_size[0]//iw, opt_size[1]//ih))\n            if j==0:\n                image_all.append(np.array(image))\n            else:\n                image_all[i] = np.hstack((image_all[i], np.array(image)))\n            del image\n            gc.collect()\n\n    image_0 = image_all[0]\n    for i in range(1,len(image_all)):\n        image_0 = np.vstack((image_0, image_all[i]))\n    image_0 = np.array(tf.image.resize(np.array(image_0), opt_size))\n    del image_all\n    gc.collect()\n    return image_0","metadata":{"execution":{"iopub.status.busy":"2022-09-19T14:38:42.942838Z","iopub.execute_input":"2022-09-19T14:38:42.943309Z","iopub.status.idle":"2022-09-19T14:38:42.955975Z","shell.execute_reply.started":"2022-09-19T14:38:42.943273Z","shell.execute_reply":"2022-09-19T14:38:42.954629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test1=[]\nopt_size = (256, 256)\nfor i in test_df['file_path']:\n    x1=preprocess_lowMem(i, opt_size)\n    test1.append(x1)\ntest1=np.array(test1)","metadata":{"execution":{"iopub.status.busy":"2022-09-19T14:38:49.108329Z","iopub.execute_input":"2022-09-19T14:38:49.108747Z","iopub.status.idle":"2022-09-19T14:49:31.796602Z","shell.execute_reply.started":"2022-09-19T14:38:49.108713Z","shell.execute_reply":"2022-09-19T14:49:31.795579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<b>3. Check resized images</b>\n","metadata":{}},{"cell_type":"code","source":"wn=2\nhn=2\nfig, ax = plt.subplots(hn, wn, figsize=(16,5*hn))\nfor i in range(len(test_df['file_path'])):\n    ax[i//wn, i%wn].imshow(Image.fromarray((test1[i]).astype(np.uint8)))\n    \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-09-19T14:50:36.578592Z","iopub.execute_input":"2022-09-19T14:50:36.579014Z","iopub.status.idle":"2022-09-19T14:50:37.110466Z","shell.execute_reply.started":"2022-09-19T14:50:36.578978Z","shell.execute_reply":"2022-09-19T14:50:37.109119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<b>4. Predict with pretrained model</b>","metadata":{}},{"cell_type":"code","source":"# load model\npath=f'../input/simple-model/VGG16_simple.h5'\nmodel=load_model(path)\ny_pred = model.predict(test1)\n","metadata":{"execution":{"iopub.status.busy":"2022-09-19T14:53:53.793055Z","iopub.execute_input":"2022-09-19T14:53:53.794195Z","iopub.status.idle":"2022-09-19T14:54:01.857221Z","shell.execute_reply.started":"2022-09-19T14:53:53.794154Z","shell.execute_reply":"2022-09-19T14:54:01.856114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub = pd.DataFrame(test_df[\"patient_id\"].copy())\nsub[\"CE\"] = y_pred\nsub[\"CE\"] = sub[\"CE\"].apply(lambda x : 0 if x<0 else x)\nsub[\"CE\"] = sub[\"CE\"].apply(lambda x : 1 if x>1 else x)\nsub[\"LAA\"] = 1- sub[\"CE\"]\nsub = sub.groupby(\"patient_id\").mean()\nsub = sub[[\"CE\", \"LAA\"]].round(6).reset_index()\nsub","metadata":{"execution":{"iopub.status.busy":"2022-09-19T14:54:24.856403Z","iopub.execute_input":"2022-09-19T14:54:24.856823Z","iopub.status.idle":"2022-09-19T14:54:24.88953Z","shell.execute_reply.started":"2022-09-19T14:54:24.856788Z","shell.execute_reply":"2022-09-19T14:54:24.888378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.to_csv(\"submission.csv\", index = False)","metadata":{"execution":{"iopub.status.busy":"2022-09-19T14:54:27.512543Z","iopub.execute_input":"2022-09-19T14:54:27.512951Z","iopub.status.idle":"2022-09-19T14:54:27.521463Z","shell.execute_reply.started":"2022-09-19T14:54:27.512903Z","shell.execute_reply":"2022-09-19T14:54:27.52033Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<b>5. Future work</b>\n\nThis method can change the size of the output image as training data. In conjunction with the following discussion of aspect ratio, it may be possible to separate the preprocessing of the input data.\nhttps://www.kaggle.com/code/matsutake94/mayo-aspect-ratio-analysis\n\nIn addition, as mentioned at the beginning of this notebbok, this method has an issue with processing speed. Even if the memory shortage is improved, further improvement is required because the processing time must be within a predetermined range.\nI hope that this Notebook will be of help to you.\n","metadata":{}}]}