{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Import libraries and load csv","metadata":{}},{"cell_type":"code","source":"# install pyvips to handle these huge images\nfrom IPython.display import clear_output\n!conda install /kaggle/input/how-to-use-pyvips-offline/*.tar.bz2\nclear_output()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:01:36.599741Z","iopub.execute_input":"2022-08-18T12:01:36.600289Z","iopub.status.idle":"2022-08-18T12:02:32.623113Z","shell.execute_reply.started":"2022-08-18T12:01:36.600176Z","shell.execute_reply":"2022-08-18T12:02:32.621633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport PIL\nimport pyvips\n\nPIL.Image.MAX_IMAGE_PIXELS = 4896084492\nplt.rcParams[\"figure.figsize\"] = (12, 8)","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:02:32.625533Z","iopub.execute_input":"2022-08-18T12:02:32.626239Z","iopub.status.idle":"2022-08-18T12:02:33.577853Z","shell.execute_reply.started":"2022-08-18T12:02:32.626197Z","shell.execute_reply":"2022-08-18T12:02:33.575833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv(\"../input/mayo-clinic-strip-ai/train.csv\")\ntest = pd.read_csv(\"../input/mayo-clinic-strip-ai/test.csv\")\nother = pd.read_csv(\"../input/mayo-clinic-strip-ai/other.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:02:33.579825Z","iopub.execute_input":"2022-08-18T12:02:33.580379Z","iopub.status.idle":"2022-08-18T12:02:33.610907Z","shell.execute_reply.started":"2022-08-18T12:02:33.580325Z","shell.execute_reply":"2022-08-18T12:02:33.609919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Train size:\", len(train), \"unique patients:\", len(train.patient_id.unique()))\nprint(\"Test size:\", len(test), \"unique patients:\", len(test.patient_id.unique()))\nprint(\"Other size:\", len(other), \"unique patients:\", len(other.patient_id.unique()))","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:02:33.612958Z","iopub.execute_input":"2022-08-18T12:02:33.614035Z","iopub.status.idle":"2022-08-18T12:02:33.63379Z","shell.execute_reply.started":"2022-08-18T12:02:33.613993Z","shell.execute_reply":"2022-08-18T12:02:33.632107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# add images path to train and get width and height of images\ntrain[\"path\"] = \"../input/mayo-clinic-strip-ai/train/\" + train[\"image_id\"] + \".tif\"\ntrain[[\"width\", \"height\"]] = train[\"path\"].apply(lambda x: pd.Series(PIL.Image.open(x).size))","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:02:33.635575Z","iopub.execute_input":"2022-08-18T12:02:33.636458Z","iopub.status.idle":"2022-08-18T12:02:50.690865Z","shell.execute_reply.started":"2022-08-18T12:02:33.636365Z","shell.execute_reply":"2022-08-18T12:02:50.68976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# initial glimpse\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:02:50.692244Z","iopub.execute_input":"2022-08-18T12:02:50.692728Z","iopub.status.idle":"2022-08-18T12:02:50.71677Z","shell.execute_reply.started":"2022-08-18T12:02:50.692682Z","shell.execute_reply":"2022-08-18T12:02:50.715466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Study training dataset\n### CLASS DISTRIBUTION: IS THE DATASET BALANCED?","metadata":{}},{"cell_type":"code","source":"ax = sns.countplot(x='label', data=train)\nax.bar_label(ax.containers[0])\nplt.xticks(fontsize=14)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:02:50.720229Z","iopub.execute_input":"2022-08-18T12:02:50.721051Z","iopub.status.idle":"2022-08-18T12:02:51.350123Z","shell.execute_reply.started":"2022-08-18T12:02:50.721011Z","shell.execute_reply":"2022-08-18T12:02:51.348946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Unbalanced dataset. CE has more than double the samples of LAA**","metadata":{}},{"cell_type":"markdown","source":"### DISTRIBUTIONS OF CLINICS AND PATIENTS","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 2)\ncenter_dis = train[\"center_id\"].value_counts()\nsns.barplot(x=center_dis.index, y=center_dis.values, order=center_dis.index, \n            color=\"tab:blue\", ax=ax[0])\nax[0].set_xlabel(\"Clinic\", fontsize=14)\nax[0].set_ylabel(\"Number of images per Clinic\", fontsize=14)\n\nsns.countplot(x=train[\"patient_id\"].value_counts().values,\n             color=\"tab:blue\", ax=ax[1])\nax[1].set_xlabel(\"Number of images\", fontsize=14)\nax[1].set_ylabel(\"Number of patients\", fontsize=14)\n\nplt.xticks(fontsize=14)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:02:51.351458Z","iopub.execute_input":"2022-08-18T12:02:51.351807Z","iopub.status.idle":"2022-08-18T12:02:51.716964Z","shell.execute_reply.started":"2022-08-18T12:02:51.351775Z","shell.execute_reply":"2022-08-18T12:02:51.715697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Image on the left shows that there is the Clinic 11 that provides the majority of images. On the other hand, Clinic 8 and 9 give few images. **Could the Clinic impact the image itself?**\n\nImage on the right shows that very few patients has more than 3 images, while a lot of pictures come from one patient. This consideration is fundamental to avoid **data leakage**.","metadata":{}},{"cell_type":"markdown","source":"### CAN A PATIENT HAVE DIFFERENT LABELS?","metadata":{}},{"cell_type":"code","source":"p_count = train[\"patient_id\"].value_counts() \np_count = p_count[p_count > 1] \np_label = train.loc[train[\"patient_id\"].isin(p_count.index)].groupby(\"patient_id\")[\"label\"].apply(list)\n\nfor i, row in p_label.iteritems():\n    if len(set(row)) > 1:\n        print(i, \"DIFFERENT LABELS\", row)\n    break","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:02:51.718701Z","iopub.execute_input":"2022-08-18T12:02:51.719678Z","iopub.status.idle":"2022-08-18T12:02:51.732001Z","shell.execute_reply.started":"2022-08-18T12:02:51.719636Z","shell.execute_reply":"2022-08-18T12:02:51.73112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"One patient could have more than one image, but all his/her images have the same label","metadata":{}},{"cell_type":"markdown","source":"## Show images\nImages have **different** sizes. It seems that there is no correlation between image size and labels","metadata":{}},{"cell_type":"code","source":"sns.scatterplot(x=\"width\", y=\"height\", data=train, hue=\"label\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:02:51.734959Z","iopub.execute_input":"2022-08-18T12:02:51.735817Z","iopub.status.idle":"2022-08-18T12:02:52.013532Z","shell.execute_reply.started":"2022-08-18T12:02:51.735779Z","shell.execute_reply":"2022-08-18T12:02:52.012267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# img example. swap axes if the image is not rectangular\ndef roll_img(img):\n    if img.shape[0] > img.shape[1]:\n        return np.rollaxis(img, 0, 2)\n    return img\n\nfig, ax = plt.subplots(1, 2)\nimg1 = pyvips.Image.thumbnail(train.iloc[0][\"path\"], 4096).numpy()\nimg2 = pyvips.Image.thumbnail(train.iloc[1][\"path\"], 4096).numpy()\n\nimg1 = roll_img(img1)\nimg2 = roll_img(img2)\n\nax[0].imshow(img1)\nax[1].imshow(img2)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:02:52.015469Z","iopub.execute_input":"2022-08-18T12:02:52.015929Z","iopub.status.idle":"2022-08-18T12:04:26.656118Z","shell.execute_reply.started":"2022-08-18T12:02:52.015884Z","shell.execute_reply":"2022-08-18T12:04:26.654902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Check \"Other\" dataset","metadata":{}},{"cell_type":"code","source":"# check label distribution\nax = sns.countplot(x='label', data=other)\nplt.xticks(fontsize=14)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:04:26.65759Z","iopub.execute_input":"2022-08-18T12:04:26.657963Z","iopub.status.idle":"2022-08-18T12:04:26.791347Z","shell.execute_reply.started":"2022-08-18T12:04:26.657929Z","shell.execute_reply":"2022-08-18T12:04:26.789693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# patients intersection between train and other\ncommon_patient = set(train[\"patient_id\"].unique()).intersection(set(other[\"patient_id\"].unique()))\nprint(\"Common patient between train and other:\", len(common_patient))","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:04:26.793897Z","iopub.execute_input":"2022-08-18T12:04:26.794962Z","iopub.status.idle":"2022-08-18T12:04:26.807166Z","shell.execute_reply.started":"2022-08-18T12:04:26.794899Z","shell.execute_reply":"2022-08-18T12:04:26.805343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Check the major causes for \"other\" images","metadata":{}},{"cell_type":"code","source":"other_specified = other[\"other_specified\"].value_counts()\nsns.barplot(x=other_specified.index, y=other_specified.values, order=other_specified.index, \n            color=\"tab:blue\")\nplt.xticks(rotation=45)\nplt.xlabel(\"Other cause\", fontsize=14)\nplt.ylabel(\"Number of cases\", fontsize=14)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-18T12:04:26.809744Z","iopub.execute_input":"2022-08-18T12:04:26.810362Z","iopub.status.idle":"2022-08-18T12:04:27.071623Z","shell.execute_reply.started":"2022-08-18T12:04:26.810311Z","shell.execute_reply":"2022-08-18T12:04:27.070454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is difficult to extract meaningful information for the other dataset. It could be interesting to find a **correlation between the \"other specified\" and the training labels**. In this way, we could check if some causes are most responsible for CE or LAA.","metadata":{}},{"cell_type":"markdown","source":"## Comment and let me know what you think!","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}