{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"from IPython.core.display import display, HTML, Javascript\nimport IPython.display\n\n\n# CSS styling for markdown\nstyling = \"\"\"\n    <style>\n        .main-heading{\n            background-color: #4d79ff;\n            color: white !important;\n            font-family: Helvetica;\n            font-size: 32px !important;\n            padding: 12px, 12px;\n            margin-bottom: 5px;\n            border-radius: 4px;\n            box-shadow: rgba(0, 0, 0, 0.19) 0px 10px 20px, rgba(0, 0, 0, 0.23) 0px 6px 6px;\n        }\n\n        .sub-heading{\n            width: auto !important;\n            background-color: #4d79ff;\n            color: white !important;\n            font-family: Helvetica;\n            font-size: 24px !important;\n            padding: 10px 12px;\n            margin-bottom: 3px;\n            box-shadow: rgba(0, 0, 0, 0.16) 0px 3px 6px, rgba(0, 0, 0, 0.23) 0px 3px 6px;\n        }\n        \n        .default-font-color{\n            color: rgba(0,0,0,0.7) !important;\n        }\n    \n        \n        .highlight-orange {\n            background: #b2481b;\n            color: white;\n            padding: 1px 3px;\n        }\n        \n        .highlight-cream{\n            background: #cd8b59;\n            color: white;\n            padding: 1px 3px;\n        }    \n        \n        .salary-diff-table tr th{\n            text-transform: upppercase;\n        }\n        \n        .salary-diff-table th{    \n            color: #444;\n            font-weight: bold !important;\n            text-transform: uppercase;\n            vertical-align: bottom !important;\n            text-align: center !important;\n            height: 10px !important;\n            padding: 0 !important;\n        }\n\n        .salary-diff-table th, .salary-diff-table td{\n            width: 120px;\n            height: 35px;\n        }\n\n        .salary-diff-table td:not(td:first-child){\n            padding: 15px;\n            font-size: 20px;\n            border: 2px solid white;\n            background: #ddd;\n            text-align: center;\n            color: #222;\n        }\n\n        .salary-diff-table td:first-child, .salary-diff-table th:first-child{\n            font-weight: bold;\n            width: 150px;\n            color: #444;\n            padding-right: 10px;\n            text-align: right !important;\n            text-transform: uppercase;\n        }\n\n        .cell-highlight-orange{\n            color: #efefef !important;\n            background: #b2481b !important;  \n        }\n\n        .cell-highlight-black{\n            color: #efefef !important;\n            background: #222 !important;  \n        }\n        \n        .related-queries-table th{\n            font-size: 14px;\n            font-weight: bold !important;\n            text-align: left;\n            background: #4d79ff !important;\n            border: none !important;\n            color: #efefef;\n            text-transform: uppercase;\n            padding: 4px;\n        }\n    \n        .related-queries-table td:nth-child(even), .related-queries-table th:nth-child(even){\n            border-right: 30px solid white !important;\n            width: 200px;\n        }\n\n        .related-queries-table td{\n            padding: 10px 10px 10px 15px !important;\n            border: 2px solid white !important;\n            border-bottom: 1px solid #666 !important;\n            border-top: 1px solid #666 !important;\n            background: #ddd;\n            font-size: 15px;\n        }\n\n        .related-queries-table tr:nth-child(even) td{\n            background: #efefef !important;\n        }\n        \n        .sidenote{\n            font-size: 20px;\n            border: 2px solid #d7d7d7;\n            padding: 5px 10px 2px;\n            box-shadow: 1px 1px 2px 1px rgba(0,0,0,0.3);\n            margin-bottom: 3px;\n        }\n    </style>\n\"\"\"\n\nHTML(styling)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-08T13:09:17.1378Z","iopub.execute_input":"2022-07-08T13:09:17.138223Z","iopub.status.idle":"2022-07-08T13:09:17.182378Z","shell.execute_reply.started":"2022-07-08T13:09:17.138137Z","shell.execute_reply":"2022-07-08T13:09:17.181061Z"},"_kg_hide-output":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class='sub-heading'>\nAbout this Notebook</div>\n<br>\n\nHi Kagglers!,\n\nI have seen that Kaggle competitions require a lot of domain knowledge and the main problems of a Kaggle beginner is not lack of Machine learning Knowledge but more often than not it is lack of domain knowledge.There are always a lot of great kernels regarding different ways of solving the problems but only a few handful address the **problems of domain knowledge** and **getting started**. In this notebook , I will start with explanation to **ischemic stroke** and its detection using **digital pathology**  and I will built on that to explain the dataset,perform exploratory data analysis.","metadata":{"execution":{"iopub.status.busy":"2022-07-07T08:48:49.826918Z","iopub.execute_input":"2022-07-07T08:48:49.827332Z","iopub.status.idle":"2022-07-07T08:48:49.835217Z","shell.execute_reply.started":"2022-07-07T08:48:49.827299Z","shell.execute_reply":"2022-07-07T08:48:49.833429Z"}}},{"cell_type":"markdown","source":"<div class='sub-heading'>Dive deep into domain knowledge¶</div>\n\n### So lets start with the domain knowledge and Address the first question\n\n\n### Q1) What is ischemic stroke?\n\nFirst Lets understand stroke, A stroke is a medical emergency. There are two types - **ischemic** and **hemorrhagic.** Ischemic stroke is the more common type. It is usually caused by a blood clot that blocks or plugs a blood vessel in the brain. This keeps blood from flowing to the brain. Within minutes, brain cells begin to die. Another cause is stenosis, or narrowing of the artery. This can happen because of atherosclerosis,\n\n\n\n![Ischemic](https://thumbs.gfycat.com/BouncyDapperChamois-size_restricted.gif \"Ischemic\")\n\nGit Scource: https://gfycat.com/gifs/search/ischemic\n\n\n\n\n\n\n","metadata":{}},{"cell_type":"code","source":"from IPython.display import YouTubeVideo\nYouTubeVideo('Ffk-L9ODEO8', end=70,width=800, height=500)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-07-08T13:09:17.18615Z","iopub.execute_input":"2022-07-08T13:09:17.187097Z","iopub.status.idle":"2022-07-08T13:09:17.321498Z","shell.execute_reply.started":"2022-07-08T13:09:17.187051Z","shell.execute_reply":"2022-07-08T13:09:17.319963Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Q2) What is whole slide imaging?\n\nWhole slide imaging is a disruptive technology where glass slides are scanned to produce digital images. There have been significant advances in whole slide scanning hardware and software that have allowed for ready access of whole slide images. The digital images, or whole slide images, can be viewed comparable to glass slides in a microscope, as digital files. Whole slide imaging has increased in adoption among pathologists, pathology departments, and scientists for clinical, educational, and research initiatives.\n\nTo know how these images are formed you can refer to below video","metadata":{}},{"cell_type":"code","source":"YouTubeVideo('Ua4Qp0ZwIYg', width=800, height=500)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:17.324342Z","iopub.execute_input":"2022-07-08T13:09:17.325171Z","iopub.status.idle":"2022-07-08T13:09:17.463312Z","shell.execute_reply.started":"2022-07-08T13:09:17.32512Z","shell.execute_reply":"2022-07-08T13:09:17.461955Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From competition discription states that Stroke remains the second-leading cause of death worldwide. Each year in the United States, over 700,000 individuals experience an ischemic stroke caused by a blood clot blocking an artery to the brain. A second stroke (23% of total events are recurrent) worsens the chances of the patient’s survival. However, subsequent strokes may be mitigated if physicians can determine stroke etiology, which influences the therapeutic management following stroke events. \n\nDuring the last decade, **mechanical thrombectomy** has become the standard of care treatment for acute ischemic stroke from large vessel occlusion. As a result, retrieved clots became amenable to analysis.\n\nRefer below video if you want to know more about mechanical thrombectomy","metadata":{}},{"cell_type":"code","source":"YouTubeVideo('eMeEsrh1oNg', width=800, height=500)","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-07-08T13:09:17.465304Z","iopub.execute_input":"2022-07-08T13:09:17.466072Z","iopub.status.idle":"2022-07-08T13:09:17.593261Z","shell.execute_reply.started":"2022-07-08T13:09:17.466026Z","shell.execute_reply":"2022-07-08T13:09:17.591929Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class='sub-heading'> Overview: </div>\n\n<br>\n\nFrom above data we understood that when every any patient get acute ischemic stroke. The sample tissue will be collected and it will be cutted into thin section slices and presented on glass slides using digital pathology  technology these same are scanned and stored in digital format to view on computer screens so that a pathologist can analyze these scans and classifiy the types of acute ischemic stroke subtypes that is cardiac and large artery atherosclerosis. \n\n\nNow the we need to build a model to assist pathalogist to identitfes the type of ischemic stroke that is either cardiac or large artery atherosclersis based on scanned whole side images","metadata":{}},{"cell_type":"markdown","source":"<div class='sub-heading'> Exploratory Data Analysis: 📊👨‍🔬</div>\n\n\n### Now Lets start with understanding data and get insights from it...","metadata":{}},{"cell_type":"markdown","source":"<div style=\";font-family:'Times';font-size:30px;color:  #4d79ff\" >📌 <b>Importing Libraries</b></div>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os\nfrom openslide import open_slide\nimport openslide\nfrom PIL import Image\nimport numpy as np\nfrom matplotlib import pyplot as plt\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:17.597175Z","iopub.execute_input":"2022-07-08T13:09:17.59803Z","iopub.status.idle":"2022-07-08T13:09:18.994546Z","shell.execute_reply.started":"2022-07-08T13:09:17.597986Z","shell.execute_reply":"2022-07-08T13:09:18.992559Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for root, dir_, files in os.walk('../input/mayo-clinic-strip-ai/'):\n    if not dir_:\n        print(root, dir_, len(files))\n    else:\n        print(root, dir_, files)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:18.996751Z","iopub.execute_input":"2022-07-08T13:09:19.001851Z","iopub.status.idle":"2022-07-08T13:09:19.358304Z","shell.execute_reply.started":"2022-07-08T13:09:19.001801Z","shell.execute_reply":"2022-07-08T13:09:19.35664Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that we have got three folders **'other', 'test', 'train'** \n\n**other** contain 396 tif files\n\n**test** contains 4 tif files,\n\n**train** contains 754 files,\n\nand four csv files - **'sample_submission.csv', 'train.csv', 'test.csv', 'other.csv'**\n\nLets explore these csv files first then will look into folder and their tif format file","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('../input/mayo-clinic-strip-ai/train.csv')\ntrain_df","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:19.370679Z","iopub.execute_input":"2022-07-08T13:09:19.377909Z","iopub.status.idle":"2022-07-08T13:09:19.472877Z","shell.execute_reply.started":"2022-07-08T13:09:19.377846Z","shell.execute_reply":"2022-07-08T13:09:19.471265Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, In train csv we have five columns and 754 rows, as mentioned in three columns names are ending with IDs these should be the unique identifies to identify the image and its patients with center ID uniquely and from the label column we can see that their are only two values we need to identify based on other columns. But what does image_num inndicates? let see in further analysis.","metadata":{}},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:19.478708Z","iopub.execute_input":"2022-07-08T13:09:19.482667Z","iopub.status.idle":"2022-07-08T13:09:19.519452Z","shell.execute_reply.started":"2022-07-08T13:09:19.482586Z","shell.execute_reply":"2022-07-08T13:09:19.517862Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['image_num'].unique()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:19.523775Z","iopub.execute_input":"2022-07-08T13:09:19.526957Z","iopub.status.idle":"2022-07-08T13:09:19.544941Z","shell.execute_reply.started":"2022-07-08T13:09:19.526911Z","shell.execute_reply":"2022-07-08T13:09:19.543417Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For **image_num** we are getting five unique number 🤔 what does these indicates ?\n\nLets find out..","metadata":{}},{"cell_type":"code","source":"train_df['image_id']","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:19.554339Z","iopub.execute_input":"2022-07-08T13:09:19.556427Z","iopub.status.idle":"2022-07-08T13:09:19.584167Z","shell.execute_reply.started":"2022-07-08T13:09:19.556381Z","shell.execute_reply":"2022-07-08T13:09:19.578842Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**The image_num are the image_id value concatenated with \"_\" char**","metadata":{}},{"cell_type":"code","source":"test_df = pd.read_csv('../input/mayo-clinic-strip-ai/test.csv')\ntest_df","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:19.591672Z","iopub.execute_input":"2022-07-08T13:09:19.592611Z","iopub.status.idle":"2022-07-08T13:09:19.619384Z","shell.execute_reply.started":"2022-07-08T13:09:19.592564Z","shell.execute_reply":"2022-07-08T13:09:19.618061Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Same as train except label column. Which our model needs to classify ","metadata":{}},{"cell_type":"code","source":"ss_df = pd.read_csv('../input/mayo-clinic-strip-ai/sample_submission.csv')\nss_df","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:19.621832Z","iopub.execute_input":"2022-07-08T13:09:19.622678Z","iopub.status.idle":"2022-07-08T13:09:19.647882Z","shell.execute_reply.started":"2022-07-08T13:09:19.622632Z","shell.execute_reply":"2022-07-08T13:09:19.646311Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the sample submission we need to provide two probabilty score for each of its class based on unique patient_id from the test csv","metadata":{}},{"cell_type":"code","source":"other_df = pd.read_csv('../input/mayo-clinic-strip-ai/other.csv')\nother_df","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:19.650337Z","iopub.execute_input":"2022-07-08T13:09:19.651316Z","iopub.status.idle":"2022-07-08T13:09:19.688013Z","shell.execute_reply.started":"2022-07-08T13:09:19.65127Z","shell.execute_reply":"2022-07-08T13:09:19.686619Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"other_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:19.690223Z","iopub.execute_input":"2022-07-08T13:09:19.691042Z","iopub.status.idle":"2022-07-08T13:09:19.737371Z","shell.execute_reply.started":"2022-07-08T13:09:19.690999Z","shell.execute_reply":"2022-07-08T13:09:19.735738Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"other_df['other_specified'].unique()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:19.740038Z","iopub.execute_input":"2022-07-08T13:09:19.740908Z","iopub.status.idle":"2022-07-08T13:09:19.76334Z","shell.execute_reply.started":"2022-07-08T13:09:19.740869Z","shell.execute_reply":"2022-07-08T13:09:19.761634Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center>\n<table class=\"related-queries-table\"> \n<tr> <th> </th> <th> other_specified<br> </th>                \n<tr> <td> 1 </td> <td>Hypercoagulable </td>        \n<tr> <td> 2 </td> <td>Dissection </td> \n<tr> <td> 3 </td> <td>Catheter </td> \n<tr> <td> 4 </td> <td>Stent thrombosis </td> \n<tr> <td> 5 </td> <td>PFO </td> \n<tr> <td> 6 </td> <td>Trauma </td> \n<tr> <td> 7 </td> <td>Takayasu vasculitis </td> \n<tr> <td> 8 </td> <td>tumor embolization </td> \n<tr> <td> 9 </td> <td>Endocarditis </td> \n\n\n</table>\n</center>\n\n<div class=\"sidenote\"> We have nine different type of other specified values</div>","metadata":{}},{"cell_type":"markdown","source":"<div class='sub-heading'> Univarient Analysis</div>","metadata":{}},{"cell_type":"code","source":"df = pd.DataFrame(train_df['label'])\nplt.rcParams[\"figure.figsize\"] = [10.00, 7.50]\nplt.rcParams[\"figure.autolayout\"] = True\n\nax = sns.countplot(x=\"label\", data=df)\n\nfor p in ax.patches:\n    ax.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.25, p.get_height()+0.01))\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:19.771516Z","iopub.execute_input":"2022-07-08T13:09:19.77385Z","iopub.status.idle":"2022-07-08T13:09:20.174312Z","shell.execute_reply.started":"2022-07-08T13:09:19.773804Z","shell.execute_reply":"2022-07-08T13:09:20.172246Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above bar plot we can clearly say that its **not a balanced dataset** as CE has more number of instances compare to LAA labels","metadata":{}},{"cell_type":"code","source":"df = pd.DataFrame(train_df['center_id'])\nplt.rcParams[\"figure.figsize\"] = [10.00, 7.50]\nplt.rcParams[\"figure.autolayout\"] = True\n\nax = sns.countplot(x=\"center_id\", data=df)\n\nfor p in ax.patches:\n    ax.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.25, p.get_height()+0.01))\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:20.179953Z","iopub.execute_input":"2022-07-08T13:09:20.183076Z","iopub.status.idle":"2022-07-08T13:09:20.797433Z","shell.execute_reply.started":"2022-07-08T13:09:20.183027Z","shell.execute_reply":"2022-07-08T13:09:20.795326Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"11th center having the highest data comparing to other center and 8th and 9th being the lowest","metadata":{}},{"cell_type":"code","source":"df = pd.DataFrame(train_df['image_num'])\nplt.rcParams[\"figure.figsize\"] = [10.00, 7.50]\nplt.rcParams[\"figure.autolayout\"] = True\n\nax = sns.countplot(x=\"image_num\", data=df)\n\nfor p in ax.patches:\n   ax.annotate('{:.1f}'.format(p.get_height()), (p.get_x()+0.25, p.get_height()+0.01))\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:20.800343Z","iopub.execute_input":"2022-07-08T13:09:20.800864Z","iopub.status.idle":"2022-07-08T13:09:21.310106Z","shell.execute_reply.started":"2022-07-08T13:09:20.800819Z","shell.execute_reply":"2022-07-08T13:09:21.308608Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Image number **0** having more count values compare to rest and image number **4** being very less in count","metadata":{}},{"cell_type":"code","source":"df = (train_df.groupby([\"patient_id\",\"label\"])[\"image_num\"].count().reset_index(name='image_count'))","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:21.317362Z","iopub.execute_input":"2022-07-08T13:09:21.31809Z","iopub.status.idle":"2022-07-08T13:09:21.345276Z","shell.execute_reply.started":"2022-07-08T13:09:21.318045Z","shell.execute_reply":"2022-07-08T13:09:21.340076Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:21.347548Z","iopub.execute_input":"2022-07-08T13:09:21.348027Z","iopub.status.idle":"2022-07-08T13:09:21.386454Z","shell.execute_reply.started":"2022-07-08T13:09:21.347981Z","shell.execute_reply":"2022-07-08T13:09:21.384673Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(train_df.image_id.unique()), len(train_df.center_id.unique()), len(train_df.patient_id.unique())","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:21.388708Z","iopub.execute_input":"2022-07-08T13:09:21.389584Z","iopub.status.idle":"2022-07-08T13:09:21.407338Z","shell.execute_reply.started":"2022-07-08T13:09:21.389508Z","shell.execute_reply":"2022-07-08T13:09:21.405103Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class='sub-heading'>Multivarient analysis</div>","metadata":{}},{"cell_type":"markdown","source":"# Patient with more than two images","metadata":{}},{"cell_type":"code","source":"temp = (train_df.groupby([\"patient_id\",\"label\"])[\"image_num\"].count().reset_index(name='image_count'))\ntemp = temp[temp['image_count']>2]\ntemp.style.background_gradient(cmap='Reds')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:21.41026Z","iopub.execute_input":"2022-07-08T13:09:21.411278Z","iopub.status.idle":"2022-07-08T13:09:21.681376Z","shell.execute_reply.started":"2022-07-08T13:09:21.411227Z","shell.execute_reply":"2022-07-08T13:09:21.679964Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only four patient having 5 images and 5 is the maximum images of any patient and 1 is the minimum image for any patient","metadata":{}},{"cell_type":"markdown","source":"# Labels with center IDs","metadata":{}},{"cell_type":"code","source":"center_grp = train_df.groupby(['center_id','label'])['image_id'].count().reset_index()\ncenter_grp.rename(columns={'image_id':'count'},inplace=True)\ncenter_grp.style.background_gradient(cmap='Reds')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-08T13:09:21.683302Z","iopub.execute_input":"2022-07-08T13:09:21.68459Z","iopub.status.idle":"2022-07-08T13:09:21.729411Z","shell.execute_reply.started":"2022-07-08T13:09:21.68454Z","shell.execute_reply":"2022-07-08T13:09:21.7279Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.rcParams['figure.dpi'] = 100\nbackground_color = '#ffffff'\nfig = plt.figure(figsize=(10,5), facecolor='#ffffff')\n\ngs = fig.add_gridspec(1,1)\ngs.update(wspace=0.3, hspace=0.4)\nlocals()[\"ax\"+str(0)] = fig.add_subplot(gs[0, 0])\nlocals()[\"ax\"+str(0)].set_facecolor(background_color)\nfor s in ['left', 'right', 'top', 'bottom']:\n    locals()[\"ax\"+str(0)].spines[s].set_visible(False)\n_ = sns.barplot(x=\"center_id\", y=\"count\", hue=\"label\", data=center_grp, ax=locals()[\"ax\"+str(0)], \n                palette=\"Reds\")\nlocals()[\"ax\"+str(0)].set_title('Number of labels by center',fontsize=14, fontweight='bold')\ngs.tight_layout(fig, rect=[0, 0, 1, 1])\nplt.show()                                                                ","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:21.735029Z","iopub.execute_input":"2022-07-08T13:09:21.735838Z","iopub.status.idle":"2022-07-08T13:09:22.528619Z","shell.execute_reply.started":"2022-07-08T13:09:21.735793Z","shell.execute_reply":"2022-07-08T13:09:22.526673Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<div class='sub-heading'>Image EDA</div>\n<br>\n\n### Now we can finally move on to Image EDA . BUT since we are complete beginners, let's first understand the format of image that is provided to us and all the image related jargons that we will be using further.\n\n# Q1) What is .tif format and Why it is used?\n\nTagged Image File Format (TIF) is a variable-resolution bitmapped image format developed by Aldus (now part of Adobe) in 1986. TIFF is very common for transporting color or gray-scale images into page layout applications, but is less suited to delivering web content.\n\n## Reasons for Usage:\n\n* IFF files are large and of very high quality. Baseline TIFF images are highly portable; most graphics, desktop publishing, and word processing applications understand them.\n* The TIFF specification is readily extensible, though this comes at the price of some of its portability. Many applications incorporate their own extensions, but a number of application-independent extensions are recognized by most programs.\n* Four types of baseline TIFF images are available: bilevel (black and white), gray scale, palette (i.e., indexed), and RGB (i.e., true color). RGB images may store up to 16.7 million colors. Palette and gray-scale images are limited to 256 colors or shades. A common extension of TIFF also allows for CMYK images.\n* TIFF files may or may not be compressed. A number of methods may be used to compress TIFF files, including the Huffman and LZW algorithms. Even compressed, TIFF files are usually much larger than similar GIF or JPEG files.\n* Because the files are so large and because there are so many possible variations of each TIFF file type, few web browsers can display them without plug-ins.\n","metadata":{"execution":{"iopub.status.busy":"2022-07-08T06:14:42.64886Z","iopub.execute_input":"2022-07-08T06:14:42.64942Z","iopub.status.idle":"2022-07-08T06:14:42.719456Z","shell.execute_reply.started":"2022-07-08T06:14:42.649384Z","shell.execute_reply":"2022-07-08T06:14:42.717787Z"}}},{"cell_type":"code","source":"YouTubeVideo('prm_GAXsWcI', width=800, height=500)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:22.530484Z","iopub.execute_input":"2022-07-08T13:09:22.531565Z","iopub.status.idle":"2022-07-08T13:09:22.664463Z","shell.execute_reply.started":"2022-07-08T13:09:22.531497Z","shell.execute_reply":"2022-07-08T13:09:22.663049Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I will be using openslide to display images as I learned it in this competition from a very informative kernel: https://www.kaggle.com/wouterbulten/getting-started-with-the-panda-dataset\nThe benefit of OpenSlide is that we can load arbitrary regions of the slide, without loading the whole image in memory. Want to interactively view a slide? We have added an interactive viewer to this notebook in the last section.\n\nYou can read more about the OpenSlide python bindings in the documentation: https://openslide.org/api/python/ and you can also refer to the below video.\n","metadata":{}},{"cell_type":"code","source":"YouTubeVideo('QntLBvUZR5c', width=800, height=500)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:22.666333Z","iopub.execute_input":"2022-07-08T13:09:22.667056Z","iopub.status.idle":"2022-07-08T13:09:22.802364Z","shell.execute_reply.started":"2022-07-08T13:09:22.66701Z","shell.execute_reply":"2022-07-08T13:09:22.800815Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import openslide\n\n'''\nExample for using Openslide to display an image\n'''\n\n\n# Open the image (does not yet read the image into memory)\nexample = openslide.OpenSlide('../input/mayo-clinic-strip-ai/train/04439c_0.tif')\n\nregion = (0, 0) # location of the top left pixel\nlevel = 0 # level of the picture (we have only 0)\nsize = (10000, 10000) # region size in pixels\n\nslide_thumb_600 = example.get_thumbnail(size=(600, 600))\nplt.figure(figsize=(20, 20))\nplt.imshow(slide_thumb_600)\nplt.show();\n\n\nregion = example.read_region(region, level, size)\nplt.figure(figsize=(20, 20))\nplt.imshow(region)\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:09:22.804357Z","iopub.execute_input":"2022-07-08T13:09:22.805466Z","iopub.status.idle":"2022-07-08T13:10:52.401667Z","shell.execute_reply.started":"2022-07-08T13:09:22.805419Z","shell.execute_reply":"2022-07-08T13:10:52.398613Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"slide = openslide.OpenSlide(\"../input/mayo-clinic-strip-ai/train/04439c_0.tif\")\n#Get slide dims at each level. Remember that whole slide images store information\n#as pyramid at various levels\ndims = slide.level_dimensions\n\nnum_levels = len(dims)\nprint(\"Number of levels in this image are:\", num_levels)\n\nprint(\"Dimensions of various levels in this image are:\", dims)\n\n#By how much are levels downsampled from the original image?\nfactors = slide.level_downsamples\nprint(\"Each level is downsampled by an amount of: \", factors)\n\n#Copy an image from a level\nlevel3_dim = dims[0]\n#Give pixel coordinates (top left pixel in the original large image)\n#Also give the level number (for level 3 we are providing a valueof 2)\n#Size of your output image\n#Remember that the output would be a RGBA image (Not, RGB)\nlevel3_img = slide.read_region((0,0), 0, level3_dim) #Pillow object, mode=RGBA\n\n#Convert the image to RGB\nlevel3_img_RGB = level3_img.convert('RGB')\nplt.imshow(level3_img_RGB)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-08T13:10:52.403934Z","iopub.execute_input":"2022-07-08T13:10:52.404638Z","iopub.status.idle":"2022-07-08T13:12:26.019592Z","shell.execute_reply.started":"2022-07-08T13:10:52.404595Z","shell.execute_reply":"2022-07-08T13:12:26.018152Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Return the best level for displaying the given downsample.\nSCALE_FACTOR = 32\nbest_level = slide.get_best_level_for_downsample(SCALE_FACTOR)\nprint(best_level)\n#Here it returns the best level to be 2 (third level)\n#If you change the scale factor to 2, it will suggest the best level to be 0 (our 1st level)\n#################################\n\n#Generating tiles for deep learning training or other processing purposes\n#We can use read_region function and slide over the large image to extract tiles\n#but an easier approach would be to use DeepZoom based generator.\n# https://openslide.org/api/python/\n\nfrom openslide.deepzoom import DeepZoomGenerator\n\n#Generate object for tiles using the DeepZoomGenerator\ntiles = DeepZoomGenerator(slide, tile_size=128, overlap=0, limit_bounds=False)\n#Here, we have divided our tif into tiles of size 128 with no overlap. \n\n#The tiles object also contains data at many levels. \n#To check the number of levels\nprint(\"The number of levels in the tiles object are: \", tiles.level_count)\n\nprint(\"The dimensions of data in each level are: \", tiles.level_dimensions)\n\n#Total number of tiles in the tiles object\nprint(\"Total number of tiles = : \", tiles.tile_count)\n\n#How many tiles at a specific level?\nlevel_num = 11\nprint(\"Tiles shape at level \", level_num, \" is: \", tiles.level_tiles[level_num])\nprint(\"This means there are \", tiles.level_tiles[level_num][0]*tiles.level_tiles[level_num][1], \" total tiles in this level\")\n\n#Dimensions of the tile (tile size) for a specific tile from a specific layer\ntile_dims = tiles.get_tile_dimensions(11, (3,4)) #Provide deep zoom level and address (column, row)\n\n\n#Tile count at the highest resolution level (level 16 in our tiles)\ntile_count_in_large_image = tiles.level_tiles[16] #126 x 151 (32001/256 = 126 with no overlap pixels)\n#Check tile size for some random tile\ntile_dims = tiles.get_tile_dimensions(16, (120,140))\n#Last tiles may not have full 256x256 dimensions as our large image is not exactly divisible by 256\ntile_dims = tiles.get_tile_dimensions(16, (125,150))\n\n\nsingle_tile = tiles.get_tile(16, (62, 70)) #Provide deep zoom level and address (column, row)\nsingle_tile_RGB = single_tile.convert('RGB')\nplt.imshow(single_tile_RGB)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:12:26.027851Z","iopub.execute_input":"2022-07-08T13:12:26.028249Z","iopub.status.idle":"2022-07-08T13:12:26.36581Z","shell.execute_reply.started":"2022-07-08T13:12:26.028211Z","shell.execute_reply":"2022-07-08T13:12:26.364571Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class='sub-heading'>Observations:</div>\n\n\n* TIF files very large files similar to google maps where we can see the whole county while we are zooming in we are able to view each tile similar to that we have our train files we need to provide the tiles to our model for training.\n* We have only dimension where we have total around 47K tiles and 17 different levels each levels correspondes to the zoomed version of the over all image.\n* Understanding domain knowledge and processsing TIF files knowledge is very important.","metadata":{}},{"cell_type":"markdown","source":"<div class='sub-heading'>   References:</div>\n\n### Some video reference:\n\n   * https://www.youtube.com/watch?v=PCdLBXb7JO8\n   * https://www.youtube.com/watch?v=JZ9JUU29r8A\n   * https://www.youtube.com/watch?v=7FR1TsKLoDI\n\n   whole slide:\n   * https://www.youtube.com/watch?v=x_9Jrsko-T8\n\n   Pahtology:\n   * https://www.youtube.com/watch?v=gtF82brtP1w","metadata":{}},{"cell_type":"markdown","source":"\n### What's Next\n\nThis is my second encounter with image data and I am also learning everything at the same data. I have learned everything about the visualizations from here:\n\n* https://www.kaggle.com/code/paulorzp/strip-ai-exploratory-data-analysis\n\n### Updates :\n\n* I have shared few domain knowledge videos and explanation about the data\n* I have done univarient and multivarient data analysis on the train column csv and visualizations.\n* I have read the tif image and displayed as per the labels \n\n### This is what you can expect in future versions of this kernel:\n\n* New,unique and improved Visualizations\n* New Insights about the Data\n* Baseline Model\n* Complete Explanation of DEEP LEARNING TECHNIQUES to be used\n\n### END NOTES\n\nThis notebook is work in progress. I am a beginner on Kaggle and this is my second Image/Computer Vision related competition.I am learning with every passing day, I try to do atleast what I am good at i.e EDA and getting Insights from data and then I build around this core by learning from some very great kernels posted on competitions. Since Kaggle provide us so many days for a competition , we can always start from zero and end up learning all of the state of the art techniques just by following along a competition.That is what my strategy is all about. I also want to emphasize the fact that there is no need to be scared of Kaggle competitions, I mean as a rookie we don't have anything to loose , we can try out different things without fear and that's what I am doing and you can do too. I will keep on updating this kernel with my new findings,learning process,explained techniques,etc in order to help everyone who is just beginning\n\n","metadata":{}},{"cell_type":"markdown","source":"<div class='sub-heading'>I hope you Liked my kernel. An upvote is a gesture of appreciation and encouragement that fills me with energy to keep improving my efforts ,be kind to show one 👍😊</div>","metadata":{}},{"cell_type":"markdown","source":"# Stay Tuned!\n\n\n# your comments and suggestions are highly appreciated ","metadata":{}},{"cell_type":"code","source":"","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"source_hidden":true}},"execution_count":null,"outputs":[]}]}