{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":37333,"databundleVersionId":3949526,"sourceType":"competition"}],"dockerImageVersionId":30587,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![mayo_baner](http://www.mf-data.fr/images/mayo_baner.png)\n\nLe but de cette [compétition](https://www.kaggle.com/competitions/mayo-clinic-strip-ai), est de développer un <b>modèle de classification d'images</b> pour différencier les étiologies des caillots sanguins dans les accidents vasculaires cérébraux ischémiques (CE ou LAA) en utilisant des images pathologiques numériques de diapositives entières...\n\n<span style=\"color:red\">**Bien-sûr, si ce Notebook peut vous aider, n'hésitez pas à upvote!**</span>\n\n# <span style=\"color: #186fb4; font-variant:small-caps;\" id=\"sommaire\">Sommaire</span>\n\n1. [Exploration des données](#section_1)  \n2. [Analyse des données d'images](#section_2) \n3. [Visualisation des images](#section_3) ","metadata":{}},{"cell_type":"markdown","source":"# <span style=\"color: #186fb4\" id=\"section_1\">Exploration des données</span>\n\nLes images sont au format TIFF et sont fournies dans les dossiers train, test et other.\n\nDes fichiers CSV (train.csv, test.csv, other.csv) fournissent des annotations pour les images, y compris les informations sur l'étiologie, le centre médical, le patient, etc.\n\nLes images sont étiquetées avec les étiologies CE et LAA, et il existe également des images avec une étiologie inconnue ou différente (dossier other).","metadata":{}},{"cell_type":"code","source":"# Importation des librairies\n\nimport os\nimport sys\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport cv2\n\nfrom glob import glob\nfrom pprint import pprint\nfrom collections import defaultdict\nimport gc\n\n# Chemin des données sur Kaggle\ninput_path = '/kaggle/input/mayo-clinic-strip-ai/'\n\n# Chargement des fichiers CSV\ntrain_df = pd.read_csv(os.path.join(input_path, 'train.csv'))\ntest_df = pd.read_csv(os.path.join(input_path, 'test.csv'))\nother_df = pd.read_csv(os.path.join(input_path, 'other.csv'))\n\n# Affichage des premières lignes des fichiers CSV\nprint(\"***- Train Data -***\")\ntrain_df.head()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-12-05T13:49:48.506835Z","iopub.execute_input":"2023-12-05T13:49:48.507365Z","iopub.status.idle":"2023-12-05T13:49:48.546484Z","shell.execute_reply.started":"2023-12-05T13:49:48.507325Z","shell.execute_reply":"2023-12-05T13:49:48.545401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"\\n***-Test Data -***\")\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:49:48.56837Z","iopub.execute_input":"2023-12-05T13:49:48.569803Z","iopub.status.idle":"2023-12-05T13:49:48.583401Z","shell.execute_reply.started":"2023-12-05T13:49:48.569759Z","shell.execute_reply":"2023-12-05T13:49:48.582146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"\\n***-Other Data -***\")\nother_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:49:48.626145Z","iopub.execute_input":"2023-12-05T13:49:48.627285Z","iopub.status.idle":"2023-12-05T13:49:48.640666Z","shell.execute_reply.started":"2023-12-05T13:49:48.627243Z","shell.execute_reply":"2023-12-05T13:49:48.639412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A présent, nous allons vérifier le <b>taux de patients uniques</b> dans chaque DataSets :\n","metadata":{}},{"cell_type":"code","source":"tx_unique_patients_train = (train_df['patient_id'].nunique()/train_df['patient_id'].count())\ntx_unique_patients_test = (test_df['patient_id'].nunique()/test_df['patient_id'].count())\ntx_unique_patients_other = (other_df['patient_id'].nunique()/other_df['patient_id'].count())\n\nprint(f\"Taux de patients uniques dans 'train' set: {round(tx_unique_patients_train*100,2)} % soit {train_df['patient_id'].nunique()} patients\")\nprint(f\"Taux de patients uniques dans 'test' set: {round(tx_unique_patients_test*100,2)} % soit {test_df['patient_id'].nunique()} patients\")\nprint(f\"Taux de patients uniques dans 'other'set: {round(tx_unique_patients_other*100,2)} % soit {other_df['patient_id'].nunique()} patients\")","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:49:48.65545Z","iopub.execute_input":"2023-12-05T13:49:48.655835Z","iopub.status.idle":"2023-12-05T13:49:48.667538Z","shell.execute_reply.started":"2023-12-05T13:49:48.655804Z","shell.execute_reply":"2023-12-05T13:49:48.666293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Puis nous vérifions la <b>répartition des étiologis</b> dans le DataSet de train","metadata":{}},{"cell_type":"code","source":"labels_train = train_df.groupby('label')['label'].count()\nlabels_other = other_df.groupby('label')['label'].count()\npatients_train = train_df.groupby('center_id')['patient_id'].count()\n\nsns.set_style(\"whitegrid\")\nfig, ax = plt.subplots(1,2, figsize=(16,6))\ng1 = sns.barplot(x=labels_train.index, y=labels_train.values, ax=ax[0], palette=\"rocket\")\nax[0].set_title(\"Distribution des étiologies Train Set\",fontsize=16)\nax[0].set_ylabel(\"Nombre\")\nax[0].set_xlabel(\"Etiologie\")\ng1.grid(True)\nfor p in g1.patches:\n    x, w, h = p.get_x(), p.get_width(), p.get_height()\n    if h > 0:\n        g1.text(x + w / 2, h, f'{h:.0f}\\n', ha='center', va='center', size=11)\ng2 = sns.barplot(x=labels_other.index, y=labels_other.values, ax=ax[1], palette=\"cubehelix\")\nax[1].set_title(\"Distribution des autres labels du Other Set\",fontsize=16)\nax[1].set_ylabel(\"Nombre\")\nax[1].set_xlabel(\"Etiologie\")\ng2.grid(True)\nfor p in g2.patches:\n    x, w, h = p.get_x(), p.get_width(), p.get_height()\n    if h > 0:\n        g2.text(x + w / 2, h, f'{h:.0f}\\n', ha='center', va='center', size=11)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:49:48.677409Z","iopub.execute_input":"2023-12-05T13:49:48.678148Z","iopub.status.idle":"2023-12-05T13:49:49.219589Z","shell.execute_reply.started":"2023-12-05T13:49:48.678109Z","shell.execute_reply":"2023-12-05T13:49:49.218379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set_style(\"white\")\nplt.figure(figsize=(15,8))\ng = sns.histplot(data=train_df, x=\"center_id\", hue=\"label\", \n                 multiple=\"dodge\", color='label',discrete=True,\n                 kde=False, shrink=0.8, palette=\"rocket\")\ng.set(xlabel='Numéro du centre clinique', ylabel='Compteur Etiologies')\ng.set(xlim=(0, 12), xticks=np.arange(0,12,1))\nfor p in g.patches:\n    x, w, h = p.get_x(), p.get_width(), p.get_height()\n    if h > 0:\n        g.text(x + w / 2, h, f'{h}\\n', ha='center', va='center', size=11)\ng.margins(y=0.07)\ng.grid(True)\ng.set_title('Distribution des labels et patients par centre clinique', fontsize=16)\n\nax2 = g.twinx()\nax2.plot(patients_train.index, patients_train.values, color='green', marker='o', linestyle='dashed', linewidth=2, markersize=8)\nax2.set_ylabel('Nombre de patients', color='green')\nax2.tick_params(axis='y', labelcolor='green')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:49:49.221587Z","iopub.execute_input":"2023-12-05T13:49:49.221929Z","iopub.status.idle":"2023-12-05T13:49:50.153514Z","shell.execute_reply.started":"2023-12-05T13:49:49.2219Z","shell.execute_reply":"2023-12-05T13:49:50.152274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Et vérifier la répartition du <b>nombre d'images par patients<b>","metadata":{}},{"cell_type":"code","source":"img_per_patients = train_df.groupby('image_num')['patient_id'].count()\nimg_per_patients.index = img_per_patients.index + 1\n\nexplode = (0, 0.1, 0.2, 0.3, 0.4)\nplt.figure(figsize=(8, 8))\nplt.pie(img_per_patients, explode=explode, labels=img_per_patients.index, autopct='%1.1f%%', \n        startangle=180, colors=sns.color_palette('Set2'))\nplt.title('Répartition du nombre d\\'images par patients')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:49:50.155317Z","iopub.execute_input":"2023-12-05T13:49:50.156017Z","iopub.status.idle":"2023-12-05T13:49:50.415231Z","shell.execute_reply.started":"2023-12-05T13:49:50.155981Z","shell.execute_reply":"2023-12-05T13:49:50.414271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color: #186fb4\" id=\"section_2\">Analyse des données d'images</span>\n\nLes images de radiographie, ou images médicales en géneral sont des images de tailles importantes avec une forte résolution. <b>OpenSlide</b> est une bibliothèque créée pour lire des images haute résolution utilisées en pathologie numérique. \nIl possède des fonctionnalités pratiques utilisées pour traiter ici ces images lourdes spécifiques.","metadata":{}},{"cell_type":"code","source":"from openslide import OpenSlide","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:49:50.417833Z","iopub.execute_input":"2023-12-05T13:49:50.418951Z","iopub.status.idle":"2023-12-05T13:49:50.423012Z","shell.execute_reply.started":"2023-12-05T13:49:50.418913Z","shell.execute_reply":"2023-12-05T13:49:50.422152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Nous allons commencer par <b>analyser les propriété des images</b> .tif utilisées ici pour avoir une idée de la masse de données à traiter et créer un <b>DataFrame global</b> regroupant ces données d'images + les données patients.","metadata":{}},{"cell_type":"code","source":"train_images = glob(\"/kaggle/input/mayo-clinic-strip-ai/train/*\")\ntest_images = glob(\"/kaggle/input/mayo-clinic-strip-ai/test/*\")\nother_images = glob(\"/kaggle/input/mayo-clinic-strip-ai/other/*\")\nprint(f\"Nombre d'images du Train set: {len(train_images)}\")\nprint(f\"Nombre d'images du Test set: {len(test_images)}\")\nprint(f\"Nombre d'images du Other set: {len(other_images)}\")","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:49:50.424574Z","iopub.execute_input":"2023-12-05T13:49:50.425438Z","iopub.status.idle":"2023-12-05T13:49:50.443391Z","shell.execute_reply.started":"2023-12-05T13:49:50.425406Z","shell.execute_reply":"2023-12-05T13:49:50.441498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_prop = defaultdict(list)\n\nfor i, path in enumerate(train_images):\n    img_path = train_images[i]\n    slide = OpenSlide(img_path)\n    resolution_x = float(slide.properties['tiff.XResolution'])\n    resolution_y = float(slide.properties['tiff.YResolution'])\n    average_resolution = (resolution_x + resolution_y) / 2.0\n    img_prop['image_id'].append(img_path[-12:-4])\n    img_prop['width'].append(slide.dimensions[0])\n    img_prop['height'].append(slide.dimensions[1])\n    img_prop['size'].append(round(os.path.getsize(img_path) / 1e6, 2))\n    img_prop['resolution'].append(average_resolution)\n    img_prop['path'].append(img_path)\n\nimg_data = pd.DataFrame(img_prop)\nimg_data['img_aspect_ratio'] = img_data['width']/img_data['height']\nimg_data.sort_values(by='image_id', inplace=True)\nimg_data.reset_index(inplace=True, drop=True)\n\nimg_data = img_data.merge(train_df, on='image_id')\nimg_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:49:50.445055Z","iopub.execute_input":"2023-12-05T13:49:50.445418Z","iopub.status.idle":"2023-12-05T13:49:56.996202Z","shell.execute_reply.started":"2023-12-05T13:49:50.445386Z","shell.execute_reply":"2023-12-05T13:49:56.995337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Traitement du jeu de test\nimg_prop_test = defaultdict(list)\n\nfor i, path in enumerate(test_images):\n    img_path = test_images[i]\n    slide = OpenSlide(img_path)\n    resolution_x = float(slide.properties['tiff.XResolution'])\n    resolution_y = float(slide.properties['tiff.YResolution'])\n    average_resolution = (resolution_x + resolution_y) / 2.0\n    img_prop_test['image_id'].append(img_path[-12:-4])\n    img_prop_test['width'].append(slide.dimensions[0])\n    img_prop_test['height'].append(slide.dimensions[1])\n    img_prop_test['size'].append(round(os.path.getsize(img_path) / 1e6, 2))\n    img_prop_test['resolution'].append(average_resolution)\n    img_prop_test['path'].append(img_path)\n\nimg_data_test = pd.DataFrame(img_prop_test)\nimg_data_test['img_aspect_ratio'] = img_data_test['width'] / img_data_test['height']\nimg_data_test.sort_values(by='image_id', inplace=True)\nimg_data_test.reset_index(inplace=True, drop=True)\nimg_data_test.head()","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:49:56.997644Z","iopub.execute_input":"2023-12-05T13:49:56.998168Z","iopub.status.idle":"2023-12-05T13:49:57.055303Z","shell.execute_reply.started":"2023-12-05T13:49:56.998137Z","shell.execute_reply":"2023-12-05T13:49:57.053957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns_to_agg = ['width', 'height', 'size', 'img_aspect_ratio','resolution']\nagg_functions = ['count', 'mean','min', 'max']\n\n# Création d'un DataFrame d'agrégation\nagg_img_df = img_data.groupby('label')[columns_to_agg].agg(agg_functions)\nagg_img_df.columns = [f\"{col}_{agg}\" for col in columns_to_agg for agg in agg_functions]\n\nagg_img_df.round(3).T","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:49:57.056983Z","iopub.execute_input":"2023-12-05T13:49:57.057461Z","iopub.status.idle":"2023-12-05T13:49:57.09198Z","shell.execute_reply.started":"2023-12-05T13:49:57.057414Z","shell.execute_reply":"2023-12-05T13:49:57.090806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ces indicateurs nous fournissent des clés importantes pour le traitement des images. Pour les visualiser au mieux, nous allons afficher un graphique sur la taille et le ratio des images :","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 2, figsize=(16, 5))\nsns.histplot(x='size', data=img_data, bins=100, ax=ax[0])\nax[0].set_title(\"Distribution du poids des images\")\nax[0].set_ylabel(\"%\")\n\n# Ajouter une ligne pour size_mean de CE et LAA\nce_mean_size = img_data.loc[img_data['label'] == 'CE', 'size'].mean()\nax[0].axvline(ce_mean_size, color='red', linestyle='dashed', linewidth=2, label='CE Mean Size')\nlaa_mean_size = img_data.loc[img_data['label'] == 'LAA', 'size'].mean()\nax[0].axvline(laa_mean_size, color='blue', linestyle='dashed', linewidth=2, label='LAA Mean Size')\n\n# Graphique pour la distribution du ratio des images\nsns.histplot(x='img_aspect_ratio', data=img_data, bins=100, ax=ax[1])\nax[1].set_title(\"Distribution du ratio des images\")\nax[1].set_ylabel(\"%\")\n\n# Ajouter une ligne pour size_mean de CE et LAA\nce_mean_ratio = img_data.loc[img_data['label'] == 'CE', 'img_aspect_ratio'].mean()\nax[1].axvline(ce_mean_ratio, color='red', linestyle='dashed', linewidth=2, label='CE Mean Ratio')\nlaa_mean_ratio = img_data.loc[img_data['label'] == 'LAA', 'img_aspect_ratio'].mean()\nax[1].axvline(laa_mean_ratio, color='blue', linestyle='dashed', linewidth=2, label='LAA Mean Ratio')\n\nax[0].legend()\nax[1].legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:49:57.093728Z","iopub.execute_input":"2023-12-05T13:49:57.094712Z","iopub.status.idle":"2023-12-05T13:49:58.449558Z","shell.execute_reply.started":"2023-12-05T13:49:57.094668Z","shell.execute_reply":"2023-12-05T13:49:58.448293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color: #186fb4\" id=\"section_3\">Visualisation des images</span>\n\nA présent, nous allons visualiser plusieurs images avec les Etiologies différentes pour avoir un aperçu des pre-traitements à mettre en place sur ces images.","metadata":{}},{"cell_type":"code","source":"from PIL import Image\nimport concurrent.futures\nImage.MAX_IMAGE_PIXELS = None","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:49:58.452873Z","iopub.execute_input":"2023-12-05T13:49:58.453256Z","iopub.status.idle":"2023-12-05T13:49:58.45838Z","shell.execute_reply.started":"2023-12-05T13:49:58.453202Z","shell.execute_reply":"2023-12-05T13:49:58.457321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <em>ETIOLOGIE CE - TRAIN</em>\nAffichage de 5 images aléatoires","metadata":{}},{"cell_type":"code","source":"CE_imgs = img_data.loc[img_data['label']=='CE','path']\nLAA_imgs = img_data.loc[img_data['label']=='LAA','path']\n\ndef minify_img(img_path):\n    img = Image.open(img_path)\n    img.thumbnail((300, 300), Image.Resampling.LANCZOS)\n    return img\n\nwith concurrent.futures.ThreadPoolExecutor() as executor:\n    CE_images = list(executor.map(minify_img, np.random.choice(CE_imgs, 5)))\n\nplt.style.use('default')\nfig, axes = plt.subplots(1, 5, figsize=(16, 16))\nfor ax, img in zip(axes.reshape(-1), CE_images):\n    ax.imshow(img)\n    ax.set_title(\"target: CE\")\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:49:58.460011Z","iopub.execute_input":"2023-12-05T13:49:58.460424Z","iopub.status.idle":"2023-12-05T13:51:06.895481Z","shell.execute_reply.started":"2023-12-05T13:49:58.460388Z","shell.execute_reply":"2023-12-05T13:51:06.894488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <em>ETIOLOGIE LAA - TRAIN</em>\nAffichage de 5 images aléatoires","metadata":{}},{"cell_type":"code","source":"def minify_img(img_path):\n    img = Image.open(img_path)\n    img.thumbnail((300, 300), Image.Resampling.LANCZOS)\n    return img\n\nwith concurrent.futures.ThreadPoolExecutor() as executor:\n    LAA_images = list(executor.map(minify_img, np.random.choice(LAA_imgs, 5)))\n\nplt.style.use('default')\nfig, axes = plt.subplots(1, 5, figsize=(16, 16))\nfor ax, img in zip(axes.reshape(-1), LAA_images):\n    ax.imshow(img)\n    ax.set_title(\"target: LAA\")\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:51:06.896845Z","iopub.execute_input":"2023-12-05T13:51:06.897605Z","iopub.status.idle":"2023-12-05T13:51:42.051163Z","shell.execute_reply.started":"2023-12-05T13:51:06.897569Z","shell.execute_reply":"2023-12-05T13:51:42.050016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <em>IMAGES DE TEST</em>\nAffichage des 4 images","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 4, figsize=(16, 16))\nfor i, ax in enumerate(axes.flatten()):\n    img_path = img_data_test.loc[i, \"path\"]\n    img = Image.open(img_path)\n    img.thumbnail((300, 300), Image.Resampling.LANCZOS)\n    ax.imshow(img)\n    ax.set_title(f\"Image {i + 1}\")\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-12-05T13:51:42.052783Z","iopub.execute_input":"2023-12-05T13:51:42.053188Z","iopub.status.idle":"2023-12-05T13:54:09.667001Z","shell.execute_reply.started":"2023-12-05T13:51:42.053151Z","shell.execute_reply":"2023-12-05T13:54:09.665733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Premières conclusions :\n<ul>\n    <li>Les tailles des images vont des plus petites aux plus haute résolution.</li>\n    <li>Les images ont des ratios assez différents et donc des qualités très variables.</li> \n    <li>Une quantité importante d'images possède un arrière-plan plus ou moins large et de couleurs différentes.</li> \n    <li>Les tissus se présentent généralement sous la forme de plusieurs petits morceaux.</li> \n    <li>Les tissus de sang ont des couleurs différentes.</li>\n</ul>","metadata":{}}]}