{"cells":[{"metadata":{"_uuid":"fc84fd93-736b-49ce-90c9-ae7597574959","_cell_guid":"60f6e4c2-02b4-4fc3-8c1a-3b1d408d2f7a","trusted":true},"cell_type":"code","source":"%%html\n<marquee style='width: 100%; color: red;'><H1>prostate-cancer-grade-assessment</H1></marquee>","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a7926317-6e88-423e-8a95-69676d5fc970","_cell_guid":"d02c153a-ccd9-4170-9e17-5f700ba7e63e","trusted":true},"cell_type":"markdown","source":"# Sommaire\n1. Objectifs\n2. Comprendre la base de données\n   * Comprendre la base de données\n3. Préparation de la base de données\n * Visualisation de données\n * Fixer quelques problèmes dans la base de données\n    * Images sans masque\n    * ISUP = 2 Gleason score = 4 + 3 \n    * remplacer \"négatif\" par \"0+0\"\n    * Quelques problèmes dans la base de données","execution_count":null},{"metadata":{"_uuid":"5854a53e-93fc-421d-9b1d-17cf0a96942c","_cell_guid":"28b619f3-7bb6-42df-a1b7-6f74e0b38a35","trusted":true},"cell_type":"markdown","source":"# 1. Objectifs\nDétecter et classer la gravité du cancer de la prostate sur des images d'échantillons de tissus prostatiques.\n\nEn pratique, les échantillons de tissus sont examinés et notés par les pathologistes selon le système de notation dit de Gleason, qui est ensuite converti en grade ISUP.","execution_count":null},{"metadata":{"_uuid":"496d6537-6a94-46de-908b-e964c8f45c05","_cell_guid":"09157560-043a-4833-bf4b-90b95ee6eb32","trusted":true},"cell_type":"markdown","source":"<img src=\"https://storage.googleapis.com/kaggle-media/competitions/PANDA/Screen%20Shot%202020-04-08%20at%202.03.53%20PM.png\" height=\"100px\">","execution_count":null},{"metadata":{"_uuid":"5a0b9abd-2132-44c9-9ec7-f8c63e8a24c5","_cell_guid":"3e2a287f-5e22-414d-a5dd-8c8508889d77","trusted":true},"cell_type":"markdown","source":"# 2.Comprendre la base de données\n\n\ntrain.csv et test.csv:\n\n* image_id: Code d'identification de l'image.\n\n* data_provider: Le nom de l'institution qui a fourni les données. L'Institut **Karolinska** et le Centre médical universitaire **Radboud** \n\n\n\n*   uniquement dans train.csv\n\n* isup_grade: La gravité du cancer sur une échelle de 0 à 5.\n\n* gleason_score: Un système alternatif d'évaluation de la gravité du cancer avec plus de niveaux que l'échelle ISUP. \n\n* train_images:\n* 10616 images de type .tiff \n  * Karolinska=5455 images\n  * Radboud=5060 images\n* test_images:\n3 images de type .tiff\n\ntrain_label_masks: Segmentation masks showing which parts of the image led to the ISUP grade. Not all training images have label masks, and there may be false positives or false negatives in the label masks for a variety of reasons. These masks are provided to assist with the development of strategies for selecting the most useful subsamples of the images. The mask values depend on the data provider:","execution_count":null},{"metadata":{"_uuid":"10a84676-7471-469b-bfe2-82453e9a8142","_cell_guid":"7cea7dc2-21cd-4fb7-b9f8-dd56eaa53376","trusted":true},"cell_type":"markdown","source":"# 3.Préparation de la base de données","execution_count":null},{"metadata":{"_uuid":"2b058ab1-d122-4774-bca2-a70717e8ea6c","_cell_guid":"c6096c9f-29ed-43a4-a0ca-c79c71477e68","trusted":true},"cell_type":"markdown","source":"## Visualisation de données","execution_count":null},{"metadata":{"_uuid":"d8e22e80-c430-419c-998c-e64565fa9422","_cell_guid":"34a5dd93-f673-45fd-bf29-cc5eb028a4ae","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport matplotlib\nimport seaborn as sns\nimport openslide\nimport os\nimport cv2\nfrom PIL import Image\nfrom sklearn.model_selection import train_test_split\nfrom keras.preprocessing.image import ImageDataGenerator\nfrom keras.applications.vgg16 import VGG16,preprocess_input\nfrom keras.models import Sequential\nfrom keras.layers import Conv2D, MaxPooling2D, Dense, Dropout, Input, Flatten,BatchNormalization,Activation\nfrom keras.layers import GlobalMaxPooling2D,GlobalAveragePooling2D\nfrom keras.models import Model\nfrom keras.optimizers import Adam, SGD, RMSprop\nfrom keras.callbacks import ModelCheckpoint, Callback, EarlyStopping\nfrom keras.callbacks.callbacks import ReduceLROnPlateau\nfrom tensorflow.keras.callbacks import EarlyStopping\nfrom sklearn.metrics import cohen_kappa_score\nimport tensorflow as tf\nfrom keras.callbacks import LearningRateScheduler\nfrom keras.metrics import *\n\ntrain_df = pd.read_csv(\"../input/prostate-cancer-grade-assessment/train.csv\")\nimage_path = \"../input/prostate-cancer-grade-assessment/train_images/\"\nPATH = \"../input/prostate-cancer-grade-assessment/\"\ntrain_df = pd.read_csv(os.path.join(PATH,'train.csv'))\ntest_df =  pd.read_csv(os.path.join(PATH,'test.csv'))\ntrain_img_path = '../input/prostate-cancer-grade-assessment/train_images'\ntrain_read_img= pd.read_csv(PATH+\"train.csv\")\nmasks = '../input/prostate-cancer-grade-assessment/train_label_masks'\nimages_train_list = os.listdir(os.path.join(PATH, 'train_images'))\nmasks_list = os.listdir(os.path.join(PATH, 'train_label_masks'))\nsns.set_style(\"darkgrid\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"DEVICE = \"TPU\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"if DEVICE != \"TPU\":\n    print(\"Using default strategy for CPU and single GPU\")\n    strategy = tf.distribute.get_strategy()\n\nif DEVICE == \"GPU\":\n    print(\"Num GPUs Available: \", len(tf.config.experimental.list_physical_devices('GPU')))\n    \nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)\nAUTO = tf.data.experimental.AUTOTUNE\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"eed03eb4-27e1-42aa-b9f6-5e7e93cf08f1","_cell_guid":"e61b4ef5-26b2-4b29-9bb1-6f604e214554","trusted":true},"cell_type":"code","source":"print(train_df)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7c39acbe-1eaa-4a6a-9af5-7a6ddcf144ea","_cell_guid":"9491ae4e-aa01-4411-8775-b87f3ddd3984","trusted":true},"cell_type":"code","source":"print(test_df)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dfbf294e-eede-41aa-bad8-3eeeb4f577a8","_cell_guid":"2cf4d1c1-134b-4b16-b366-a1b869c2224a","trusted":true},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"31883da2-5cb8-42be-8837-6611ab72cb56","_cell_guid":"f27adfb7-514f-4a4a-9d3b-53d8a5ce837c","trusted":true},"cell_type":"code","source":"test_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d347d94b-5cfb-46ec-af10-3f64bc8cc461","_cell_guid":"cc0e83d5-b6ac-4fb2-b367-cbefc19ef47a","trusted":true},"cell_type":"code","source":"fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(20,5))\nsns.countplot(ax=ax1, x=\"data_provider\", data=train_df)\nax1.set_title(\"distribution de data_provider  dans  training data\")\nsns.countplot(ax=ax2, x=\"data_provider\", data=test_df)\nax2.set_title(\"distribution de data_provider dans test data\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c83e5464-cc80-40d3-8513-34d1dd5b44ba","_cell_guid":"4b9f6180-34bd-4650-9e45-97e58fdc546a","trusted":true},"cell_type":"code","source":"fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(20,5))\nsns.countplot(ax=ax1, x=\"isup_grade\", data=train_df)\nax1.set_title(\"ISUP Grade distribution dans  Data Provider\")\nsns.countplot(ax=ax2, x=\"gleason_score\", data=train_df)\nax2.set_title(\"Gleason_Score distribution dans  Data Provider\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8c188b42-ba94-4fa1-b6f0-9f9ac114bffb","_cell_guid":"9d9cb83d-246f-4f55-8535-2ef5d95f3311","trusted":true},"cell_type":"code","source":"from tqdm import tqdm\n\nimg_dim= []\n\nfor i,row in tqdm(train_df.iterrows()):\n    slide = openslide.OpenSlide(os.path.join(train_img_path, train_df.image_id.iloc[i]+'.tiff'))\n    img_dim.append(slide.dimensions)\n    slide.close()\n    \nwidth = [dimensions[0] for dimensions in img_dim] \nheight = [dimensions[1] for dimensions in img_dim] \n\ntrain_df['width'] = width\ntrain_df['height'] = height","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"671764f9-dbab-45df-bd9c-416046d48636","_cell_guid":"198dcbc6-6deb-4232-8b54-31504f949aab","trusted":true},"cell_type":"code","source":"fig = plt.figure(figsize=(20,5))\nax = sns.scatterplot(x='width', y='height', data=train_df, hue='data_provider', alpha=0.70)\nax.tick_params(labelsize=10)\n\nplt.title('Dimensions des images')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c5bb7f5a-6451-48f0-bd4a-14bd3f05d350","_cell_guid":"ff0284eb-3253-482d-a074-110e5eb0ecb6","trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(1, 2)\nfig.set_size_inches(20, 5)\n\nsns.stripplot(train_df['width'],train_df['data_provider'],ax=ax[0],jitter=True)\nsns.stripplot(train_df['height'],train_df['data_provider'],ax=ax[1],jitter=True)\n\nax[0].tick_params(labelsize=10)\nax[1].tick_params(labelsize=10)\nax[0].tick_params(labelrotation=90)\nax[1].tick_params(labelrotation=90)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c0934906-e512-43c7-9864-89b185ae60ce","_cell_guid":"861b375a-388b-4cf8-b464-ecdc6fb060ae","trusted":true},"cell_type":"code","source":"data_file_masks = pd.Series(masks_list).to_frame()\ndata_file_masks.columns = ['mask_file_name']\ndata_file_masks.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"29d987ba-0ecd-462e-b8de-707944fd8fc3","_cell_guid":"c9a78487-dba3-4fe9-a8c3-c45a68ee4179","trusted":true},"cell_type":"code","source":"data_file_masks['image_id'] =data_file_masks.mask_file_name.apply(lambda x: x.split('_')[0])\ndata_file_masks.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fc9e81b2-f12e-4baf-9786-1abefdbba3b1","_cell_guid":"134464bc-cc99-46a6-85c1-fd6b1caace6b","trusted":true},"cell_type":"code","source":"train_df = pd.merge(train_df, data_file_masks, on='image_id', how='outer')\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6aaabafc-739c-4429-ad61-1955294f2174","_cell_guid":"65e970df-3123-4076-909b-dc8ff33f80fd","trusted":true},"cell_type":"markdown","source":"## Fixer quelques problèmes dans la base de données","execution_count":null},{"metadata":{"_uuid":"11815ae2-1c3d-4b8b-bfd5-205304411079","_cell_guid":"a3cc9f44-13f5-47d7-ac9d-e2db3a900f2f","trusted":true},"cell_type":"markdown","source":"# Images sans masque\nil y a des images sans masque dans la base de ddonnées","execution_count":null},{"metadata":{"_uuid":"784062dc-f571-426f-b207-0c77a3230d55","_cell_guid":"d10c6d24-baf0-43ff-b42e-3e5668343ac0","trusted":true},"cell_type":"code","source":"del data_file_masks\nprint(f\"Il y a {len(train_df[train_df.mask_file_name.isna()])} images sans masque.\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a7f779e8-8526-4af9-b9d2-8c340d6af71a","_cell_guid":"ed96ac2a-27b2-437d-aace-ef7baf143ed9","trusted":true},"cell_type":"code","source":"print(f\"Train data avant la réduction: {len(train_df)}\")\ndf_train_reduction= train_df[~train_df.mask_file_name.isna()]\nprint(f\"Train data après la réduction: {len(df_train_reduction)}\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c7181047-b6af-4806-ad4f-c06a3b5f1133","_cell_guid":"5ee95c78-e1f3-4799-b35b-d2979b476b21","trusted":true},"cell_type":"code","source":"fig,ax=plt.subplots(1,2,figsize=(20,5))\ntrain_df['data_provider'].value_counts().plot.pie(autopct='%1.1f%%',ax=ax[0])\nax[0].set_ylabel('')\ndf_train_reduction['data_provider'].value_counts().plot.pie(autopct='%1.1f%%',ax=ax[1])\nax[1].set_ylabel('')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"601a1ff6-6ca4-4fac-a0f3-2581f6b5b079","_cell_guid":"2e349714-50e1-4d06-8563-57647ca8950f","trusted":true},"cell_type":"markdown","source":"* le test data contient uniquement 3 images , donc  je vais créer un autre fichier new_test.csv  avec les 100 images que j'ai supprimé (images sans masque)","execution_count":null},{"metadata":{"_uuid":"cde152f7-004e-4890-9a48-c8102125dcc1","_cell_guid":"e47a89b8-d55c-4818-a6dc-c706b50ab84e","trusted":true},"cell_type":"code","source":"\"\"\"\nimages_without_masks=train_df[train_df.mask_file_name.isna()]\nwithout_masks=images_without_masks.groupby('image_id').data_provider.unique().to_frame()\nwithout_masks.to_csv(\"new_test.csv\",index=False)\nwithout_masks\n\"\"\"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c067b7e0-3550-4ba0-928a-5b32fe0f6ef2","_cell_guid":"a2f011f2-5619-4963-afbb-e51682097637","trusted":true},"cell_type":"markdown","source":"inspiré de : [Links](https://medium.com/@kvnamipara/a-better-visualisation-of-pie-charts-by-matplotlib-935b7667d77f)","execution_count":null},{"metadata":{"_uuid":"1d0b5a1a-9db9-48a2-9399-30b46d7923ab","_cell_guid":"796d1630-c672-40e1-9bcd-0b483fe245e5","trusted":true},"cell_type":"markdown","source":"* 1. ISUP grade = 0  Gleason score 0+0 or negative.\n* 1. ISUP grade = 1  Gleason score 3+3.\n* 1. ISUP grade = 2  Gleason score 3+4.\n* 1. ISUP grade = 3  Gleason score 4+3.\n* 1. ISUP grade = 4  Gleason score 4+4 (majority), 3+5 or 5+3.\n* 1. ISUP grade = 5  Gleason score 4+5 (majority), 5+4 or 5+5.","execution_count":null},{"metadata":{"_uuid":"5c42e882-b445-4362-8a08-e366b881ec76","_cell_guid":"73f37bc4-4843-4dc9-92fd-a8ccca5e88b1","trusted":true},"cell_type":"code","source":"df_train_reduction.groupby('isup_grade').gleason_score.unique().to_frame()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3601aa0e-6511-4f6b-b8ba-87ff8178f086","_cell_guid":"9863a6ab-c13c-48ba-a8f2-75c07b2c26d2","trusted":true},"cell_type":"markdown","source":"# ISUP = 2 Gleason score = 4 + 3 \n** Il n'y a pas de ISUP = 2 , Gleason score = 4+3 dans le système de notation Gleason + il n'y a qu'une seule image de ce type et elle semble être une erreur, je vais donc la supprimer.**","execution_count":null},{"metadata":{"_uuid":"624075a6-ef37-4f82-9b06-25f2d9d4771a","_cell_guid":"ecf15552-20f4-4b9e-8e26-3de4e414529b","trusted":true},"cell_type":"code","source":"df_train_reduction[(df_train_reduction.isup_grade == 2) & (df_train_reduction.gleason_score == '4+3')].reset_index()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1d683692-be9c-409f-b721-ae36ff1b38d9","_cell_guid":"e01cf395-e1d8-464a-9909-f70c2c11292b","trusted":true},"cell_type":"code","source":"df_train_reduction.reset_index(inplace=True)\ndf_train_reduction = df_train_reduction[df_train_reduction.image_id !='b0a92a74cb53899311acc30b7405e101']","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"292c975d-e085-467c-a34e-0227fffb3e12","_cell_guid":"162a2b78-cd84-4f04-ad87-61a5e53ed566","trusted":true},"cell_type":"code","source":"df_train_reduction[(df_train_reduction.isup_grade == 2) & (df_train_reduction.gleason_score == '4+3')].reset_index()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dfbda04a-6a7d-42da-8087-9088ff992970","_cell_guid":"5f731547-1abb-44a9-a298-405a008bf089","trusted":true},"cell_type":"code","source":"df_train_reduction.groupby('isup_grade').gleason_score.unique().to_frame()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c9ecca22-df38-4e17-8bcb-a311c3a850d3","_cell_guid":"0a7c6dbd-4d8f-401e-9a42-ce6597f71f2d","trusted":true},"cell_type":"code","source":"temp = df_train_reduction.groupby('isup_grade').count()['image_id'].reset_index().sort_values(by='image_id',ascending=False)\ntemp.style.background_gradient(cmap='Purples')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"23bee60e-dbd3-4a74-8c2e-af66c5e4483e","_cell_guid":"6418279c-6dd1-461e-bc29-703b443f47ea","trusted":true},"cell_type":"code","source":"temp = df_train_reduction.groupby('gleason_score').count()['image_id'].reset_index().sort_values(by='image_id',ascending=False)\ntemp.style.background_gradient(cmap='Reds')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8ab5a6dd-dacd-4994-a0f4-7ed7b17abb91","_cell_guid":"e213a378-758d-405c-8661-f0eb818cb6d6","trusted":true},"cell_type":"markdown","source":"# remplacer \"negative\" par \"0+0\"","execution_count":null},{"metadata":{"_uuid":"90164be3-c3c6-4563-91c4-79736766506a","_cell_guid":"f5d16a30-c095-4735-827c-b068863c163f","trusted":true},"cell_type":"code","source":"df_train_reduction[(df_train_reduction.isup_grade == 0) & (df_train_reduction.gleason_score =='negative')].reset_index()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e8e04c3d-f802-4653-9d7b-70c4e8fbd0cb","_cell_guid":"7af08ef0-2394-4ae6-ab1b-8ffc59f89705","trusted":true},"cell_type":"code","source":"sns.set_style(\"darkgrid\")\nfig= plt.subplots(figsize=(20,5))\nsns.countplot(x='gleason_score', hue=\"data_provider\", data=df_train_reduction)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7d76c482-5a44-4ed3-9b78-4f9cc46acb80","_cell_guid":"67414289-068d-4b46-8a18-5417bfbb87b6","trusted":true},"cell_type":"markdown","source":"*   nous pouvons voir que radboud n'a pas de valeurs \"0+0\" alors que karolinska n'a pas de valeurs \"negative\".\n*    conclusion : \"negative\" correspond à la façon dont le radbound représente \"0+0\" (c'est-à-dire l'absence de cancer) ; il serait donc plus logique de remplacer \"negative\" par \"0+0\".","execution_count":null},{"metadata":{"_uuid":"ab58d7de-ec4a-4680-a7c4-dabb55505cf9","_cell_guid":"8cda8234-2b8f-4c01-8464-db9eac3e22de","trusted":true},"cell_type":"code","source":"df_train_reduction[\"gleason_score\"]= df_train_reduction[\"gleason_score\"].replace(\"negative\", \"0+0\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9b981cec-9b44-4bde-a1c9-61d24059d56d","_cell_guid":"bd909551-c1a1-4f2e-8994-4d145d56f71a","trusted":true},"cell_type":"code","source":"df_train_reduction.groupby('isup_grade').gleason_score.unique().to_frame()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0543c469-e456-4c65-a96a-759eaa22c4ce","_cell_guid":"070b14ad-0e34-4a0c-9187-165cfeabd004","trusted":true},"cell_type":"code","source":"temp = df_train_reduction.groupby('gleason_score').count()['image_id'].reset_index().sort_values(by='image_id',ascending=False)\ntemp.style.background_gradient(cmap='Reds')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"97d8f9a0-877f-45b0-9e29-3b28c79b8287","_cell_guid":"797363e9-e421-4a27-929f-71c1bffd3037","trusted":true},"cell_type":"markdown","source":"# Affichage de quelques images","execution_count":null},{"metadata":{"_uuid":"78dc548c-0b82-4031-8153-3bc63a145a82","_cell_guid":"b2bba995-1c81-4b64-b15d-c056588b038d","trusted":true},"cell_type":"code","source":"def show_images(df, read_region=(1780,1950)):\n    \n    data = df\n    f, ax = plt.subplots(3,3, figsize=(20,20))\n    for i,data_row in enumerate(data.iterrows()):\n        image = str(data_row[1][0])+'.tiff'\n        image_path = os.path.join(PATH,\"train_images\",image)\n        image = openslide.OpenSlide(image_path)\n        spacing = 1 / (float(image.properties['tiff.XResolution']) / 10000)\n        patch = image.read_region(read_region, 0, (256, 256))\n        ax[i//3, i%3].imshow(patch) \n        image.close()       \n        ax[i//3, i%3].axis('off')\n        ax[i//3, i%3].set_title('ID: {}\\nSource: {} ISUP: {} Gleason: {}'.format(\n                data_row[1][0], data_row[1][1], data_row[1][2], data_row[1][3]))\n\n    plt.show()\n    \nimages = [\n    '07a7ef0ba3bb0d6564a73f4f3e1c2293',\n    '037504061b9fba71ef6e24c48c6df44d',\n    '035b1edd3d1aeeffc77ce5d248a01a53',\n    '059cbf902c5e42972587c8d17d49efed',\n    '06a0cbd8fd6320ef1aa6f19342af2e68',\n    '06eda4a6faca84e84a781fee2d5f47e1',\n    '0a4b7a7499ed55c71033cefb0765e93d',\n    '0838c82917cd9af681df249264d2769c',\n    '046b35ae95374bfb48cdca8d7c83233f'\n]\ndata_sample = train_df.loc[train_df.image_id.isin(images)]\nshow_images(data_sample)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6df6796a-6835-4de9-87fe-fe2e529a15b8","_cell_guid":"dcb28666-fa8e-44cb-aa13-cd8208e7ecdb","trusted":true},"cell_type":"code","source":"fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(20,5))\nsns.countplot(ax=ax1, x=\"isup_grade\", data=df_train_reduction)\nax1.set_title(\"ISUP Grade Count by Data Provider\")\nsns.countplot(ax=ax2, x=\"gleason_score\", data=df_train_reduction)\nax2.set_title(\"Gleason_Score Count by Data Provider\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a1409558-2654-4d40-9865-bf791905bead","_cell_guid":"db52a2c1-94fe-4d13-b383-1bc7d7a04e91","trusted":true},"cell_type":"markdown","source":"# Affichage de quelques masques pour loacaliser le cancer et comprendre chaque grade de la maladie","execution_count":null},{"metadata":{"_uuid":"949108ce-98ce-4cca-86d9-846f0373c8ef","_cell_guid":"73c77ec6-7d20-4ffd-bc1a-dae9cf0b9e3c","trusted":true},"cell_type":"code","source":"def show_masks(slides): \n    f, ax = plt.subplots(5,3, figsize=(18,22))\n    for i, slide in enumerate(slides):\n        mask = openslide.OpenSlide(os.path.join(mask_dir, f'{slide}_mask.tiff'))\n        mask_data = mask.read_region((0,0), mask.level_count - 1, mask.level_dimensions[-1])\n        cmap = matplotlib.colors.ListedColormap(['black', 'gray', 'green', 'yellow', 'orange', 'red'])\n        ax[i//3, i%3].imshow(np.asarray(mask_data)[:,:,0], cmap=cmap, interpolation='nearest', vmin=0, vmax=5) \n        mask.close()       \n        ax[i//3, i%3].axis('off')    \n        image_id = slide\n        data_provider = data_sample_mask.loc[slide, 'data_provider']\n        isup_grade = data_sample_mask.loc[slide, 'isup_grade']\n        gleason_score = data_sample_mask.loc[slide, 'gleason_score']\n        ax[i//3, i%3].set_title(f\"ID: {image_id}\\nSource: {data_provider} ISUP: {isup_grade} Gleason: {gleason_score}\")\n        f.tight_layout()\n        \n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"45eadf49-44e2-489d-a351-866ec3b43892","_cell_guid":"3f5ccf67-6214-487c-953c-7082a0735946","trusted":true},"cell_type":"code","source":"images_mask  = [\n    '07a7ef0ba3bb0d6564a73f4f3e1c2293',\n    '037504061b9fba71ef6e24c48c6df44d',\n    '035b1edd3d1aeeffc77ce5d248a01a53',\n    '059cbf902c5e42972587c8d17d49efed',\n    '06a0cbd8fd6320ef1aa6f19342af2e68',\n    '06eda4a6faca84e84a781fee2d5f47e1',\n    '0a4b7a7499ed55c71033cefb0765e93d',\n    '0838c82917cd9af681df249264d2769c',\n    '028098c36eb49a8c6aa6e76e365dd055',\n    '0280f8b612771801229e2dde52371141',\n    '028dc05d52d1dd336952a437f2852a0a',\n    '02a2dcd6ad8bc1d9ad7fdc04ffb6dff3',\n    '049031b0ea0dede1ca1e5ca470c1332d',\n    '05f4e9415af9fdabc19109c980daf5ad',\n    '07fd8d4f02f9b95d86da4bc89563e077'\n]\n\nmask_dir = os.path.join(PATH,\"train_label_masks\")\ndata_sample_mask = df_train_reduction.set_index('image_id')\nshow_masks(images_mask)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"820bd10d-b282-416b-a4f4-c1461db241b0","_cell_guid":"9d2f78a5-a625-4237-b7da-586b39c01308","trusted":true},"cell_type":"markdown","source":"# Affichage de quelques images et leurs masques","execution_count":null},{"metadata":{"_uuid":"77b06b2b-7d17-4749-be71-2aa61da9bbbc","_cell_guid":"226a87ab-9451-4c75-8db0-f80595459638","trusted":true},"cell_type":"code","source":"def mask_img(image,max_size=(600,400)):\n    slide = openslide.OpenSlide(os.path.join(train_img_path, f'{image}.tiff'))\n    mask =  openslide.OpenSlide(os.path.join(mask_dir, f'{image}_mask.tiff'))\n    f,ax =  plt.subplots(1,2 ,figsize=(18,22))\n    spacing = 1 / (float(slide.properties['tiff.XResolution']) / 10000)\n    img = slide.get_thumbnail(size=(600,400)) \n    mask_data = mask.read_region((0,0), mask.level_count - 1, mask.level_dimensions[-1])\n    cmap = matplotlib.colors.ListedColormap(['black', 'gray', 'green', 'yellow', 'orange', 'red'])\n    \n    \n    ax[0].imshow(img)\n    ax[1].imshow(np.asarray(mask_data)[:,:,0], cmap=cmap, interpolation='nearest', vmin=0, vmax=5) \n    \n    image_id = image\n    data_provider = data_sample_mask.loc[image, 'data_provider']\n    isup_grade = data_sample_mask.loc[image, 'isup_grade']\n    gleason_score = data_sample_mask.loc[image, 'gleason_score']\n    ax[0].set_title(f\"ID: {image_id}\\nSource: {data_provider} ISUP: {isup_grade} Gleason: {gleason_score} IMAGE\")\n    ax[1].set_title(f\"ID: {image_id}\\nSource: {data_provider} ISUP: {isup_grade} Gleason: {gleason_score} IMAGE_MASK\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6b0947c4-f83a-46e4-b166-dc4a54c61f90","_cell_guid":"882daf6d-b17f-4e1f-bd59-9b7b191f5166","trusted":true},"cell_type":"code","source":"images1= [\n    '08ab45297bfe652cc0397f4b37719ba1',\n    '090a77c517a7a2caa23e443a77a78bc7',\n    '07fd8d4f02f9b95d86da4bc89563e077'\n]\n\nfor image in images1:\n    mask_img(image)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6f916ce1-8671-4f64-8c63-32d14ccd39e0","_cell_guid":"3a706eef-8377-4f71-9c73-8b78687f56e8","trusted":true},"cell_type":"markdown","source":"\npanda-resized-train-data-512x512 , code source : [Links](https://www.kaggle.com/xhlulu/panda-resize-and-save-train-data)","execution_count":null},{"metadata":{"_uuid":"ca644467-25c8-4218-9eb4-ac27b7be3665","_cell_guid":"7023c390-3f0e-4133-b35b-0a92d0430da8","trusted":true},"cell_type":"code","source":"train_df=df_train_reduction\nAccuracies_list=[]\nlabels=[]\ndata=[]\ndata_dir='../input/panda-resized-train-data-512x512/train_images/train_images/'\nfor i in range(train_df.shape[0]):\n    data.append(data_dir + train_df['image_id'].iloc[i]+'.png')\n    labels.append(train_df['isup_grade'].iloc[i])\ndf=pd.DataFrame(data)\ndf.columns=['images']\ndf['isup_grade']=labels","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d28ea76d-0c1d-428d-81aa-3289e45a6c20","_cell_guid":"84c2fd8b-97ab-4c4e-9b15-c97081429c1b","trusted":true},"cell_type":"code","source":"df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c4122294-0102-4217-8a49-3cdc6b86c6bc","_cell_guid":"9ed48424-5a57-4e3f-8175-afc69738cd0d","trusted":true},"cell_type":"code","source":"print(len(df))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9f97c959-e766-4b1d-8807-9343b04b4d8e","_cell_guid":"684d2250-d4be-4ba5-a42a-11c12aa30d9d","trusted":true},"cell_type":"code","source":"print(labels)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3eb63ff5-1481-40c3-8312-6014e2bdde5f","_cell_guid":"47164b19-c4eb-48f8-94d1-c510af4e45bf","trusted":true},"cell_type":"markdown","source":"### diviser notre data set","execution_count":null},{"metadata":{"_uuid":"022b3481-a6d8-4741-a90e-be6d993f552e","_cell_guid":"29370d66-250c-4dc8-a45a-89feb45c31b3","trusted":true},"cell_type":"code","source":"X_train, X_val, y_train, y_val = train_test_split(df['images'],df['isup_grade'], test_size=0.1, random_state=42)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"232d7974-6af6-45b6-96fe-83dda17f70e5","_cell_guid":"5003ce3e-986f-49e0-bace-da583c15ea93","trusted":true},"cell_type":"code","source":"train=pd.DataFrame(X_train)\ntrain.columns=['images']\ntrain['isup_grade']=y_train\n\nvalidation=pd.DataFrame(X_val)\nvalidation.columns=['images']\nvalidation['isup_grade']=y_val\n\ntrain['isup_grade']=train['isup_grade'].astype(str)\nvalidation['isup_grade']=validation['isup_grade'].astype(str)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c08843c3-1ac0-49d5-b84c-f22f8d9e1a59","_cell_guid":"3f2e137b-c040-4db9-919a-25787a50750b","trusted":true},"cell_type":"code","source":"print(\"train size \",len(train))\nprint(\"validation size \",len(validation))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8c83f0b1-6464-4db1-a56c-a42b635f838a","_cell_guid":"f3a53e99-d472-4cbe-813a-b3f8dc72eec6","trusted":true},"cell_type":"code","source":"print(train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a94fea0a-5b35-4a07-9564-946fa5158521","_cell_guid":"ba3864f6-a492-4878-a47b-6a9f3b0fb4ee","trusted":true},"cell_type":"code","source":"print(validation)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"09a269ee-79a2-4b99-a7ea-72a947a2d1e0","_cell_guid":"ea150e5a-16b7-41a5-a703-9185bd4492d1","trusted":true},"cell_type":"markdown","source":"### après le divisiment de notre data set","execution_count":null},{"metadata":{"_uuid":"0a35a989-7f73-40a2-8527-d4778138c5bc","_cell_guid":"6500d7b6-7c57-4ba1-906b-bd2e145511f4","trusted":true},"cell_type":"code","source":"sns.set(style=\"darkgrid\")\na = ['TRAIN DATA ','TEST DATA ']\nb = [len((train)),len((validation))]\nax = sns.barplot(x=a, y=b)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"257269b8-b1af-407b-99bb-62b2b9b2ad4c","_cell_guid":"ec669219-29d5-47a6-bfca-af145be41ba5","trusted":true},"cell_type":"code","source":"fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(20,5))\nsns.countplot(ax=ax1, x=\"isup_grade\", data=train)\nax1.set_title(\"distribution de Grade ISUP dans le TRAIN DATA après le divisiment\")\nsns.countplot(ax=ax2, x=\"isup_grade\", data=validation)\nax2.set_title(\"distribution de Grade ISUP dans le TEST DATA après le divisiment\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f7b33f66-af7a-4220-a4bb-4661b026d98f","_cell_guid":"ccfed954-e94d-4509-9e5b-efd71b892f16","trusted":true},"cell_type":"markdown","source":"### data  augmentation","execution_count":null},{"metadata":{"_uuid":"e096d52f-6b2d-44c2-872c-6c592d969c5f","_cell_guid":"86b93286-c3df-4cc2-ade7-611cd4733a26","trusted":true},"cell_type":"code","source":"train_datagen = ImageDataGenerator(rescale=1./255,rotation_range=45,\n    featurewise_center=True,\n    featurewise_std_normalization=True,\n    zoom_range=[0.8, 1.2],        \n    horizontal_flip=True, vertical_flip = True,\n    brightness_range=[0.9, 1.1],\n    width_shift_range=1.0,\n    height_shift_range=1.0)\n\nval_datagen=train_datagen = ImageDataGenerator(rescale=1./255)\ntrain_generator = train_datagen.flow_from_dataframe(\n    dataframe=train,\n    x_col='images',\n    y_col='isup_grade',\n    target_size=(224, 224),\n    batch_size=32,\n    seed=2020,\n    shuffle = True,\n    class_mode='categorical')\n\nvalidation_generator = val_datagen.flow_from_dataframe(\n    dataframe=validation,\n    x_col='images',\n    y_col='isup_grade',\n    target_size=(224, 224),\n    batch_size=32,\n    seed=2020,\n    class_mode='categorical')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3bfab2dc-ad7b-4e9d-b704-70091873c9e2","_cell_guid":"de13d525-957c-4a46-9730-9d0006114971","trusted":true},"cell_type":"markdown","source":"#### METRIC","execution_count":null},{"metadata":{"_uuid":"0ddb3288-5c6a-4b60-8e5b-0970c3f8c928","_cell_guid":"372aa6a1-08e1-4ea9-aaf2-36cffd1a8c1e","trusted":true},"cell_type":"code","source":"METRICS = [\n      TruePositives(name='tp'),\n      FalsePositives(name='fp'),\n      TrueNegatives(name='tn'),\n      FalseNegatives(name='fn'), \n      BinaryAccuracy(name='accuracy'),\n      Precision(name='precision'),\n      Recall(name='recall'),\n      AUC(name='auc'),\n]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7390a2b3-3bc2-4ffe-a1f8-9e9ae7d0198a","_cell_guid":"36041f13-1555-4b04-82de-f04eabbdc551","trusted":true},"cell_type":"code","source":"\"\"\"\n%load_ext tensorboard\nlogdir = \"logs/scalars/\"\n\"\"\"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7acc8e3b-56e3-49bc-9737-ab7ac05e7642","_cell_guid":"0a8389b1-ee6e-4af1-9c6f-d1435f01b77a","trusted":true},"cell_type":"code","source":"\"\"\"\nimport keras\ndef lr_schedule(epoch):\n  \n  learning_rate = 0.2\n  if epoch > 10:\n    learning_rate = 0.02\n  if epoch > 20:\n    learning_rate = 0.001\n  if epoch > 50:\n    learning_rate = 0.0005\n\n  tf.summary.scalar('learning rate', data=learning_rate, step=epoch)\n  return learning_rate\n\n\nlr_callback = keras.callbacks.LearningRateScheduler(lr_schedule)\ntensorboard_callback = keras.callbacks.TensorBoard(log_dir=logdir)\n\"\"\"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cd14dab2-b1bd-4abb-a94e-0acbbbb2adab","_cell_guid":"0f65758e-38c9-43e1-9f43-a82659476ca5","trusted":true},"cell_type":"code","source":"\"\"\"\ncallbacks = [ReduceLROnPlateau(monitor='val_loss', patience=1, verbose=1, factor=0.5),\n             EarlyStopping(monitor='val_loss', patience=3),\n             ModelCheckpoint(filepath='best_model.h5', monitor='val_loss', save_best_only=True)]\n\"\"\"\n\n\n#earlyStopping = EarlyStopping(monitor='val_loss', patience=2, verbose=1, mode='auto')\n#mcp_save = ModelCheckpoint(filepath='best_model.h5', save_best_only=True, monitor='val_loss', mode='auto')\n#reduce_lr_loss = ReduceLROnPlateau(monitor='val_loss', factor=0.1, patience=2, verbose=1, epsilon=1e-4, mode='auto')\n\n\nearlyStopping =tf.keras.callbacks.EarlyStopping(monitor='val_loss', min_delta=0, patience=3, verbose=1, mode='auto')\nmcp_save  = tf.keras.callbacks.ModelCheckpoint(filepath='best_model.h5',save_best_only=True,monitor='val_loss', mode='auto')\n#reduce_lr_loss = tf.keras.callbacks.ReduceLROnPlateau( monitor='val_loss', factor=0.1, patience=2, verbose=1, mode='auto', min_delta=0.0001,epsilon=1e-4)\n\n\n\ndef lrfn(epoch):\n    LR_START          = 0.000005\n    LR_MAX            = 0.000020 * strategy.num_replicas_in_sync\n    LR_MIN            = 0.000001\n    LR_RAMPUP_EPOCHS = 5\n    LR_SUSTAIN_EPOCHS = 0\n    LR_EXP_DECAY = .8\n    \n    if epoch < LR_RAMPUP_EPOCHS:\n        lr = (LR_MAX - LR_START) / LR_RAMPUP_EPOCHS * epoch + LR_START\n    elif epoch < LR_RAMPUP_EPOCHS + LR_SUSTAIN_EPOCHS:\n        lr = LR_MAX\n    else:\n        lr = (LR_MAX - LR_MIN) * LR_EXP_DECAY**(epoch - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS) + LR_MIN\n    return lr","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"lr_schedule = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose=1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3ec536a6-dc74-4a52-a1f6-8570996476f6","_cell_guid":"bb5f3735-94e7-4e7c-8ab5-9194c1e0815b","trusted":true},"cell_type":"code","source":"def vgg16_model( num_classes=None):\n\n    model = VGG16(weights='../input/keras-pretrained-models/vgg16_weights_tf_dim_ordering_tf_kernels_notop.h5', include_top=False, input_shape=(224, 224, 3))\n    x=Dropout(0.3)(model.output)\n    x=Flatten()(x)\n    x=Dense(32, activation = 'relu')(x)\n    x=Dropout(0.2)(x)\n    output=Dense(num_classes,activation='softmax')(x)\n    model=Model(model.input,output)\n    return model\nvgg16_conv=vgg16_model(6)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fc401798-0d8c-4272-a18a-5c9a57be7f4d","_cell_guid":"44f66757-fd0a-4fe7-8af4-750aee91fc88","trusted":true},"cell_type":"code","source":"vgg16_conv.summary()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9e8b3852-e062-4daa-bb49-b1f094b71c02","_cell_guid":"43cfc623-297c-49d8-9b86-d5156337f1a8","trusted":true},"cell_type":"code","source":"\"\"\"\ndef kappa_score(y_true, y_pred):\n    \n    y_true=tf.math.argmax(y_true)\n    y_pred=tf.math.argmax(y_pred)\n    return tf.compat.v1.py_func(cohen_kappa_score ,(y_true, y_pred),tf.double)\n\"\"\"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1e689e5b-6eae-4f32-84fd-92dd861ce4de","_cell_guid":"4a925766-1a32-4c9e-a202-059b8fb5f394","trusted":true},"cell_type":"markdown","source":"lr= 0.0005, momentum=0.9,decay=1e-4 ,inspiré de : [Links](https://machinelearningmastery.com/understand-the-dynamics-of-learning-rate-on-deep-learning-neural-networks/)","execution_count":null},{"metadata":{"_uuid":"84bb948b-136f-406a-823d-0863d0db04e4","_cell_guid":"fe3e0dc9-68e3-431a-965e-dbc67add49cb","trusted":true},"cell_type":"code","source":"#opt =SGD(lr= 0.0005, momentum=0.9,decay=1e-4)\nvgg16_conv.compile(optimizer='adam',\n    loss = 'binary_crossentropy',\n    metrics=['accuracy'])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d0e00835-b37d-4609-aafe-0363c7c98d70","_cell_guid":"dff8d6f9-80ea-4c7e-be75-3075f35c105c","trusted":true},"cell_type":"code","source":"nb_epochs =10\nbatch_size=32\nnb_train_steps = train.shape[0]//batch_size\nnb_val_steps=validation.shape[0]//batch_size\nprint(\"Number of training and validation steps: {} and {}\".format(nb_train_steps,nb_val_steps))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"\"\"\"\ndef data_augment(image, label):\n    image = tf.image.random_flip_left_right(image)\n    image = tf.image.random_flip_up_down(image)\n    image = tf.image.random_hue(image, 0.01)\n    image = tf.image.random_saturation(image, 0.7, 1.3)\n    image = tf.image.random_contrast(image, 0.8, 1.2)\n    image = tf.image.random_brightness(image, 0.1)\n    return image, label   \n\ndef get_training_dataset():\n    dataset = load_dataset(TRAINING_FILENAMES, labeled=True)\n    dataset = dataset.map(data_augment, num_parallel_calls=AUTO)\n    dataset = dataset.repeat() \n    dataset = dataset.shuffle(2048)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.prefetch(AUTO) =\n    return dataset\n\ndef get_validation_dataset(ordered=False):\n    dataset = load_dataset(VALIDATION_FILENAMES, labeled=True, ordered=ordered)\n    dataset = dataset.batch(BATCH_SIZE)\n    dataset = dataset.cache()\n    dataset = dataset.prefetch(AUTO)\n    return dataset\n\"\"\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f36c2403-e140-4fb2-b9a5-af397bfff35b","_cell_guid":"df189ed6-f503-4a21-9993-3ff9c2530a8d","trusted":true},"cell_type":"code","source":"vgg16_history=vgg16_conv.fit_generator(train_generator,\n                                       steps_per_epoch=nb_train_steps,\n                                       epochs=nb_epochs,\n                                       validation_data=validation_generator,\n                                       validation_steps=nb_val_steps,\n                                       callbacks=[earlyStopping,mcp_save,lr_schedule])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ae9883a8-f6bb-4233-a15b-b101d0063d0f","_cell_guid":"9b9ef6b9-d348-4883-bd2d-0478492ee2b1","trusted":true},"cell_type":"code","source":"vgg16_conv.save('prostate_cancer_vgg16_model.h5')\nvgg16_weights =vgg16_conv.save_weights('vgg16_weights.h5')\nAccuracies_list.append(['vgg16', vgg16_history])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"93e14fc3-f950-430c-80b4-a3b69c1d2bb6","_cell_guid":"c0f8fcda-a92f-4f3d-890c-10c54564ed8b","trusted":true},"cell_type":"code","source":"def show_history(history):\n    fig, ax = plt.subplots(1, 3, figsize=(20,5))\n    ax[0].set_title('loss')\n    ax[0].plot(history.epoch, history.history[\"loss\"], label=\"Train loss\")\n    ax[0].plot(history.epoch, history.history[\"val_loss\"], label=\"Validation loss\")\n    ax[1].set_title('AUC')\n    ax[1].plot(history.epoch, history.history[\"auc\"], label=\"Train AUC\")\n    ax[1].plot(history.epoch, history.history[\"val_auc\"], label=\"Validation AUC\")\n    ax[2].set_title('Accuracy')\n    ax[2].plot(history.epoch, history.history[\"accuracy\"], label=\"Train accuracy\")\n    ax[2].plot(history.epoch, history.history[\"val_accuracy\"], label=\"Validation accuracy\")\n    ax[0].legend()\n    ax[1].legend()\n    ax[2].legend()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ec76a3f8-c75a-401f-8fb8-3c13c63626fe","_cell_guid":"660d9bda-28b6-48cf-ab79-cb1b38926370","trusted":true},"cell_type":"code","source":"show_history(vgg16_history)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"81c3f224-259c-4e85-8ca4-ef316904f6f2","_cell_guid":"0bf3c99e-2793-4c12-bc16-1604beebf711","trusted":true},"cell_type":"code","source":"from keras.applications.resnet50 import ResNet50\ndef ResNet50_model(num_classes = None):\n    #model = ResNet50(weights='imagenet', include_top = False, input_shape = (224,224,3))\n    model = ResNet50(weights='imagenet', include_top=False, input_shape=(224, 224, 3))\n    #x=Dropout(0.2)(model.output)\n    #x = GlobalAveragePooling2D()(model.output)\n    x=Flatten()(model.output)\n    #x =Dropout(0.2)(x)\n    x =Dense(16, activation = 'relu')(x)\n    x =Dropout(0.2)(x)\n    output=Dense(num_classes,activation='softmax')(x)\n    model=Model(model.input,output)\n    return model\nResNet50_conv = ResNet50_model(6)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"773486c8-27cc-42b1-bb72-26f74cbf809a","_cell_guid":"f1b0d4d1-b081-4609-b0e3-60ae840ff971","trusted":true},"cell_type":"code","source":"ResNet50_conv.summary()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"671bd655-1b95-495a-be23-d24814c4fd92","_cell_guid":"58d1b444-6541-41d9-a658-a513a07fe9e0","trusted":true},"cell_type":"code","source":"ResNet50_conv.compile(loss='binary_crossentropy',optimizer=opt,metrics=METRICS)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"686a97e8-13ae-4a91-9e81-4d71e6e3acad","_cell_guid":"011db8f1-ee0f-44d1-b0de-d9faa4ac45a6","trusted":true},"cell_type":"code","source":"RN_50_history=ResNet50_conv.fit_generator(train_generator,\n                                          steps_per_epoch=nb_train_steps,\n                                          epochs=nb_epochs,\n                                          validation_data=validation_generator,\n                                          validation_steps=nb_val_steps,\n                                          callbacks=[earlyStopping, mcp_save,lr_schedule])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"34c34b78-72d9-4d3e-b517-f5690419af1e","_cell_guid":"3fde5195-83df-4ea6-85ee-632e313701d8","trusted":true},"cell_type":"code","source":"ResNet50_conv.save('prostate_cancer_ResNet50_conv.h5')\nAccuracies_list.append(['ResNet50', RN_50_history])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5f6d6c41-8729-4fef-9e51-864a409ab4b6","_cell_guid":"716926c1-77f2-4aae-a533-b29fe79a9837","trusted":true},"cell_type":"code","source":"def show_history(history):\n    fig, ax = plt.subplots(1, 3, figsize=(20,5))\n    ax[0].set_title('loss')\n    ax[0].plot(history.epoch, history.history[\"loss\"], label=\"Train loss\")\n    ax[0].plot(history.epoch, history.history[\"val_loss\"], label=\"Validation loss\")\n    ax[1].set_title('AUC')\n    ax[1].plot(history.epoch, history.history[\"auc\"], label=\"Train AUC\")\n    ax[1].plot(history.epoch, history.history[\"val_auc\"], label=\"Validation AUC\")\n    ax[2].set_title('Accuracy')\n    ax[2].plot(history.epoch, history.history[\"accuracy\"], label=\"Train accuracy\")\n    ax[2].plot(history.epoch, history.history[\"val_accuracy\"], label=\"Validation accuracy\")\n    ax[0].legend()\n    ax[1].legend()\n    ax[2].legend()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d195af2e-8f2d-4f44-8cba-0110bac6e745","_cell_guid":"015aff3b-4a40-4444-8eae-613001c0dccb","trusted":true},"cell_type":"code","source":"show_history(RN_50_history)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"01a9282c-8298-4da1-9b57-f77348b5b10b","_cell_guid":"fe603f0d-32cc-44e8-9a8f-b00b7dfd910f","trusted":true},"cell_type":"code","source":"from keras.applications.vgg19 import VGG19\ndef vgg19_model(num_classes = None):\n    model = VGG19(weights='imagenet', include_top=False, input_shape=(224, 224, 3))\n    x=Dropout(0.3)(model.output)\n    x=Flatten()(x)\n    x =Dense(32, activation = 'relu')(x)\n    x =Dropout(0.2)(x)\n    output=Dense(num_classes,activation='softmax')(x)\n    model=Model(model.input,output)\n    return model\nvgg19_conv = vgg19_model(6)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5931f29d-02e0-4dc5-8b8e-96614e001dae","_cell_guid":"c28a1551-b32c-4951-b2ce-48f472f18417","trusted":true},"cell_type":"code","source":"vgg19_conv.summary()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c367bed4-7cce-412c-b161-84920a7f2b9a","_cell_guid":"2d5b6ff1-9829-4585-a582-5c550d0d2c22","trusted":true},"cell_type":"code","source":"vgg19_conv.compile(loss='binary_crossentropy',optimizer=opt,metrics=METRICS)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5d8b979b-052e-45bb-933d-cd43cb7d7b79","_cell_guid":"ac03c736-8892-46e3-a72f-779bf16b239a","trusted":true},"cell_type":"code","source":"vgg19_history=vgg19_conv.fit_generator(train_generator,\n                                       steps_per_epoch=nb_train_steps,\n                                       epochs=nb_epochs,\n                                       validation_data=validation_generator,\n                                       validation_steps=nb_val_steps,\n                                       callbacks=[earlyStopping, mcp_save,lr_schedule])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c3bc3f71-116e-4664-ac2b-b7787266e7f7","_cell_guid":"b40b1f29-d1cb-40a0-9ad1-6e3416d4f6c6","trusted":true},"cell_type":"code","source":"vgg19_conv.save('prostate_cancer_vgg19_conv.h5')\nAccuracies_list.append(['vgg19', vgg19_history])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cb6bbf7f-ccda-4f97-90f0-4e0e713c77da","_cell_guid":"114f36e8-4e19-4ab8-b322-880a420aa622","trusted":true},"cell_type":"code","source":"def show_history(history):\n    fig, ax = plt.subplots(1, 3, figsize=(20,5))\n    ax[0].set_title('loss')\n    ax[0].plot(history.epoch, history.history[\"loss\"], label=\"Train loss\")\n    ax[0].plot(history.epoch, history.history[\"val_loss\"], label=\"Validation loss\")\n    ax[1].set_title('AUC')\n    ax[1].plot(history.epoch, history.history[\"auc\"], label=\"Train AUC\")\n    ax[1].plot(history.epoch, history.history[\"val_auc\"], label=\"Validation AUC\")\n    ax[2].set_title('Accuracy')\n    ax[2].plot(history.epoch, history.history[\"accuracy\"], label=\"Train accuracy\")\n    ax[2].plot(history.epoch, history.history[\"val_accuracy\"], label=\"Validation accuracy\")\n    ax[0].legend()\n    ax[1].legend()\n    ax[2].legend()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"12526949-bf01-4da9-8bd1-1e9a48373d20","_cell_guid":"54f66a83-d4ad-4ea5-aee4-e96ead726998","trusted":true},"cell_type":"code","source":"show_history(vgg19_history)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"68450e45-12e1-4aa5-8714-541830d0757d","_cell_guid":"c1f74c22-a3e1-441b-a8f6-328ec1025dda","trusted":true},"cell_type":"code","source":"from keras.applications.inception_v3 import InceptionV3\ndef InceptionV3_model(num_classes = None):\n    InceptionV3_weights = '../input/keras-pretrained-models/inception_v3_weights_tf_dim_ordering_tf_kernels_notop.h5'\n    model = InceptionV3(weights= InceptionV3_weights, include_top=False, input_shape=(224, 224, 3))\n    x=Dropout(0.3)(model.output)\n    x=Flatten()(x)\n    x =Dense(32, activation = 'relu')(x)\n    x =Dropout(0.2)(x)\n    output=Dense(num_classes,activation='softmax')(x)\n    model=Model(model.input,output)\n    return model\nInceptionV3_conv = InceptionV3_model(6)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"78c1168f-bd83-4aff-9d7c-88249d98bc0c","_cell_guid":"06449b01-05b6-4e0d-bd67-4f4698c85003","trusted":true},"cell_type":"code","source":"InceptionV3_conv.summary()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c45b4474-f608-4d93-bb54-7d3a4d32edaf","_cell_guid":"cae28a23-7472-490f-bf76-0afb08410002","trusted":true},"cell_type":"code","source":"InceptionV3_conv.compile(loss='binary_crossentropy',optimizer=opt,metrics=METRICS)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"48ed1493-2b01-4229-9eef-045ae4824448","_cell_guid":"9b310831-e300-4129-9157-56f5f65b9057","trusted":true},"cell_type":"code","source":"InceptionV3_history=InceptionV3_conv.fit_generator( train_generator,\n                                           steps_per_epoch=nb_train_steps,\n                                           epochs=nb_epochs,\n                                           validation_data=validation_generator,\n                                           validation_steps=nb_val_steps,\n                                           callbacks=[earlyStopping, mcp_save,lr_schedule])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"58b98ab3-07f9-41cb-b5ae-2c0d671e4f8f","_cell_guid":"370b0af3-abf3-4dc4-a22a-7206df478fd5","trusted":true},"cell_type":"code","source":"InceptionV3_conv.save('prostate_cancer_vgg19_conv.h5')\nAccuracies_list.append(['InceptionV3',InceptionV3_history])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a1298108-d874-4e77-83fb-eb9049044c2b","_cell_guid":"b4a466ab-5b4f-42ee-a07f-fe9f4247b60f","trusted":true},"cell_type":"code","source":"def show_history(history):\n    fig, ax = plt.subplots(1, 3, figsize=(20,5))\n    ax[0].set_title('loss')\n    ax[0].plot(history.epoch, history.history[\"loss\"], label=\"Train loss\")\n    ax[0].plot(history.epoch, history.history[\"val_loss\"], label=\"Validation loss\")\n    ax[1].set_title('AUC')\n    ax[1].plot(history.epoch, history.history[\"auc\"], label=\"Train AUC\")\n    ax[1].plot(history.epoch, history.history[\"val_auc\"], label=\"Validation AUC\")\n    ax[2].set_title('Accuracy')\n    ax[2].plot(history.epoch, history.history[\"accuracy\"], label=\"Train accuracy\")\n    ax[2].plot(history.epoch, history.history[\"val_accuracy\"], label=\"Validation accuracy\")\n    ax[0].legend()\n    ax[1].legend()\n    ax[2].legend()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"861980b6-65d9-49b9-8b2a-c98ad8c893bf","_cell_guid":"c08a208a-ff65-4480-b777-ffd5cd1523b5","trusted":true},"cell_type":"code","source":"show_history(InceptionV3_history)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ec4616a6-9641-4728-922b-99270fa421b6","_cell_guid":"f60b88e3-3b57-4452-9a65-4e9a5bd5587b","trusted":true},"cell_type":"code","source":"Accuracies_list = np.array(Accuracies_list)\nmodel_names = Accuracies_list[:, 0]\nhistories = Accuracies_list[:, 1]\n\nfig, ax = plt.subplots(2, 2, figsize=(20, 20))\nsns.barplot(x=model_names, y=list(map(lambda x: x.history.get('auc')[-1], histories)), ax=ax[0, 0], palette='Spectral')\nsns.barplot(x=model_names, y=list(map(lambda x: x.history.get('val_auc')[-1], histories)), ax=ax[0, 1], palette='gist_yarg')\nsns.barplot(x=model_names, y=list(map(lambda x: x.history.get('accuracy')[-1], histories)), ax=ax[1, 0], palette='rocket')\nsns.barplot(x=model_names, y=list(map(lambda x: x.history.get('val_accuracy')[-1], histories)), ax=ax[1, 1], palette='ocean_r')\nax[0, 0].set_title('Model Training AUC scores')\nax[0, 1].set_title('Model Validation AUC scores')\nax[1, 0].set_title('Model Training Accuracies')\nax[1, 1].set_title('Model Validation Accuracies')\nfig.suptitle('Model Comparisions')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0dc4c554-c838-4d70-84ab-54bdbb834f0f","_cell_guid":"28fb495c-ada6-420d-9f3c-dc1c6eb3477d","trusted":true},"cell_type":"code","source":"metric_dataframe = pd.DataFrame({\n    'Model Names': model_names,\n    'True Positives': list(map(lambda x: x.history.get('tp')[-1], histories)),\n    'False Positives': list(map(lambda x: x.history.get('fp')[-1], histories)),\n    'True Negatives': list(map(lambda x: x.history.get('tn')[-1], histories)),\n    'False Negatives': list(map(lambda x: x.history.get('fn')[-1], histories))\n})\nfig, ax = plt.subplots(2, 2, figsize=(20, 20))\nsns.barplot(x='Model Names', y='True Positives', data=metric_dataframe, ax=ax[0, 0], palette='BrBG')\nsns.barplot(x='Model Names', y='False Positives', data=metric_dataframe, ax=ax[0, 1], palette='icefire_r')\nsns.barplot(x='Model Names', y='True Negatives', data=metric_dataframe, ax=ax[1, 0], palette='PuBu_r')\nsns.barplot(x='Model Names', y='False Negatives', data=metric_dataframe, ax=ax[1, 1], palette='YlOrBr')\nax[0, 0].set_title('True Positives of Models')\nax[0, 1].set_title('False Positives of Models')\nax[1, 0].set_title('True Negatives of Models')\nax[1, 1].set_title('False Negatives of Models')\nfig.suptitle('Confusion Matrix comparision of Models', size=16)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ba82bbe1-04bc-4ee2-9d6e-344319a95fc5","_cell_guid":"735cb8d0-0fd1-4777-b4af-caaa97d7a207","trusted":true},"cell_type":"code","source":"vgg16_conv.load_weights(\"best_model.h5\")\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import skimage.io\ndef predict_isup_grade(df, path):\n    \n    df[\"image_path\"] = [path+image_id+\".tiff\" for image_id in df[\"image_id\"]]\n    df[\"isup_grade\"] = 0\n    predictions = []\n    for idx, row in df.iterrows():\n        print(row.image_path)\n        img=skimage.io.imread(str(row.image_path))\n        img = cv2.resize(img, (224,224))\n        img = cv2.resize(img, (224,224))\n        img = img.astype(np.float32)/255.\n        img=np.reshape(img,(1,224,224,3))\n        prediction=vgg16_conv.predict(img)\n        predictions.append(np.argmax(prediction))\n            \n    df[\"isup_grade\"] = predictions\n    df = df.drop('image_path', 1)\n    return df[[\"image_id\",\"isup_grade\"]]\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"training_df_val = pd.read_csv(\"../input/prostate-cancer-grade-assessment/train.csv\")[:20]\npredict_isup_grade(training_df_val, image_path)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"47663c49-fdd9-46e7-98b1-0b0dbc6d172c","_cell_guid":"c59cf687-b339-490f-a744-ca259900be34","trusted":true},"cell_type":"code","source":"\"\"\"\ntest_path = \"../input/prostate-cancer-grade-assessment/test_images/\"\ntest_df = pd.read_csv(\"../input/prostate-cancer-grade-assessment/test.csv\")[:20]\npredict_isup_grade(test_df, test_path, passes=5)\npredict_isup_grade.head()\n\"\"\"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"466723e6-46e9-407d-8edd-b1588022a49c","_cell_guid":"60c1ed8f-736c-4552-80f1-2746a80dca46","trusted":true},"cell_type":"markdown","source":"## prochain travail\n1. créer une nouvelle data-set a partir des images supprimées pour le test après ,car il y a juste le test.csv , les images pour le test n'existe pas\n1. essayer de nouvelle techniques ,il est possible que j'utilise pytorch","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}