{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport os\nimport pydicom as dcm\nimport cv2\nimport tensorflow as tf\nimport keras\nfrom tensorflow.keras.applications.resnet50 import ResNet50\nfrom tensorflow.keras.layers import Dense, GlobalAveragePooling2D\nfrom keras.models import Model\nfrom keras.losses import BinaryCrossentropy\n\ndcm.config.WARN = 0\n\nTARGET_PATH = \"/kaggle/input/rsna-intracranial-hemorrhage-detection/rsna-intracranial-hemorrhage-detection/stage_2_train.csv\"\nTRAIN_IMG_DIR = \"/kaggle/input/rsna-intracranial-hemorrhage-detection/rsna-intracranial-hemorrhage-detection/stage_2_train/\"\nTEST_IMG_DIR = \"/kaggle/input/rsna-intracranial-hemorrhage-detection/rsna-intracranial-hemorrhage-detection/stage_2_test/\"\n\n!python --version","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:30:08.655921Z","iopub.execute_input":"2023-05-31T05:30:08.656843Z","iopub.status.idle":"2023-05-31T05:30:19.369978Z","shell.execute_reply.started":"2023-05-31T05:30:08.656768Z","shell.execute_reply":"2023-05-31T05:30:19.368511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### *Hemorrhage Types*\n> Hemorrhage in the head (intracranial hemorrhage) is a relatively common condition that has many causes ranging from trauma, stroke, aneurysm, vascular malformations, high blood pressure, illicit drugs and blood clotting disorders. The neurologic consequences also vary extensively depending upon the size, type of hemorrhage and location ranging from headache to death. The role of the Radiologist is to detect the hemorrhage, characterize the hemorrhage subtype, its size and to determine if the hemorrhage might be jeopardizing critical areas of the brain that might require immediate surgery.\n>\n> While all acute (i.e. new) hemorrhages appear dense (i.e. white) on computed tomography (CT), the primary imaging features that help Radiologists determine the subtype of hemorrhage are the location, shape and proximity to other structures (see table).\n>\n> Intraparenchymal hemorrhage is blood that is located completely within the brain itself; intraventricular or subarachnoid hemorrhage is blood that has leaked into the spaces of the brain that normally contain cerebrospinal fluid (the ventricles or subarachnoid cisterns). Extra-axial hemorrhages are blood that collects in the tissue coverings that surround the brain (e.g. subdural or epidural subtypes). ee figure.) Patients may exhibit more than one type of cerebral hemorrhage, which c may appear on the same image. While small hemorrhages are less morbid than large hemorrhages typically, even a small hemorrhage can lead to death because it is an indicator of another type of serious abnormality (e.g. cerebral aneurysm). \n>\n> ![image.png](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F603584%2F56162e47358efd77010336a373beb0d2%2Fsubtypes-of-hemorrhage.png?generation=1568657910458946&alt=media)\n>\n> ![image.png](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F603584%2Fda30220341a8c77a9925868023698b8f%2FMeninges-en.png?generation=1566848157664036&alt=media)\n>\n> -- <cite>[RSNA Intracranial Hemorrhage Detection | Kaggle](https://www.kaggle.com/competitions/rsna-intracranial-hemorrhage-detection/overview/hemorrhage-types)</cite>\n\n### Objective\nOur goal is to check if there is any hemorrage and classify its sub-type","metadata":{}},{"cell_type":"markdown","source":"# Load Data","metadata":{}},{"cell_type":"code","source":"y_train = pd.read_csv(TARGET_PATH)\ny_train","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:30:19.373469Z","iopub.execute_input":"2023-05-31T05:30:19.37473Z","iopub.status.idle":"2023-05-31T05:30:23.819519Z","shell.execute_reply.started":"2023-05-31T05:30:19.374691Z","shell.execute_reply":"2023-05-31T05:30:23.818439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"⇒ For each **patient_id**, e.g. *ID_12cadc6af*, there is a **sub-type**, e.g. *epidural* and the **Label** indicates the diagnosis *True/False*","metadata":{}},{"cell_type":"code","source":"id_train = y_train.ID.str.rsplit(\"_\", n=1, expand=True)\ny_train = pd.concat([id_train, y_train.Label], axis=1)\ny_train","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:30:23.821356Z","iopub.execute_input":"2023-05-31T05:30:23.821798Z","iopub.status.idle":"2023-05-31T05:30:36.668072Z","shell.execute_reply.started":"2023-05-31T05:30:23.821753Z","shell.execute_reply":"2023-05-31T05:30:36.666972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train = y_train.rename(columns={0: \"id\", 1: \"sub_type\", \"Label\": \"label\"})\ny_train","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:30:36.672528Z","iopub.execute_input":"2023-05-31T05:30:36.676941Z","iopub.status.idle":"2023-05-31T05:30:36.949205Z","shell.execute_reply.started":"2023-05-31T05:30:36.676895Z","shell.execute_reply":"2023-05-31T05:30:36.947579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_img_filenames = os.listdir(TRAIN_IMG_DIR)","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:30:36.955581Z","iopub.execute_input":"2023-05-31T05:30:36.958363Z","iopub.status.idle":"2023-05-31T05:31:11.949372Z","shell.execute_reply.started":"2023-05-31T05:30:36.958318Z","shell.execute_reply":"2023-05-31T05:31:11.948182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Target exploration ","metadata":{}},{"cell_type":"markdown","source":"## Data quality","metadata":{}},{"cell_type":"code","source":"y_train.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:11.966996Z","iopub.execute_input":"2023-05-31T05:31:11.967854Z","iopub.status.idle":"2023-05-31T05:31:12.793939Z","shell.execute_reply.started":"2023-05-31T05:31:11.967814Z","shell.execute_reply":"2023-05-31T05:31:12.792921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train.id.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:12.795609Z","iopub.execute_input":"2023-05-31T05:31:12.796733Z","iopub.status.idle":"2023-05-31T05:31:14.583004Z","shell.execute_reply.started":"2023-05-31T05:31:12.79669Z","shell.execute_reply":"2023-05-31T05:31:14.581845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"4 same IDs ⇒ The 4*6 rows could be duplicated","metadata":{}},{"cell_type":"code","source":"len(train_img_filenames)","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:14.584722Z","iopub.execute_input":"2023-05-31T05:31:14.585145Z","iopub.status.idle":"2023-05-31T05:31:14.594078Z","shell.execute_reply.started":"2023-05-31T05:31:14.585105Z","shell.execute_reply":"2023-05-31T05:31:14.592677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train = y_train.drop_duplicates()  # drop all duplicated rows\ny_train.id.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:14.596224Z","iopub.execute_input":"2023-05-31T05:31:14.596607Z","iopub.status.idle":"2023-05-31T05:31:18.420563Z","shell.execute_reply.started":"2023-05-31T05:31:14.596567Z","shell.execute_reply":"2023-05-31T05:31:18.419349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observation\n⇒ 0 null value \\\n⇒ 752 803 patients \\\n⇒ 4 same IDs as expected\\\n⇒ 0 missing line\n\n**List of sub-types :**\n* Epidural\n* Intraparenchymal\n* Intraventricular\n* Subarachnoid\n* Subdural\n* Any = 1 if any sub-type is 1","metadata":{}},{"cell_type":"markdown","source":"## Distribution","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(1, 2, figsize=(10, 5))\n\nany_label = y_train.loc[y_train[\"sub_type\"] == \"any\", \"label\"]\nsns.countplot(x=any_label, ax=axs[0])\naxs[0].set_title(\"Patients mismatch\")\naxs[0].set_xlabel('patient positivity')\n\nsns.countplot(x=y_train.label, ax=axs[1])\naxs[1].set_title(\"All labels distribution\")","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:18.422385Z","iopub.execute_input":"2023-05-31T05:31:18.422825Z","iopub.status.idle":"2023-05-31T05:31:19.835899Z","shell.execute_reply.started":"2023-05-31T05:31:18.422766Z","shell.execute_reply":"2023-05-31T05:31:19.834792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sum_by_sub_type = y_train.groupby(\"sub_type\").sum()\nsns.barplot(x=sum_by_sub_type.index, y=sum_by_sub_type.label)\nplt.title(\"Sub-types balance\")\nplt.xticks(rotation=30)\n\nsub_types = sum_by_sub_type.index","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:19.837545Z","iopub.execute_input":"2023-05-31T05:31:19.837952Z","iopub.status.idle":"2023-05-31T05:31:20.928845Z","shell.execute_reply.started":"2023-05-31T05:31:19.83791Z","shell.execute_reply":"2023-05-31T05:31:20.927767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sum_by_id = y_train.groupby(\"id\").sum()\nsns.countplot(x=sum_by_id.label, palette=\"flare\")\nplt.title('Nb of positive labels / patient (\"any\" included)')","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:20.930883Z","iopub.execute_input":"2023-05-31T05:31:20.931664Z","iopub.status.idle":"2023-05-31T05:31:24.907384Z","shell.execute_reply.started":"2023-05-31T05:31:20.931621Z","shell.execute_reply":"2023-05-31T05:31:24.90622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Observation\n⇒ **High mismatch** \\\n⇒ 14% of positive patients i.e. ~108 000 \\\n⇒ 5,5% of positive labels i.e. ~256 000 \\\n⇒ epidural is the rarest case ~3 000 (3% of IH)","metadata":{}},{"cell_type":"markdown","source":"## Spilt into train, dev and test set","metadata":{}},{"cell_type":"code","source":"split_1 = int(0.96 * len(train_img_filenames))\nsplit_2 = int(0.98 * len(train_img_filenames))\n\ndev_img_filenames = train_img_filenames[split_1:split_2]\ntest_img_filenames = train_img_filenames[split_2:]\ntrain_img_filenames = train_img_filenames[:split_1]","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:24.912194Z","iopub.execute_input":"2023-05-31T05:31:24.914629Z","iopub.status.idle":"2023-05-31T05:31:24.944086Z","shell.execute_reply.started":"2023-05-31T05:31:24.914587Z","shell.execute_reply":"2023-05-31T05:31:24.942072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"split_1 = len(train_img_filenames) * 6\nsplit_2 = len(train_img_filenames) * 6 + len(dev_img_filenames) * 6\n\ny_test = y_train[split_2:]\ny_dev = y_train[split_1:split_2]\ny_train = y_train[:split_1]","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:24.955508Z","iopub.execute_input":"2023-05-31T05:31:24.958332Z","iopub.status.idle":"2023-05-31T05:31:24.968063Z","shell.execute_reply.started":"2023-05-31T05:31:24.958278Z","shell.execute_reply":"2023-05-31T05:31:24.966938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:24.973752Z","iopub.execute_input":"2023-05-31T05:31:24.976279Z","iopub.status.idle":"2023-05-31T05:31:24.99568Z","shell.execute_reply.started":"2023-05-31T05:31:24.976237Z","shell.execute_reply":"2023-05-31T05:31:24.994761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Check distribution of test and dev set respect data distribution","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(1, 2, figsize=(10, 5))\n\nany_label = y_dev.loc[y_dev[\"sub_type\"] == \"any\", \"label\"]\nsns.countplot(x=any_label, ax=axs[0])\naxs[0].set_title(\"Patients mismatch\")\naxs[0].set_xlabel('patient positivity')\n\nsns.countplot(x=y_dev.label, ax=axs[1])\naxs[1].set_title(\"All labels distribution\")","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:24.999984Z","iopub.execute_input":"2023-05-31T05:31:25.002515Z","iopub.status.idle":"2023-05-31T05:31:25.415383Z","shell.execute_reply.started":"2023-05-31T05:31:25.00247Z","shell.execute_reply":"2023-05-31T05:31:25.414202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sum_by_sub_type = y_dev.groupby(\"sub_type\").sum()\nsns.barplot(x=sum_by_sub_type.index, y=sum_by_sub_type.label)\nplt.title(\"Sub-types balance\")\nplt.xticks(rotation=30)\n\nsub_types = sum_by_sub_type.index","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:25.416876Z","iopub.execute_input":"2023-05-31T05:31:25.417966Z","iopub.status.idle":"2023-05-31T05:31:25.692966Z","shell.execute_reply.started":"2023-05-31T05:31:25.41792Z","shell.execute_reply":"2023-05-31T05:31:25.691827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sum_by_id = y_dev.groupby(\"id\").sum()\nsns.countplot(x=sum_by_id.label, palette=\"flare\")\nplt.title('Nb of positive labels / patient (\"any\" included)')","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:25.694824Z","iopub.execute_input":"2023-05-31T05:31:25.695246Z","iopub.status.idle":"2023-05-31T05:31:25.991241Z","shell.execute_reply.started":"2023-05-31T05:31:25.695201Z","shell.execute_reply":"2023-05-31T05:31:25.990168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axs = plt.subplots(1, 2, figsize=(10, 5))\n\nany_label = y_test.loc[y_test[\"sub_type\"] == \"any\", \"label\"]\nsns.countplot(x=any_label, ax=axs[0])\naxs[0].set_title(\"Patients mismatch\")\naxs[0].set_xlabel('patient positivity')\n\nsns.countplot(x=y_test.label, ax=axs[1])\naxs[1].set_title(\"All labels distribution\")","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:25.992938Z","iopub.execute_input":"2023-05-31T05:31:25.993622Z","iopub.status.idle":"2023-05-31T05:31:26.357076Z","shell.execute_reply.started":"2023-05-31T05:31:25.993579Z","shell.execute_reply":"2023-05-31T05:31:26.35582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sum_by_sub_type = y_test.groupby(\"sub_type\").sum()\nsns.barplot(x=sum_by_sub_type.index, y=sum_by_sub_type.label)\nplt.title(\"Sub-types balance\")\nplt.xticks(rotation=30)\n\nsub_types = sum_by_sub_type.index","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:26.358624Z","iopub.execute_input":"2023-05-31T05:31:26.359568Z","iopub.status.idle":"2023-05-31T05:31:26.649885Z","shell.execute_reply.started":"2023-05-31T05:31:26.359523Z","shell.execute_reply":"2023-05-31T05:31:26.648512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sum_by_id = y_test.groupby(\"id\").sum()\nsns.countplot(x=sum_by_id.label, palette=\"flare\")\nplt.title('Nb of positive labels / patient (\"any\" included)')","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:26.651426Z","iopub.execute_input":"2023-05-31T05:31:26.65209Z","iopub.status.idle":"2023-05-31T05:31:26.987858Z","shell.execute_reply.started":"2023-05-31T05:31:26.652034Z","shell.execute_reply":"2023-05-31T05:31:26.986834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualization","metadata":{}},{"cell_type":"code","source":"data = dcm.dcmread(TRAIN_IMG_DIR + train_img_filenames[9])\nprint(data)","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:26.989217Z","iopub.execute_input":"2023-05-31T05:31:26.989897Z","iopub.status.idle":"2023-05-31T05:31:27.012093Z","shell.execute_reply.started":"2023-05-31T05:31:26.989854Z","shell.execute_reply":"2023-05-31T05:31:27.010753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.imshow(data.pixel_array, cmap=plt.cm.bone)","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:27.014787Z","iopub.execute_input":"2023-05-31T05:31:27.015623Z","iopub.status.idle":"2023-05-31T05:31:27.306513Z","shell.execute_reply.started":"2023-05-31T05:31:27.01558Z","shell.execute_reply":"2023-05-31T05:31:27.305499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Utils\nhttps://www.kaggle.com/code/redwankarimsony/ct-scans-dicom-files-windowing-explained\n\nThese files contain a 16-bit relative grayscale image and represents the X-ray absorbance of the tissue. Hence, you can identify the type of the tissue (Air, Fat, Water, Soft tissue, Blood, Bone).\n\nYou need to get the absolute density in [HU level](https://en.wikipedia.org/wiki/Hounsfield_scale) to normalize the dataset. It is obtained by linear transformation with the coefficients present in the file","metadata":{}},{"cell_type":"code","source":"def get_img(path_img, *, windowing=True):\n    def get_first_field(x):\n        return np.int16(x[0]) if isinstance(x, dcm.multival.MultiValue) else np.int16(x)\n\n    def get_windowing_info(data: dcm.dataset.FileDataset):\n        return {\"window_center\": get_first_field(data.WindowCenter),\n                \"window_width\": get_first_field(data.WindowWidth),\n                \"intercept\": data.RescaleIntercept,\n                \"slope\": data.RescaleSlope}\n\n    def window_image(img, info):\n        \"\"\"linear transformation and min/max windowing\"\"\"\n        img = img*info[\"slope\"] + info[\"intercept\"]  # for translation adjustments given in the file\n        img_min = info[\"window_center\"] - info[\"window_width\"]//2, # minimum HU level,\n        img_max = info[\"window_center\"] + info[\"window_width\"]//2  # maximum HU level\n        return np.clip(img, img_min, img_max)\n    \n    data = dcm.dcmread(path_img)\n    image = data.pixel_array\n    \n    if windowing:\n        info = get_windowing_info(data)\n        image = window_image(image, info)\n\n    return image","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:27.307707Z","iopub.execute_input":"2023-05-31T05:31:27.308058Z","iopub.status.idle":"2023-05-31T05:31:27.321319Z","shell.execute_reply.started":"2023-05-31T05:31:27.308022Z","shell.execute_reply":"2023-05-31T05:31:27.320078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img = get_img(TRAIN_IMG_DIR + train_img_filenames[9])\nplt.imshow(img, cmap=plt.cm.bone)\nimg","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:27.323866Z","iopub.execute_input":"2023-05-31T05:31:27.324261Z","iopub.status.idle":"2023-05-31T05:31:27.626092Z","shell.execute_reply.started":"2023-05-31T05:31:27.324219Z","shell.execute_reply":"2023-05-31T05:31:27.624727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img = get_img(TRAIN_IMG_DIR + train_img_filenames[9], windowing=False)\nplt.imshow(img, cmap=plt.cm.bone)\nimg","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:27.627689Z","iopub.execute_input":"2023-05-31T05:31:27.629075Z","iopub.status.idle":"2023-05-31T05:31:27.913778Z","shell.execute_reply.started":"2023-05-31T05:31:27.629032Z","shell.execute_reply":"2023-05-31T05:31:27.912731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_patient_ids(sub_type=None, num=10):\n    if sub_type is not None:\n        # 2 times faster than y_train.loc[(y_train[\"sub_type\"] == sub_type) & (y_train[\"label\"] == 1)]\n        return y_train.query(\"sub_type == @sub_type and label == 1\").iloc[:num, 0]\n    else:\n        return y_train.iloc[:num, 0]","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:27.915521Z","iopub.execute_input":"2023-05-31T05:31:27.916289Z","iopub.status.idle":"2023-05-31T05:31:27.922881Z","shell.execute_reply.started":"2023-05-31T05:31:27.916244Z","shell.execute_reply":"2023-05-31T05:31:27.921508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata = {}\nbins = np.linspace(0.30, 0.70, num=35)\n\nfor sub_type in sub_types:\n    metadata[sub_type] = {\"rows\": [], \"columns\": [], \"pixel_spacing\": []}\n    patient_ids = get_patient_ids(sub_type, 1000)\n    \n    for patient_id in patient_ids:\n        data = dcm.dcmread(TRAIN_IMG_DIR + patient_id + \".dcm\")\n        metadata[sub_type][\"rows\"].append(data.Rows)\n        metadata[sub_type][\"columns\"].append(data.Columns)\n        metadata[sub_type][\"pixel_spacing\"].append(float(data.PixelSpacing[0]))\n        \n    print(sub_type, \":\")\n    print(\"    list of unique row size :\", np.unique(metadata[sub_type][\"rows\"]))\n    print(\"    list of unique column size :\", np.unique(metadata[sub_type][\"columns\"]))\n    plt.hist(metadata[sub_type][\"pixel_spacing\"], bins, alpha=0.5, label=sub_type)\n\nplt.title(\"pixel_spacing distribution\")\nplt.xlabel(\"pixel_spacing\")\nplt.legend(loc=\"upper right\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:31:27.924622Z","iopub.execute_input":"2023-05-31T05:31:27.925405Z","iopub.status.idle":"2023-05-31T05:32:32.308203Z","shell.execute_reply.started":"2023-05-31T05:31:27.92535Z","shell.execute_reply":"2023-05-31T05:32:32.306996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pixels = {}\nplt.figure(figsize=(15,5))\n\nfor sub_type in sub_types:\n    pixels[sub_type] = []\n    patient_ids = get_patient_ids(sub_type, 15)\n    \n    for patient_id in patient_ids:\n        pixels[sub_type].extend(get_img(TRAIN_IMG_DIR + patient_id + \".dcm\").flatten())\n        \n    sns.histplot(pixels[sub_type], label=sub_type)\n\nplt.title(\"Pixels distribution\")\nplt.xlabel(\"Pixels greyness (HU unit)\")\nplt.legend(loc=\"upper right\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:32:32.309774Z","iopub.execute_input":"2023-05-31T05:32:32.31025Z","iopub.status.idle":"2023-05-31T05:33:15.392051Z","shell.execute_reply.started":"2023-05-31T05:32:32.31021Z","shell.execute_reply":"2023-05-31T05:33:15.390798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"TODO : Corélation entre WindowCenter et sub-type","metadata":{}},{"cell_type":"markdown","source":"# Data preparation","metadata":{}},{"cell_type":"code","source":"test_df = y_train.set_index([\"id\", \"sub_type\"]).unstack(level=-1)","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:33:15.393807Z","iopub.execute_input":"2023-05-31T05:33:15.394627Z","iopub.status.idle":"2023-05-31T05:33:21.322746Z","shell.execute_reply.started":"2023-05-31T05:33:15.394575Z","shell.execute_reply":"2023-05-31T05:33:21.32166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def preprocess(path_img, img_size):\n    def resize(img, size):\n        return cv2.resize(img, size[:2])\n    \n    def rescale(img):\n        \"\"\"rescaling to 0-1\"\"\"\n        img = (img + 40) / 160\n        return np.clip(img, 0, 1)\n    \n    def fill_channels(img):\n        return np.stack((img,)*3, axis=-1)  # to 3 channel\n    \n    img = get_img(path_img)\n    img = resize(img, img_size)\n    img = rescale(img)\n    img = fill_channels(img)\n    \n    return img","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:33:21.324346Z","iopub.execute_input":"2023-05-31T05:33:21.325014Z","iopub.status.idle":"2023-05-31T05:33:21.333741Z","shell.execute_reply.started":"2023-05-31T05:33:21.324966Z","shell.execute_reply":"2023-05-31T05:33:21.332657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preprocess(TRAIN_IMG_DIR + train_img_filenames[9], (4,4))","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:33:21.335444Z","iopub.execute_input":"2023-05-31T05:33:21.336148Z","iopub.status.idle":"2023-05-31T05:33:21.383969Z","shell.execute_reply.started":"2023-05-31T05:33:21.335988Z","shell.execute_reply":"2023-05-31T05:33:21.382859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"https://www.tensorflow.org/api_docs/python/tf/keras/utils/Sequence","metadata":{}},{"cell_type":"code","source":"class DataSequence(keras.utils.Sequence):\n\n    def __init__(self, img_dir, ids, labels=None, batch_size=32, img_size=(224, 224, 3),\n                 *args, **kwargs):\n\n        self.img_dir = img_dir\n        self.ids = ids\n        self.labels = labels\n        self.batch_size = batch_size\n        self.img_size = img_size\n\n    def __len__(self):\n        return len(self.ids)//self.batch_size + 1\n\n    def __getitem__(self, index):\n        ids_batch = self.ids[index*self.batch_size:(index+1)*self.batch_size]\n        \n        if self.labels is not None:\n            X = self.generate_img_batch(ids_batch)\n            Y = self.generate_labels_batch(ids_batch)\n            return X, Y\n        else:\n            X = self.generate_img_batch(ids_batch)\n            return X\n    \n    def generate_img_batch(self, ids_batch):\n        X = np.zeros((self.batch_size, *self.img_size))\n        for i, id_p in enumerate(ids_batch):\n            X[i] = preprocess(self.img_dir+id_p+\".dcm\", self.img_size)\n        return X\n    \n    def generate_labels_batch(self, ids_batch):\n        Y = np.zeros((self.batch_size, self.labels.shape[1]), dtype=np.float32)\n        for i, id_p in enumerate(ids_batch):\n            Y[i] = self.labels.loc[id_p].values\n        return Y","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:33:21.385228Z","iopub.execute_input":"2023-05-31T05:33:21.385631Z","iopub.status.idle":"2023-05-31T05:33:21.399327Z","shell.execute_reply.started":"2023-05-31T05:33:21.38558Z","shell.execute_reply":"2023-05-31T05:33:21.398107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"example_gen = DataSequence(TRAIN_IMG_DIR, test_df.index, test_df, batch_size=2)","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:33:21.401025Z","iopub.execute_input":"2023-05-31T05:33:21.401906Z","iopub.status.idle":"2023-05-31T05:33:21.41084Z","shell.execute_reply.started":"2023-05-31T05:33:21.40186Z","shell.execute_reply":"2023-05-31T05:33:21.409659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"example_gen[0]","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:33:21.413464Z","iopub.execute_input":"2023-05-31T05:33:21.413786Z","iopub.status.idle":"2023-05-31T05:33:21.757995Z","shell.execute_reply.started":"2023-05-31T05:33:21.413756Z","shell.execute_reply":"2023-05-31T05:33:21.756832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.imshow(example_gen[0][0][0])","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:33:21.759718Z","iopub.execute_input":"2023-05-31T05:33:21.760456Z","iopub.status.idle":"2023-05-31T05:33:22.048822Z","shell.execute_reply.started":"2023-05-31T05:33:21.760409Z","shell.execute_reply":"2023-05-31T05:33:22.047773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"https://keras.io/api/applications/","metadata":{}},{"cell_type":"code","source":"base_model = ResNet50(include_top=False, weights='imagenet')\n\nfor layer in base_model.layers:\n    layer.trainable = False\n\nx = base_model.output\nx = GlobalAveragePooling2D()(x)\n\npredictions = Dense(6, activation='sigmoid')(x)\n\nmodel = Model(inputs=base_model.input, outputs=predictions)\n\nmodel.compile(optimizer='adam', loss='binary_crossentropy')  # multi-label classifications","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:33:22.051485Z","iopub.execute_input":"2023-05-31T05:33:22.051798Z","iopub.status.idle":"2023-05-31T05:33:29.601272Z","shell.execute_reply.started":"2023-05-31T05:33:22.051768Z","shell.execute_reply":"2023-05-31T05:33:29.600154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_gen = DataSequence(TRAIN_IMG_DIR, test_df.index, test_df, batch_size=32, img_size=(224, 224, 3))\n\nhistory = model.fit(train_gen, steps_per_epoch=32, epochs=5)\n\nmodel.save('/kaggle/working/myModel')","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:33:29.602689Z","iopub.execute_input":"2023-05-31T05:33:29.603092Z","iopub.status.idle":"2023-05-31T05:35:03.113131Z","shell.execute_reply.started":"2023-05-31T05:33:29.603052Z","shell.execute_reply":"2023-05-31T05:35:03.11202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(history.epoch, history.history[\"loss\"])","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:35:03.117282Z","iopub.execute_input":"2023-05-31T05:35:03.117667Z","iopub.status.idle":"2023-05-31T05:35:03.376737Z","shell.execute_reply.started":"2023-05-31T05:35:03.117633Z","shell.execute_reply":"2023-05-31T05:35:03.375717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = y_test.set_index([\"id\", \"sub_type\"]).unstack(level=-1)\ntest_gen = DataSequence(TRAIN_IMG_DIR, test_df.index, test_df, batch_size=32)","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:35:03.37871Z","iopub.execute_input":"2023-05-31T05:35:03.379842Z","iopub.status.idle":"2023-05-31T05:35:03.455624Z","shell.execute_reply.started":"2023-05-31T05:35:03.379779Z","shell.execute_reply":"2023-05-31T05:35:03.454524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Test for 1 batch\ndef cost(test_gen):\n    \n    pred = model.predict(test_gen, batch_size=32, verbose='auto', steps=1)\n    bce = tf.keras.losses.BinaryCrossentropy()\n    cost = bce(y_test.iloc[0:32*6].label.values.reshape((32,6)), pred)\n    \n    return cost.numpy\n","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:35:03.463028Z","iopub.execute_input":"2023-05-31T05:35:03.463366Z","iopub.status.idle":"2023-05-31T05:35:03.470075Z","shell.execute_reply.started":"2023-05-31T05:35:03.463333Z","shell.execute_reply":"2023-05-31T05:35:03.468896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cost(test_gen)","metadata":{"execution":{"iopub.status.busy":"2023-05-31T05:35:03.471644Z","iopub.execute_input":"2023-05-31T05:35:03.472739Z","iopub.status.idle":"2023-05-31T05:35:06.633256Z","shell.execute_reply.started":"2023-05-31T05:35:03.472696Z","shell.execute_reply":"2023-05-31T05:35:06.630971Z"},"trusted":true},"execution_count":null,"outputs":[]}]}