{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\n##### The main focus of this notebook is the exploratory data analysis (EDA) for the RSNA 2023 Abdominal Trauma Detection competition. EDA helps us gain a deep understanding of the dataset's characteristics, such as its size, structure, and distribution of variables, which is crucial for selecting appropriate analytical techniques. EDA also enables us to identify missing values, possible outliers, or inconsistencies in the data, which could adversely affect the accuracy of predictive models. Finally, EDA can aid in the selection of relevant features for model development, potentially improving the model's predictive performance while reducing computational costs. \n\n##### The notebook covers inspection of the folders and files, tabular data and images in DICOM format which is the international standard in medical imaging.","metadata":{}},{"cell_type":"markdown","source":"# Table of Contents\n\n* [1. Importing libraries](#1)\n* [2. Exploring files and folders](#2)\n* [3. Tabular data inspection](#3)\n    * [3.1. Data distribution](#3.1)\n    * [3.2. Multiple injuries](#3.2)\n* [4. Exploring images](#4)\n    * [4.1. Displaying random DICOM image](#4.1)\n    * [4.2. Converting DICOM image to PyTorch tensor](#4.2)","metadata":{}},{"cell_type":"markdown","source":"# 1. Importing libraries <a id=\"1\"></a>\n\nHere we load some standard and third-party libraries that will be helpful to perform EDA.","metadata":{}},{"cell_type":"code","source":"import os\n\nimport pandas as pd\nimport numpy as np\nimport random\nimport matplotlib.pyplot as plt\nimport plotly.express as px\nimport plotly.graph_objects as go\nimport pydicom\nimport torch","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:32:00.839152Z","iopub.execute_input":"2023-08-30T16:32:00.839762Z","iopub.status.idle":"2023-08-30T16:32:05.232523Z","shell.execute_reply.started":"2023-08-30T16:32:00.839728Z","shell.execute_reply":"2023-08-30T16:32:05.231539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Exploring files and folders <a id=\"2\"></a>\n\nLet's define the base directory of our project:","metadata":{}},{"cell_type":"code","source":"BASE_DIR = \"/kaggle/input/rsna-2023-abdominal-trauma-detection\"","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:32:05.234542Z","iopub.execute_input":"2023-08-30T16:32:05.235415Z","iopub.status.idle":"2023-08-30T16:32:05.242966Z","shell.execute_reply.started":"2023-08-30T16:32:05.235373Z","shell.execute_reply":"2023-08-30T16:32:05.241779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can investigate the content of the base directory. Let's define a simple function that will print all the files and folders in a given path. This function could also be useful in the future.","metadata":{}},{"cell_type":"code","source":"def list_files_and_folders(path: str) -> None:\n    \"\"\"\n    List all files and folders in the specified directory path, excluding subdirectories.\n\n    This function prints the names of files and folders directly within the given path.\n    Subdirectories and their contents are not listed.\n\n    Args:\n        path (str): string of the path to exlore.\n    Returns:\n        None: The function prints the names of files and folders in the specified path.\n    \"\"\"\n    if not os.path.exists(path):\n        print(f\"Error: The path {path} does not exist\")\n        return\n    \n    for entry in os.listdir(path):\n        entry_path = os.path.join(path, entry)\n        if os.path.isfile(entry_path):\n            print(f\"File:     {entry}\")\n        elif os.path.isdir(entry_path):\n            print(f\"Folder:   {entry}\")","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:32:05.244601Z","iopub.execute_input":"2023-08-30T16:32:05.245628Z","iopub.status.idle":"2023-08-30T16:32:05.255602Z","shell.execute_reply.started":"2023-08-30T16:32:05.245589Z","shell.execute_reply":"2023-08-30T16:32:05.25471Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list_files_and_folders(BASE_DIR)","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:32:05.258833Z","iopub.execute_input":"2023-08-30T16:32:05.259261Z","iopub.status.idle":"2023-08-30T16:32:05.274639Z","shell.execute_reply.started":"2023-08-30T16:32:05.259231Z","shell.execute_reply":"2023-08-30T16:32:05.273423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Nice! We programmatically fetched all files and folders inside `/kaggle/input/rsna-2023-abdominal-trauma-detection`. We can see the same results by navigating to the *Data->Input* section of the Notebook panel. Also, it was possible to obtain the same outcome using the standard linux `ls` command:","metadata":{}},{"cell_type":"code","source":"!ls -la /kaggle/input/rsna-2023-abdominal-trauma-detection","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:32:05.277918Z","iopub.execute_input":"2023-08-30T16:32:05.278291Z","iopub.status.idle":"2023-08-30T16:32:06.369504Z","shell.execute_reply.started":"2023-08-30T16:32:05.278261Z","shell.execute_reply":"2023-08-30T16:32:06.368117Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The raw files in DICOM format (`.dcm`) should be located in `train_images` and `test_images` folders. Let's print out how many subfolders are in each of them. Note that each subfolder contains data for one patient.","metadata":{}},{"cell_type":"code","source":"train_images_path = os.path.join(BASE_DIR, \"train_images\")\ntest_images_path = os.path.join(BASE_DIR, \"test_images\")\n\nprint(f\"Train data: {train_images_path} --> {len(os.listdir(train_images_path))} patients\")\nprint(f\"Test data: {test_images_path} --> {len(os.listdir(test_images_path))} patients\")","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:32:06.372395Z","iopub.execute_input":"2023-08-30T16:32:06.373346Z","iopub.status.idle":"2023-08-30T16:32:06.48245Z","shell.execute_reply.started":"2023-08-30T16:32:06.373305Z","shell.execute_reply":"2023-08-30T16:32:06.481339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since each subfolder in `train_images` can contain multiple scans, and each scan contains many images, let's find out the distribution of those. The simplest way to do so is to iterate over all nested folders and count the relevant values. ","metadata":{}},{"cell_type":"code","source":"def count_files_in_folder(path: str) -> int:\n    \"\"\"\n    Count the number of files in a given folder.\n    \n    This function recursively counts the number of files present in the specified folder and all its subfolders.\\\n\n    Args:\n        path (str): The path to the folder.\n    Returns:\n        int: The total number of files found within the folder and its subfolders.\n    \"\"\"\n    file_count = 0\n    for _, _, files in os.walk(path):\n        file_count += len(files)\n    return file_count","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:32:06.484289Z","iopub.execute_input":"2023-08-30T16:32:06.484719Z","iopub.status.idle":"2023-08-30T16:32:06.491456Z","shell.execute_reply.started":"2023-08-30T16:32:06.484668Z","shell.execute_reply":"2023-08-30T16:32:06.490293Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scans_per_patient, images_per_scan = [], []\nfor patient in os.listdir(train_images_path):\n    patient_folder = os.path.join(train_images_path, patient)\n    scans = os.listdir(patient_folder)\n    scans_per_patient.append(len(scans))\n    for scan in scans:\n        scan_folder = os.path.join(patient_folder, scan)\n        file_count = count_files_in_folder(scan_folder)\n        images_per_scan.append(file_count)","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:32:06.492761Z","iopub.execute_input":"2023-08-30T16:32:06.493098Z","iopub.status.idle":"2023-08-30T16:38:04.539636Z","shell.execute_reply.started":"2023-08-30T16:32:06.493043Z","shell.execute_reply":"2023-08-30T16:38:04.538691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,2,figsize=(10,4))\nbin_edges = np.arange(min(scans_per_patient) - 0.5, max(scans_per_patient) + 1.5, 1)\nbin_centers = np.arange(min(scans_per_patient), max(scans_per_patient) + 1, 1)\nax[0].hist(scans_per_patient, bins=bin_edges, align=\"mid\", rwidth=0.8, edgecolor=\"black\")\nax[0].set_xticks(bin_centers)\nax[0].set_xlabel(\"Scans\")\nax[0].set_ylabel(\"Occurences\")\nax[0].set_title(\"Scans/patient\")\n\nbin_edges = np.arange(0, max(images_per_scan), 50)\nax[1].hist(images_per_scan, bins=bin_edges)\nax[1].set_xlabel(\"Images\")\nax[1].set_ylabel(\"Occurences\")\nax[1].set_title(\"Images/scan\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:04.54099Z","iopub.execute_input":"2023-08-30T16:38:04.541834Z","iopub.status.idle":"2023-08-30T16:38:05.132681Z","shell.execute_reply.started":"2023-08-30T16:38:04.541802Z","shell.execute_reply":"2023-08-30T16:38:05.131447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, the patient's folder can have either 1 or 2 scans. The distribution of images per scan is bimodal with the majority of scan having around 200 images. \n\nFinally, we could compute the total amount of `.dcm` files throughout the dataset, and the average number of `.dcm` images per patient.","metadata":{}},{"cell_type":"code","source":"total_files = count_files_in_folder(train_images_path)\nprint(f\"The total number of DICOM files in the training dataset: {total_files}\")","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:05.137898Z","iopub.execute_input":"2023-08-30T16:38:05.138276Z","iopub.status.idle":"2023-08-30T16:38:14.438321Z","shell.execute_reply.started":"2023-08-30T16:38:05.138244Z","shell.execute_reply":"2023-08-30T16:38:14.43717Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mean_per_patient = total_files / len(os.listdir(train_images_path))\nprint(f\"The mean number of DICOM files per patient: {mean_per_patient:.1f}\")","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:14.440049Z","iopub.execute_input":"2023-08-30T16:38:14.440718Z","iopub.status.idle":"2023-08-30T16:38:14.448762Z","shell.execute_reply.started":"2023-08-30T16:38:14.440677Z","shell.execute_reply":"2023-08-30T16:38:14.448035Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Tabular data inspection <a id=\"3\"></a>","metadata":{}},{"cell_type":"markdown","source":"In this section we are going to inspect the training dataset.\nFirst, we have to load `train.csv` file and explore it using standard `pandas` tools.","metadata":{}},{"cell_type":"code","source":"train_data_path = os.path.join(BASE_DIR, \"train.csv\")\ntrain_data = pd.read_csv(train_data_path)","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:14.449928Z","iopub.execute_input":"2023-08-30T16:38:14.450575Z","iopub.status.idle":"2023-08-30T16:38:14.47911Z","shell.execute_reply.started":"2023-08-30T16:38:14.450545Z","shell.execute_reply":"2023-08-30T16:38:14.478285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"Length of train_data: {len(train_data)} items\")","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:14.480662Z","iopub.execute_input":"2023-08-30T16:38:14.481381Z","iopub.status.idle":"2023-08-30T16:38:14.487192Z","shell.execute_reply.started":"2023-08-30T16:38:14.481342Z","shell.execute_reply":"2023-08-30T16:38:14.486098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `train.csv` contains the same number of entries as the number of folders in `train_images`.","metadata":{}},{"cell_type":"code","source":"train_data.columns","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:14.488926Z","iopub.execute_input":"2023-08-30T16:38:14.489365Z","iopub.status.idle":"2023-08-30T16:38:14.500007Z","shell.execute_reply.started":"2023-08-30T16:38:14.489327Z","shell.execute_reply":"2023-08-30T16:38:14.498988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It will be useful to have separate list of injury-related columns:","metadata":{}},{"cell_type":"code","source":"injury_columns = train_data.columns[1:]","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:14.501316Z","iopub.execute_input":"2023-08-30T16:38:14.501695Z","iopub.status.idle":"2023-08-30T16:38:14.511468Z","shell.execute_reply.started":"2023-08-30T16:38:14.501666Z","shell.execute_reply":"2023-08-30T16:38:14.510204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.sample(5)","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:14.513093Z","iopub.execute_input":"2023-08-30T16:38:14.514119Z","iopub.status.idle":"2023-08-30T16:38:14.545972Z","shell.execute_reply.started":"2023-08-30T16:38:14.514059Z","shell.execute_reply":"2023-08-30T16:38:14.544878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.1. Data distribution <a id=\"3.1\"></a>\nNow we are going to display the data distribution within each category. Since columns represent patients' conditions in the binary format, we can easily compute how many patients with each condition. To do so, we can count all the values in each column.\n\nNote that since the column `any_injury` does not a counterpart, we can manually compute this value.","metadata":{}},{"cell_type":"code","source":"label_counts = {}\nfor col in train_data.columns[1:]:\n    label_counts[col] = train_data[col].sum()\n    \nlabel_counts[\"any_healthy\"] = len(train_data) - label_counts[\"any_injury\"]\nlabel_counts","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:14.547396Z","iopub.execute_input":"2023-08-30T16:38:14.547695Z","iopub.status.idle":"2023-08-30T16:38:14.558055Z","shell.execute_reply.started":"2023-08-30T16:38:14.54767Z","shell.execute_reply":"2023-08-30T16:38:14.557279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, we have to prepare a dictionary contatining all available subcategories for each category.","metadata":{}},{"cell_type":"code","source":"subcategories = {}\nfor key in label_counts.keys():\n    category, subcategory = key.split(\"_\")\n    if category not in subcategories.keys():\n        subcategories[category] = []\n    subcategories[category].append(subcategory)\n\n# the following line simply reverses the subcategories for `any` category to make it consistent with the other\nsubcategories[\"any\"] = subcategories[\"any\"][::-1]\nsubcategories","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:14.559378Z","iopub.execute_input":"2023-08-30T16:38:14.559706Z","iopub.status.idle":"2023-08-30T16:38:14.572467Z","shell.execute_reply.started":"2023-08-30T16:38:14.559678Z","shell.execute_reply":"2023-08-30T16:38:14.571478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we have erything we need to display data distribution. Among possible visualizations I opt for donut chart using `matplotlib`.","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(2, 3, figsize=(9, 6))\naxs = axs.flatten()\n\nfor idx, category in enumerate(subcategories.keys()):\n    subcategory_values = [label_counts[f\"{category}_{sub}\"] for sub in subcategories[category]]\n    total = sum(subcategory_values)\n    \n    wedges, texts, autotexts = axs[idx].pie(\n        subcategory_values,\n        labels=subcategories[category],\n        autopct=lambda pct: f\"{pct:.1f}%\\n({int(pct / 100 * total)})\" if pct > 5.0  else \"\",\n        startangle=0,\n        colors=[\"#a8e6a0\", \"#ff9900\", \"#ff3333\"],\n        wedgeprops={\"edgecolor\": \"gray\"},\n    )\n    \n    axs[idx].set_title(f\"Category: {category.upper()}\")\n    \n    # Draw circle in the center to create the donut shape\n    centre_circle = plt.Circle((0, 0), 0.5, fc=\"white\")\n    axs[idx].add_artist(centre_circle)\n    \nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:14.575184Z","iopub.execute_input":"2023-08-30T16:38:14.575925Z","iopub.status.idle":"2023-08-30T16:38:15.429667Z","shell.execute_reply.started":"2023-08-30T16:38:14.575884Z","shell.execute_reply":"2023-08-30T16:38:15.428811Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As one can see there are 855 patients with injury of any kind. Most of them are liver and spleen injuries, while there is a handful of data with bowel injuries. \n\nNext we are going to see the distribution of data within `any_injury`. This is because some patient can have multiple injuries.","metadata":{}},{"cell_type":"markdown","source":"## 3.2. Multiple injuries <a id=\"3.2\"></a>","metadata":{}},{"cell_type":"code","source":"# all injury columns except for `any_injury`\nindividual_injury_columns = injury_columns[:-1]\nindividual_injury_columns","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:15.430883Z","iopub.execute_input":"2023-08-30T16:38:15.431422Z","iopub.status.idle":"2023-08-30T16:38:15.437776Z","shell.execute_reply.started":"2023-08-30T16:38:15.431392Z","shell.execute_reply":"2023-08-30T16:38:15.436973Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# all individual injury columns that define some injury types (binary or tertiary)\npositive_injury_columns = [col for col in individual_injury_columns if not \"healthy\" in col]\npositive_injury_columns","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:15.439186Z","iopub.execute_input":"2023-08-30T16:38:15.439717Z","iopub.status.idle":"2023-08-30T16:38:15.450279Z","shell.execute_reply.started":"2023-08-30T16:38:15.439689Z","shell.execute_reply":"2023-08-30T16:38:15.449391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here are a going to make a `Pandas.Series` that has all contibutions from individual injury types.","metadata":{}},{"cell_type":"code","source":"injury_summary = train_data[positive_injury_columns].apply(lambda row: \",\".join(row.index[row == 1]), axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:15.451671Z","iopub.execute_input":"2023-08-30T16:38:15.452009Z","iopub.status.idle":"2023-08-30T16:38:15.866998Z","shell.execute_reply.started":"2023-08-30T16:38:15.451982Z","shell.execute_reply":"2023-08-30T16:38:15.865826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"injury_summary = injury_summary[injury_summary != \"\"]\ninjury_summary","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:15.868594Z","iopub.execute_input":"2023-08-30T16:38:15.869101Z","iopub.status.idle":"2023-08-30T16:38:15.879721Z","shell.execute_reply.started":"2023-08-30T16:38:15.869045Z","shell.execute_reply":"2023-08-30T16:38:15.878656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This summary confirms that some patiens has more than more injury. Let's plot this distribution as histogram and put all injury combination with less than 5 occurrences to the 'Other' category.","metadata":{}},{"cell_type":"code","source":"threshold = 5\ninjury_counts = injury_summary.value_counts()\ninjury_counts[\"Other\"] = injury_counts[injury_counts < threshold].sum()\ninjury_counts = injury_counts[injury_counts >= threshold]","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:15.881386Z","iopub.execute_input":"2023-08-30T16:38:15.882058Z","iopub.status.idle":"2023-08-30T16:38:15.897451Z","shell.execute_reply.started":"2023-08-30T16:38:15.882018Z","shell.execute_reply":"2023-08-30T16:38:15.896512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10, 6))\ninjury_counts.plot(kind=\"barh\").invert_yaxis()\nplt.title(\"Distribution of injury types\")\nplt.xlabel(\"Count\")\nplt.ylabel(\"Injury type\")\nplt.xticks(rotation=90)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:15.899147Z","iopub.execute_input":"2023-08-30T16:38:15.899515Z","iopub.status.idle":"2023-08-30T16:38:16.408899Z","shell.execute_reply.started":"2023-08-30T16:38:15.899485Z","shell.execute_reply":"2023-08-30T16:38:16.408102Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here we have analysed the patients with injury of any time (27.2% of the training data, or 855 samples). We conclude that most patient have only one injury (with `liver_low` being the most frequent case). Substantial part of patients have simultaneously 2 injuries. Patients with more than 2 types of injuries fall into the 'Other' category and represent minority class.","metadata":{}},{"cell_type":"markdown","source":"# 4. Exploring images <a id=\"4\"></a>\n\nLet's first define a function to fetch some random image from the dataset.","metadata":{}},{"cell_type":"code","source":"def get_random_image(seed, mode=\"train\"):\n    \"\"\"\n    Get a random image file path from the specified dataset mode.\n    \n    This function takes a seed value and an optional 'mode' argument to determine whether\n    to retrieve a random image from the 'train' or 'test' dataset. It generates a random\n    image path by following the hierarchy: dataset mode -> patient ID -> scan ID -> image ID.\n\n    Args:\n        seed (int): Seed value used for randomization.\n        mode (str, optional): The dataset mode. Either 'train' or 'test'. Defaults to 'train'.\n\n    Returns:\n        str: File path to a randomly selected image.\n    \"\"\"\n    random.seed(seed)\n    \n    if mode == \"train\":\n        random_patient_id = random.choice(os.listdir(train_images_path))\n        random_patient_path = os.path.join(train_images_path, random_patient_id)\n    elif mode == \"test\":\n        random_patient_id = random.choice(os.listdir(test_images_path))\n        random_patient_path = os.path.join(test_images_path, random_patient_id)\n    \n    random_scan_id = random.choice(os.listdir(random_patient_path))\n    random_scan_path = os.path.join(random_patient_path, random_scan_id)\n    \n    random_image_id = random.choice(os.listdir(random_scan_path))\n    random_image_path = os.path.join(random_scan_path, random_image_id)\n    \n    return random_image_path","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:16.410017Z","iopub.execute_input":"2023-08-30T16:38:16.410987Z","iopub.status.idle":"2023-08-30T16:38:16.41901Z","shell.execute_reply.started":"2023-08-30T16:38:16.410955Z","shell.execute_reply":"2023-08-30T16:38:16.418095Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_path = get_random_image(seed=42)\nimage_path","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:16.420253Z","iopub.execute_input":"2023-08-30T16:38:16.420561Z","iopub.status.idle":"2023-08-30T16:38:16.435267Z","shell.execute_reply.started":"2023-08-30T16:38:16.420534Z","shell.execute_reply":"2023-08-30T16:38:16.434316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To work with the DICOM files, we can make use of `pydicom` library. In a nutshell, this library allows to access the main attributes of the DICOM instance via the keys of a dictionary. The detailed guide can be found elsewhere (https://pydicom.github.io/pydicom/dev/old/pydicom_user_guide.html).","metadata":{}},{"cell_type":"code","source":"dicom_data = pydicom.dcmread(image_path)\ndicom_data","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:16.441866Z","iopub.execute_input":"2023-08-30T16:38:16.44264Z","iopub.status.idle":"2023-08-30T16:38:16.464498Z","shell.execute_reply.started":"2023-08-30T16:38:16.442608Z","shell.execute_reply":"2023-08-30T16:38:16.463538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As one can see, the DICOM file contains a lot of useful information and associated metadata. We are interested in the 'Pixel Data' as it contains the image raw pixels. Note also that to convert pixel values to the standard 8-bit representation, it is neccesary to apply rescaling and windowing operations. For further details, check this link: https://pydicom.github.io/pydicom/stable/reference/handlers.pixel_data.html.","metadata":{}},{"cell_type":"markdown","source":"## 4.1. Displaying random DICOM image <a id=\"4.1\"></a>","metadata":{}},{"cell_type":"code","source":"def dicom_to_gray(dicom_data: pydicom.dataset.FileDataset) -> np.ndarray:\n    \"\"\"\n    Convert a DICOM image to grayscale.\n\n    This function takes DICOM pixel data and performs the following steps to convert it to grayscale:\n    1. Rescales the pixel values using pydicom's apply_rescale function.\n    2. Applies windowing to adjust pixel values using pydicom's apply_windowing function.\n    3. Normalizes pixel values to the range [0, 255] and converts them to uint8.\n\n    Args:\n        dicom_data (pydicom.dataset.FileDataset): The input DICOM data containing pixel_array and metadata.\n\n    Returns:\n        numpy.ndarray: A 2D numpy array containing the grayscale image data.\n    \"\"\"\n    dicom_image = dicom_data.pixel_array\n    dicom_image_rescaled = pydicom.pixel_data_handlers.apply_rescale(dicom_image, dicom_data)\n    dicom_image_windowed = pydicom.pixel_data_handlers.apply_windowing(dicom_image_rescaled, dicom_data)\n    \n    min_value = dicom_image_windowed.min()\n    max_value = dicom_image_windowed.max()\n    image_gray = ((dicom_image_windowed - min_value) / (max_value - min_value) * 255).astype(\"uint8\")\n    \n    return image_gray","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:16.465993Z","iopub.execute_input":"2023-08-30T16:38:16.466875Z","iopub.status.idle":"2023-08-30T16:38:16.474595Z","shell.execute_reply.started":"2023-08-30T16:38:16.466836Z","shell.execute_reply":"2023-08-30T16:38:16.47341Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def fetch_annotation(patient_id):\n    \"\"\"\n    This function parses a row from the metadata table as a multiline string.\n    Args:\n        patient_id (int): patient ID.\n    Returns:\n        str: multiline annotation string \n    \"\"\"\n    data = train_data[train_data[\"patient_id\"] == patient_id].squeeze()\n    annotation_string = \"\"\n    for k, v in data.items():    \n        annotation_string += f\"{k} --- {v}<br>\"\n    return annotation_string","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:16.476319Z","iopub.execute_input":"2023-08-30T16:38:16.476715Z","iopub.status.idle":"2023-08-30T16:38:16.4889Z","shell.execute_reply.started":"2023-08-30T16:38:16.476678Z","shell.execute_reply":"2023-08-30T16:38:16.487944Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def display_image(path: str, add_annotation: bool=True) -> None:\n    \"\"\"\n    Display a grayscale DICOM image with optional annotation using Plotly.\n    \n    Args:\n        path (str): Path to the DICOM file.\n        add_annotation (bool, optional): Whether to add annotation to the image. Default is True.\n        \n    Returns:\n        None\n    \"\"\"\n    if not os.path.exists(path):\n        raise FileNotFoundError(f\"File not found at path: {path}\")\n    if not os.path.isfile(path):\n        raise IsADirectoryError(f\"The provided path '{path}' points to a directory.\")\n    \n    dicom_data = pydicom.dcmread(path)\n    image = dicom_to_gray(dicom_data)\n    \n    fig = go.Figure()\n    fig.add_trace(px.imshow(image, binary_string=True).data[0])\n    fig.update_layout(\n        xaxis={\"visible\": False, \"showticklabels\": False},\n        yaxis={\"visible\": False, \"showticklabels\": False},\n        title_text=f\"Patient ID: {dicom_data.PatientID} | Scan ID: {dicom_data.SeriesInstanceUID.split('.')[-1]} | Instance ID: {dicom_data.InstanceNumber}\",\n        title_x=0.5\n    )\n    \n    if add_annotation:\n        annotation_text = fetch_annotation(int(dicom_data.PatientID))\n        annotation=[\n            dict(\n                text=annotation_text,\n                align=\"left\",\n                showarrow=False,\n                xref=\"paper\",\n                yref=\"paper\",\n                x=1.0,\n                y=0.5,\n                bordercolor=\"black\",\n                borderwidth=1,\n                font=dict(family=\"Courier New\", size=18)\n            )\n        ]\n        fig.update_layout(annotations=annotation)\n        \n    \n    fig.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:16.492359Z","iopub.execute_input":"2023-08-30T16:38:16.49265Z","iopub.status.idle":"2023-08-30T16:38:16.50495Z","shell.execute_reply.started":"2023-08-30T16:38:16.492625Z","shell.execute_reply":"2023-08-30T16:38:16.50398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have made some helper functions for processing the DICOM raw data and fetching annotation from the `.csv` file. Now we can display grayscale images together with its annotation using the interactive `plotly` library.","metadata":{}},{"cell_type":"code","source":"image_path = get_random_image(seed=42)\ndisplay_image(image_path, add_annotation=True)","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:16.506257Z","iopub.execute_input":"2023-08-30T16:38:16.506619Z","iopub.status.idle":"2023-08-30T16:38:18.293106Z","shell.execute_reply.started":"2023-08-30T16:38:16.506592Z","shell.execute_reply":"2023-08-30T16:38:18.292204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.2. Converting DICOM image to PyTorch tensor <a id=\"4.2\"></a>\n\nFinally, we can make a function that takes the DICOM file and convert the image into PyTorch tensor with value in the range [0, 1].","metadata":{}},{"cell_type":"code","source":"def dicom_to_tensor(dicom_data: pydicom.dataset.FileDataset) -> torch.Tensor:\n    \"\"\"\n    Convert a DICOM image to PyTorch Tensor.\n\n    This function takes DICOM pixel data and performs the following steps to convert it to tensor:\n    1. Rescales the pixel values using pydicom's apply_rescale function.\n    2. Applies windowing to adjust pixel values using pydicom's apply_windowing function.\n    3. Normalizes pixel values to the range [0, 1].\n    4. Convert numpy array to tensor with float32 datatype.\n\n    Args:\n        dicom_data (pydicom.dataset.FileDataset): The input DICOM data containing pixel_array and metadata.\n\n    Returns:\n        torch.Tensor: PyTorch Tensor containing image data.\n    \"\"\"\n    dicom_image = dicom_data.pixel_array\n    dicom_image_rescaled = pydicom.pixel_data_handlers.apply_rescale(dicom_image, dicom_data)\n    dicom_image_windowed = pydicom.pixel_data_handlers.apply_windowing(dicom_image_rescaled, dicom_data)\n    \n    min_value = dicom_image_windowed.min()\n    max_value = dicom_image_windowed.max()\n    np_array = (dicom_image_windowed - min_value) / (max_value - min_value)\n    tensor = torch.tensor(np_array, dtype=torch.float32)\n    \n    return tensor","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:18.294351Z","iopub.execute_input":"2023-08-30T16:38:18.294654Z","iopub.status.idle":"2023-08-30T16:38:18.302027Z","shell.execute_reply.started":"2023-08-30T16:38:18.294627Z","shell.execute_reply":"2023-08-30T16:38:18.301037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_path = get_random_image(seed=42)\ndicom_data = pydicom.dcmread(image_path)\ndicom_tensor = dicom_to_tensor(dicom_data)\ndicom_tensor","metadata":{"execution":{"iopub.status.busy":"2023-08-30T16:38:18.303308Z","iopub.execute_input":"2023-08-30T16:38:18.303594Z","iopub.status.idle":"2023-08-30T16:38:18.456695Z","shell.execute_reply.started":"2023-08-30T16:38:18.303568Z","shell.execute_reply":"2023-08-30T16:38:18.455925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If you find this notebook useful, please upvote! <br>\nIn case you have remarks, questions or suggestions, do not hesitate to leave a comment!","metadata":{}}],"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}}