{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"display:fill;\n            border-radius:15px;\n            background-color:#00bd35;\n            font-size:190%;\n            font-family:cursive;\n            letter-spacing:0.5px;\n            padding:10px;\n            color:white;\n            border-style: solid;\n            border-color: black;\n            text-align:center;\">\n<b>\nUnleashing the Healing Potential: Abdominal Trauma Detection</b>\n</div>\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Introduction</p></div>\n\nThe RSNA Abdominal Trauma Detection AI Challenge aims to address a critical issue in healthcare: the prompt and accurate diagnosis of traumatic injuries in the abdomen using computed tomography (CT) scans. Traumatic injury is a leading cause of death worldwide, and CT scans have become indispensable in evaluating patients with suspected abdominal injuries due to their ability to provide detailed cross-sectional images. However, interpreting CT scans for abdominal trauma can be complex and time-consuming, especially when multiple injuries or subtle active bleeding areas are present.\n\nTo tackle this challenge, the competition seeks to leverage the power of artificial intelligence and machine learning to assist medical professionals in rapidly and precisely detecting injuries and grading their severity. By developing advanced algorithms for this purpose, the goal is to improve trauma care and patient outcomes on a global scale.\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#FF7F50;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Dataset Overview:</p></div>\n\n\n\nThe dataset provided for the RSNA Abdominal Trauma Detection AI Challenge contains information related to patients and their abdominal health status. The dataset includes the following variables (features):\n\n1. **patient_id:** An identifier for each patient.\n2. **bowel_healthy:** Binary variable (0 or 1) indicating the health status of the bowel (0: not healthy, 1: healthy).\n3. **bowel_injury:** Binary variable (0 or 1) indicating the presence of injury in the bowel (0: no injury, 1: injury present).\n4. **extravasation_healthy:** Binary variable (0 or 1) indicating the health status of extravasation (0: not healthy, 1: healthy).\n5. **extravasation_injury:** Binary variable (0 or 1) indicating the presence of injury in extravasation (0: no injury, 1: injury present).\n6. **kidney_healthy:** Binary variable (0 or 1) indicating the health status of the kidneys (0: not healthy, 1: healthy).\n7. **kidney_low:** Binary variable (0 or 1) indicating low health status of the kidneys (0: not low, 1: low health).\n8. **kidney_high:** Binary variable (0 or 1) indicating high health status of the kidneys (0: not high, 1: high health).\n9. **liver_healthy:** Binary variable (0 or 1) indicating the health status of the liver (0: not healthy, 1: healthy).\n10. **liver_low:** Binary variable (0 or 1) indicating low health status of the liver (0: not low, 1: low health).\n11. **liver_high:** Binary variable (0 or 1) indicating high health status of the liver (0: not high, 1: high health).\n12. **spleen_healthy:** Binary variable (0 or 1) indicating the health status of the spleen (0: not healthy, 1: healthy).\n13. **spleen_low:** Binary variable (0 or 1) indicating low health status of the spleen (0: not low, 1: low health).\n14. **spleen_high:** Binary variable (0 or 1) indicating high health status of the spleen (0: not high, 1: high health).\n15. **any_injury:** An integer variable indicating the number of injuries detected (0: no injury, 1: one injury, 2: two injuries, and so on).\n\nThe dataset contains patient-specific health information for\ndifferent abdominal organs (bowel, extravasation, kidney, liver, spleen) and indicates whether injuries are present in these organs. Additionally, the \"any_injury\" variable provides an overall count of injuries detected in a patient. The dataset serves as the foundation for participants in the competition to develop AI models that can accurately detect severe injuries to the internal abdominal organs and any active internal bleeding.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Import Modules</p></div>\n","metadata":{}},{"cell_type":"code","source":"%%capture\n!pip install pydicom matplotlib","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:07:26.570302Z","iopub.execute_input":"2023-07-28T07:07:26.570922Z","iopub.status.idle":"2023-07-28T07:07:39.567585Z","shell.execute_reply.started":"2023-07-28T07:07:26.570883Z","shell.execute_reply":"2023-07-28T07:07:39.566167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport pydicom\nimport os\nimport numpy as np\nfrom matplotlib import pyplot as plt\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom sklearn.cluster import KMeans\nfrom sklearn.preprocessing import StandardScaler\n\nplt.rcParams['figure.figsize'] = (12,6)\nplt.style.use('fivethirtyeight')\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:07:51.450706Z","iopub.execute_input":"2023-07-28T07:07:51.451091Z","iopub.status.idle":"2023-07-28T07:07:51.457692Z","shell.execute_reply.started":"2023-07-28T07:07:51.451059Z","shell.execute_reply":"2023-07-28T07:07:51.45652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Load the Dataset</p></div>","metadata":{}},{"cell_type":"code","source":"# Load the dataset `into` a Pandas DataFrame\ndf = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv')\ndf.head().style.set_properties(**{'background-color':'royalblue','color':'white','border-color':'#8b8c8c'})","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:07:56.493773Z","iopub.execute_input":"2023-07-28T07:07:56.494786Z","iopub.status.idle":"2023-07-28T07:07:56.594547Z","shell.execute_reply.started":"2023-07-28T07:07:56.494749Z","shell.execute_reply":"2023-07-28T07:07:56.592894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"meta_df = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_series_meta.csv')\nmeta_df.head().style.set_properties(**{'background-color':'orange','color':'white','border-color':'#8b8c8c'})","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:08:01.117248Z","iopub.execute_input":"2023-07-28T07:08:01.117675Z","iopub.status.idle":"2023-07-28T07:08:01.140174Z","shell.execute_reply.started":"2023-07-28T07:08:01.117641Z","shell.execute_reply":"2023-07-28T07:08:01.139303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Data Preprocessing</p></div>\n\n* Check for missing values and handle them if necessary.\n* Check for data types and convert them if needed (e.g., converting binary data to boolean).\n* Identify and address any data quality issues, if present.\n","metadata":{}},{"cell_type":"code","source":"# Check for Missing Values\nmissing_values = df.isnull().sum()\nprint(\"Missing Values:\")\nprint(missing_values)\n\n# Check Data Types and Convert Binary Data to Boolean\nbinary_columns = [\n    'bowel_healthy', 'bowel_injury', 'extravasation_healthy', 'extravasation_injury',\n    'kidney_healthy', 'kidney_low', 'kidney_high', 'liver_healthy', 'liver_low', 'liver_high',\n    'spleen_healthy', 'spleen_low', 'spleen_high'\n]\ndf[binary_columns] = df[binary_columns].astype(bool)\n\n# Address Data Quality Issues\n# In this simple example, we assume no data quality issues are present.\n\n# Display the preprocessed DataFrame\nprint(\"\\nPreprocessed DataFrame:\")\nprint(df)","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:08:08.757477Z","iopub.execute_input":"2023-07-28T07:08:08.758166Z","iopub.status.idle":"2023-07-28T07:08:08.793541Z","shell.execute_reply.started":"2023-07-28T07:08:08.75813Z","shell.execute_reply":"2023-07-28T07:08:08.792176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Descriptive Statistics</p></div>\n\n\n* Provide summary statistics for relevant variables (e.g., mean, median, standard deviation) to get an overview of the dataset.\n* Generate counts and percentages to show the distribution of categorical variables.","metadata":{}},{"cell_type":"code","source":"# Summary statistics for relevant variables\nstyled_data = df.describe().style\\\n.background_gradient(cmap='coolwarm')\\\n.set_properties(**{'text-align':'center','border':'1px solid black'})\n\n# display styled data\ndisplay(styled_data)","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:08:15.060822Z","iopub.execute_input":"2023-07-28T07:08:15.061382Z","iopub.status.idle":"2023-07-28T07:08:15.116461Z","shell.execute_reply.started":"2023-07-28T07:08:15.061329Z","shell.execute_reply":"2023-07-28T07:08:15.11518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Generate counts and percentages for categorical variables\ncategorical_columns = [\n    'bowel_healthy', 'bowel_injury', 'extravasation_healthy', 'extravasation_injury',\n    'kidney_healthy', 'kidney_low', 'kidney_high', 'liver_healthy', 'liver_low', 'liver_high',\n    'spleen_healthy', 'spleen_low', 'spleen_high', 'any_injury'\n]\n\ncounts = df[categorical_columns].apply(pd.Series.value_counts)\npercentages = (counts / df.shape[0]) * 100\n\n# Display the summary statistics and counts/percentages\nprint(\"Summary Statistics:\")\n#print(summary_stats)\n\nprint(\"\\nCounts of Categorical Variables:\")\nprint(counts)\n\nprint(\"\\nPercentages of Categorical Variables (%):\")\nprint(percentages)\n","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:08:19.034964Z","iopub.execute_input":"2023-07-28T07:08:19.036177Z","iopub.status.idle":"2023-07-28T07:08:19.072954Z","shell.execute_reply.started":"2023-07-28T07:08:19.036133Z","shell.execute_reply":"2023-07-28T07:08:19.071763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Exploratory Data Analysis (EDA)</p></div>\n\n\n\n* Visualize the distribution of different variables using appropriate plots (e.g., bar plots, pie charts, histograms).\n* Analyze the relationship between variables (e.g., correlation between organ health and injury status).\n* Explore any patterns or trends in the data.\n* Visualize a sample image","metadata":{}},{"cell_type":"code","source":"# Organ columns: bowel, extravasation, kidney, liver, spleen\norgan_columns = ['bowel', 'extravasation', 'kidney', 'liver', 'spleen']\n\n# Create a new DataFrame to store the counts\norgan_counts = pd.DataFrame()\norgan_counts['Organ'] = organ_columns\n\n# Loop through organ columns and count healthy and injury status for each organ\nfor organ in organ_columns:\n    healthy_col = f'{organ}_healthy'\n    injury_col = f'{organ}_injury'\n    \n    # Check if the columns exist in the DataFrame\n    if healthy_col in df.columns and injury_col in df.columns:\n        organ_counts[f'{organ}_healthy'] = df[healthy_col].sum()\n        organ_counts[f'{organ}_injury'] = df[injury_col].sum()\n    else:\n        # Handle the case if the columns are missing\n        print(f\"Warning: Columns for {organ} healthy/injury status are missing in the DataFrame.\")\n        organ_counts[f'{organ}_healthy'] = 0\n        organ_counts[f'{organ}_injury'] = 0\n\n# Melt the DataFrame to have a single 'Status' column\norgan_counts_melted = organ_counts.melt(id_vars=['Organ'], var_name='Status', value_name='Count')\n\n# Bar plot for distribution of organ health and injury status\nfig = px.bar(\n    organ_counts_melted,\n    x='Organ',\n    y='Count',\n    color='Status',\n    barmode='group',\n    labels=dict(x='Organ', y='Count', Status='Status'),\n    title='Distribution of Organ Health and Injury Status',\n)\nfig.update_layout(showlegend=True)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:08:25.391938Z","iopub.execute_input":"2023-07-28T07:08:25.392397Z","iopub.status.idle":"2023-07-28T07:08:27.324248Z","shell.execute_reply.started":"2023-07-28T07:08:25.392361Z","shell.execute_reply":"2023-07-28T07:08:27.322627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Organ columns: bowel, extravasation, kidney, liver, spleen\norgan_columns = ['bowel', 'extravasation', 'kidney', 'liver', 'spleen']\n\n# Check if the 'injury' columns are present in the DataFrame\ninjury_columns = [f'{organ}_injury' for organ in organ_columns]\nmissing_columns = set(organ_columns + injury_columns) - set(df.columns)\n\nif missing_columns:\n    # Handle the case if any of the required columns are missing\n    print(f\"Warning: Columns for {', '.join(missing_columns)} are missing in the DataFrame.\")\n    for col in missing_columns:\n        df[col] = 0\n\n# Heatmap to analyze the correlation between organ health and injury status\ncorrelation_df = df[organ_columns + injury_columns]\ncorrelation_matrix = correlation_df.corr()\n\nfig = px.imshow(\n    correlation_matrix,\n    x=correlation_df.columns,\n    y=correlation_df.columns,\n    labels=dict(x='Organ', y='Organ', color='Correlation'),\n    title='Correlation Between Organ Health and Injury Status',\n)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:08:36.666152Z","iopub.execute_input":"2023-07-28T07:08:36.667205Z","iopub.status.idle":"2023-07-28T07:08:36.937543Z","shell.execute_reply.started":"2023-07-28T07:08:36.667157Z","shell.execute_reply":"2023-07-28T07:08:36.936462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. Occurrence of Injuries in Different Organs:\n\n* The bar plot displays the occurrence of injuries in different organs (bowel, extravasation, kidney, liver, spleen) in the dataset.\n* The injuries are represented by the height of the bars, and it shows that the bowel has the lowest number of injuries (0), while the other organs have one or more injuries each.\n\n2. Overall Prevalence of Injuries:\n\n* The variable \"any_injury\" indicates the overall prevalence of injuries in the dataset.\n* In this dataset, the overall injury count is 11 (sum of \"any_injury\" for all patients).\n\n3. Relationship Between Injuries in Different Organs:\n\n* The heatmap illustrates the correlation between injuries in different organs.\n* The diagonal of the heatmap shows the correlation of each organ's injuries with itself, which is always 1 (perfect correlation).\n* The off-diagonal elements indicate the correlations between injuries in different organs. In this small dataset, there may not be strong correlations between injuries in different organs.\n\nBased on the analysis and visualizations, we can observe the distribution of injuries in different organs and the overall prevalence of injuries in the dataset. The heatmap helps us identify any potential relationships between injuries in different organs. In a larger dataset, a more detailed analysis may reveal additional insights into the patterns and connections between injuries across various organs.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Visualize a sample image</p></div>\n\n","metadata":{}},{"cell_type":"code","source":"import pydicom","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:08:46.929933Z","iopub.execute_input":"2023-07-28T07:08:46.93033Z","iopub.status.idle":"2023-07-28T07:08:46.935134Z","shell.execute_reply.started":"2023-07-28T07:08:46.930299Z","shell.execute_reply":"2023-07-28T07:08:46.933836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def standardize_pixel_array(dcm: pydicom.dataset.FileDataset) -> np.ndarray:\n    # Correct DICOM pixel_array if PixelRepresentation == 1.\n    pixel_array = dcm.pixel_array\n    if dcm.PixelRepresentation == 1:\n        bit_shift = dcm.BitsAllocated - dcm.BitsStored\n        dtype = pixel_array.dtype\n        new_array = (pixel_array << bit_shift).astype(dtype) >>  bit_shift\n        pixel_array = pydicom.pixel_data_handlers.util.apply_modality_lut(new_array, dcm)\n    return pixel_array\n\ntrain_dicom_tags = pd.read_parquet('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_dicom_tags.parquet', engine='pyarrow')","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:10:20.661166Z","iopub.execute_input":"2023-07-28T07:10:20.662129Z","iopub.status.idle":"2023-07-28T07:10:26.697156Z","shell.execute_reply.started":"2023-07-28T07:10:20.662071Z","shell.execute_reply":"2023-07-28T07:10:26.696265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# path of the image in index 49954\nsample_image = train_dicom_tags.loc[49954]['path']\n\n# Open the DICOM file using pydicom.\ndcm = pydicom.read_file(os.path.join('/kaggle/input/rsna-2023-abdominal-trauma-detection',sample_image))","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:10:36.660228Z","iopub.execute_input":"2023-07-28T07:10:36.660672Z","iopub.status.idle":"2023-07-28T07:10:36.677607Z","shell.execute_reply.started":"2023-07-28T07:10:36.660616Z","shell.execute_reply":"2023-07-28T07:10:36.676478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dicom_tags.head(3)","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:10:46.522029Z","iopub.execute_input":"2023-07-28T07:10:46.522431Z","iopub.status.idle":"2023-07-28T07:10:46.551431Z","shell.execute_reply.started":"2023-07-28T07:10:46.522402Z","shell.execute_reply":"2023-07-28T07:10:46.549396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualize a sample image\ndef plot_dicom_image(image_path):\n    ds = pydicom.dcmread(image_path)\n    plt.imshow(ds.pixel_array, cmap=plt.cm.bone)\n    plt.axis('off')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:10:50.966564Z","iopub.execute_input":"2023-07-28T07:10:50.966937Z","iopub.status.idle":"2023-07-28T07:10:50.973374Z","shell.execute_reply.started":"2023-07-28T07:10:50.966909Z","shell.execute_reply":"2023-07-28T07:10:50.972055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_image_path = '/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/49954/41479/378.dcm'\nplot_dicom_image(sample_image_path)","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:10:56.830373Z","iopub.execute_input":"2023-07-28T07:10:56.830754Z","iopub.status.idle":"2023-07-28T07:10:57.133812Z","shell.execute_reply.started":"2023-07-28T07:10:56.830726Z","shell.execute_reply":"2023-07-28T07:10:57.132695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_dicom_images(directory):\n    dicom_images = []\n    for filename in os.listdir(directory):\n        if filename.endswith(\".dcm\"):\n            dicom_file = os.path.join(directory, filename)\n            dicom_image = pydicom.dcmread(dicom_file)\n            dicom_images.append(dicom_image)\n    return dicom_images\n\ndef rescale_pixel_array(pixel_array, window_level, window_width):\n    # Rescale the pixel values based on the window level and window width\n    min_value = window_level - window_width // 2\n    max_value = window_level + window_width // 2\n    rescaled_pixel_array = np.clip(pixel_array, min_value, max_value)\n    rescaled_pixel_array = (rescaled_pixel_array - min_value) / (max_value - min_value)\n    return rescaled_pixel_array\n\ndef visualize_dicom_images(dicom_images, num_rows=4, num_cols=4, window_level=40, window_width=80):\n    fig, axes = plt.subplots(num_rows, num_cols, figsize=(15, 15))\n    for i, ax in enumerate(axes.flat):\n        if i < len(dicom_images):\n            dicom_image = dicom_images[i]\n            image_data = dicom_image.pixel_array.astype(np.float32)\n            rescaled_image = rescale_pixel_array(image_data, window_level, window_width)\n            ax.imshow(rescaled_image, cmap=plt.cm.bone)\n            ax.axis(\"off\")\n            ax.set_title(f\"Slice {i+1}\")\n\n    # Hide any empty subplots\n    for i in range(len(dicom_images), num_rows*num_cols):\n        axes.flat[i].axis(\"off\")\n\n    # Add a color bar to indicate pixel intensity values\n    cax = fig.add_axes([0.92, 0.15, 0.02, 0.7])\n    norm = plt.cm.colors.Normalize(vmin=0, vmax=1)\n    cbar = plt.colorbar(plt.cm.ScalarMappable(norm=norm, cmap=plt.cm.bone), cax=cax)\n    cbar.ax.set_ylabel(\"Pixel Intensity\")\n\n    plt.tight_layout()\n    plt.show()\n\nif __name__ == \"__main__\":\n    # Replace 'path_to_directory' with the actual path where your DICOM images are located\n    path_to_directory = \"/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/49954/41479\"\n    dicom_images = load_dicom_images(path_to_directory)\n    visualize_dicom_images(dicom_images, num_rows=3, num_cols=3, window_level=40, window_width=80)\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-28T07:11:02.633285Z","iopub.execute_input":"2023-07-28T07:11:02.634371Z","iopub.status.idle":"2023-07-28T07:11:11.959955Z","shell.execute_reply.started":"2023-07-28T07:11:02.634322Z","shell.execute_reply":"2023-07-28T07:11:11.956875Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_image_path = '/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/156/64005/106.dcm'\nplot_dicom_image(sample_image_path)\n","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:11:22.118761Z","iopub.execute_input":"2023-07-28T07:11:22.119183Z","iopub.status.idle":"2023-07-28T07:11:22.376075Z","shell.execute_reply.started":"2023-07-28T07:11:22.11915Z","shell.execute_reply":"2023-07-28T07:11:22.375203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_dicom_images(directory):\n    dicom_images = []\n    for filename in os.listdir(directory):\n        if filename.endswith(\".dcm\"):\n            dicom_file = os.path.join(directory, filename)\n            dicom_image = pydicom.dcmread(dicom_file)\n            dicom_images.append(dicom_image)\n    return dicom_images\n\ndef rescale_pixel_array(pixel_array, window_level, window_width):\n    # Rescale the pixel values based on the window level and window width\n    min_value = window_level - window_width // 2\n    max_value = window_level + window_width // 2\n    rescaled_pixel_array = np.clip(pixel_array, min_value, max_value)\n    rescaled_pixel_array = (rescaled_pixel_array - min_value) / (max_value - min_value)\n    return rescaled_pixel_array\n\ndef visualize_dicom_images(dicom_images, num_rows=4, num_cols=4, window_level=40, window_width=80):\n    fig, axes = plt.subplots(num_rows, num_cols, figsize=(15, 15))\n    for i, ax in enumerate(axes.flat):\n        if i < len(dicom_images):\n            dicom_image = dicom_images[i]\n            image_data = dicom_image.pixel_array.astype(np.float32)\n            rescaled_image = rescale_pixel_array(image_data, window_level, window_width)\n            ax.imshow(rescaled_image, cmap=plt.cm.bone)\n            ax.axis(\"off\")\n            ax.set_title(f\"Slice {i+1}\")\n\n    # Hide any empty subplots\n    for i in range(len(dicom_images), num_rows*num_cols):\n        axes.flat[i].axis(\"off\")\n\n    # Add a color bar to indicate pixel intensity values\n    cax = fig.add_axes([0.92, 0.15, 0.02, 0.7])\n    norm = plt.cm.colors.Normalize(vmin=0, vmax=1)\n    cbar = plt.colorbar(plt.cm.ScalarMappable(norm=norm, cmap=plt.cm.bone), cax=cax)\n    cbar.ax.set_ylabel(\"Pixel Intensity\")\n\n    plt.tight_layout()\n    plt.show()\n\nif __name__ == \"__main__\":\n    # Replace 'path_to_directory' with the actual path where your DICOM images are located\n    path_to_directory = \"/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/156/64005\"\n    dicom_images = load_dicom_images(path_to_directory)\n    visualize_dicom_images(dicom_images, num_rows=3, num_cols=3, window_level=40, window_width=80)\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-28T07:11:27.728951Z","iopub.execute_input":"2023-07-28T07:11:27.729356Z","iopub.status.idle":"2023-07-28T07:11:31.766203Z","shell.execute_reply.started":"2023-07-28T07:11:27.729323Z","shell.execute_reply":"2023-07-28T07:11:31.764916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Multi-Planar Reconstruction (MPR):\n\nMPR involves displaying slices in different planes (e.g., axial, sagittal, and coronal) simultaneously. You can use libraries like matplotlib or pyvista to create interactive MPR visualizations.\n\nMPR using matplotlib:","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\ndef display_mpr(dicom_images):\n    fig, axes = plt.subplots(1, 3, figsize=(12, 4))\n\n    # Display axial slice\n    axes[0].imshow(dicom_images[0].pixel_array, cmap='gray')\n    axes[0].set_title('Axial')\n\n    # Display sagittal slice\n    axes[1].imshow(dicom_images[1].pixel_array.T, cmap='gray', origin='lower')\n    axes[1].set_title('Sagittal')\n\n    # Display coronal slice\n    axes[2].imshow(dicom_images[2].pixel_array.T, cmap='gray', origin='lower')\n    axes[2].set_title('Coronal')\n\n    for ax in axes:\n        ax.axis('off')\n\n    plt.show()\n\n# Assuming you have 3 DICOM images for axial, sagittal, and coronal planes, respectively\n# Replace 'dicom_images' with your actual DICOM image data\ndisplay_mpr(dicom_images)","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:11:43.698056Z","iopub.execute_input":"2023-07-28T07:11:43.698489Z","iopub.status.idle":"2023-07-28T07:11:44.159342Z","shell.execute_reply.started":"2023-07-28T07:11:43.698453Z","shell.execute_reply":"2023-07-28T07:11:44.158176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Analysis of Organ Health</p></div>\n\n* Compare the prevalence of healthy and low/high health conditions for each organ (bowel, extravasation, kidney, liver, spleen).\n* Identify the organs with the highest and lowest health status.\n","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv')\n\n# List of organ columns\norgan_columns = ['bowel', 'extravasation', 'kidney', 'liver', 'spleen']\n\n# Initialize lists to store counts\nhealthy_count = [df[f'{organ}_healthy'].sum() for organ in organ_columns]\nlow_count = []\nhigh_count = []\n\n# Check if '_low' and '_high' columns exist before accessing them\nfor organ in organ_columns:\n    if f'{organ}_low' in df.columns:\n        low_count.append(df[f'{organ}_low'].sum())\n    else:\n        low_count.append(0)\n\n    if f'{organ}_high' in df.columns:\n        high_count.append(df[f'{organ}_high'].sum())\n    else:\n        high_count.append(0)\n        \n# Create the Plotly plot to compare organ health prevalence\nfig = go.Figure()\nfig.add_trace(go.Bar(x=organ_columns, y=healthy_count, name='Healthy'))\nfig.add_trace(go.Bar(x=organ_columns, y=low_count, name='Low Health'))\nfig.add_trace(go.Bar(x=organ_columns, y=high_count, name='High Health'))\n\nfig.update_layout(\n    title='Comparison of Organ Health Prevalence',\n    xaxis_title='Organ',\n    yaxis_title='Count',\n    barmode='stack',\n    showlegend=True,\n)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:11:51.752026Z","iopub.execute_input":"2023-07-28T07:11:51.752438Z","iopub.status.idle":"2023-07-28T07:11:51.782373Z","shell.execute_reply.started":"2023-07-28T07:11:51.752407Z","shell.execute_reply":"2023-07-28T07:11:51.781095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The plot above compares the prevalence of healthy, low health, and high health conditions for each organ (bowel, extravasation, kidney, liver, spleen) in the dataset. The bars represent the count of occurrences for each health status.\n\n**Observations:**\n\n- **Bowel:** All patients have healthy bowels (represented by the \"Healthy\" bar).\n- **Extravasation:** The majority of patients have healthy extravasation, with only one patient showing an injury (represented by the \"Healthy\" and \"Low Health\" bars).\n- **Kidney:** All patients have healthy kidneys (represented by the \"Healthy\" bar), and none have low or high health status.\n- **Liver:** Most patients have healthy livers, while one patient has both low and high health status (represented by the \"Healthy,\" \"Low Health,\" and \"High Health\" bars).\n- **Spleen:** The majority of patients have healthy spleens, with one patient having both low and high health status (represented by the \"Healthy,\" \"Low Health,\" and \"High Health\" bars).\n\nThe plot highlights that overall, the dataset predominantly consists of healthy organ conditions, with only a few instances of low or high health status. This may indicate that the dataset is biased toward healthy patients or that the study focuses on detecting injuries in predominantly healthy patients.\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Analysis of Injuries</p></div>\n\n* Analyze the occurrence of injuries in different organs.\n* Determine the overall prevalence of injuries in the dataset.\n* Investigate if there is any relationship between injuries in different organs.","metadata":{}},{"cell_type":"code","source":"df.columns","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:12:00.553938Z","iopub.execute_input":"2023-07-28T07:12:00.554328Z","iopub.status.idle":"2023-07-28T07:12:00.561759Z","shell.execute_reply.started":"2023-07-28T07:12:00.554299Z","shell.execute_reply":"2023-07-28T07:12:00.560405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Analyze the occurrence of injuries in different organs\norgan_columns = ['bowel_injury', 'extravasation_injury', 'any_injury']  # Include the '_injury' suffix\ninjury_counts = [df[column].sum() for column in organ_columns]\n\n# Determine the overall prevalence of injuries in the dataset\noverall_injury_count = df['any_injury'].sum()\n\n# Create the Plotly plot to visualize the occurrence of injuries in different organs\nfig = go.Figure()\nfig.add_trace(go.Bar(x=organ_columns, y=injury_counts))\nfig.update_layout(title='Occurrence of Injuries in Different Organs', xaxis_title='Organ', yaxis_title='Injury Count')\nfig.show()\n\n# Investigate the relationship between injuries in different organs\ninjury_relationship_df = df[organ_columns]  # Use only the 'organ_columns' for injury_relationship_df\ninjury_correlation_matrix = injury_relationship_df.corr()\n\n# Create the Plotly heatmap to visualize the correlation between injuries in different organs\nfig = go.Figure(data=go.Heatmap(\n    z=injury_correlation_matrix.values,\n    x=injury_correlation_matrix.columns,\n    y=injury_correlation_matrix.columns,\n    colorscale='Viridis',\n))\nfig.update_layout(title='Correlation Between Injuries in Different Organs')\nfig.show()\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-28T07:12:04.340451Z","iopub.execute_input":"2023-07-28T07:12:04.340847Z","iopub.status.idle":"2023-07-28T07:12:04.366523Z","shell.execute_reply.started":"2023-07-28T07:12:04.340815Z","shell.execute_reply":"2023-07-28T07:12:04.365321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. Occurrence of Injuries in Different Organs:\n\n* The bar plot displays the occurrence of injuries in different organs (bowel_injury,extravasation_injury,any_injury) in the dataset.\n* The injuries are represented by the height of the bars, and it shows that the bowel has the lowest number of injuries (0), while the other organs have one or more injuries each.\n\n2. Overall Prevalence of Injuries:\n\n* The variable \"any_injury\" indicates the overall prevalence of injuries in the dataset.\n* In this dataset, the overall injury count is 11 (sum of \"any_injury\" for all patients).\n\n3. Relationship Between Injuries in Different Organs:\n\n* The heatmap illustrates the correlation between injuries in different organs.\n* The diagonal of the heatmap shows the correlation of each organ's injuries with itself, which is always 1 (perfect correlation).\n* The off-diagonal elements indicate the correlations between injuries in different organs. In this small dataset, there may not be strong correlations between injuries in different organs.\n\nBased on the analysis and visualizations, we can observe the distribution of injuries in different organs and the overall prevalence of injuries in the dataset. The heatmap helps us identify any potential relationships between injuries in different organs. In a larger dataset, a more detailed analysis may reveal additional insights into the patterns and connections between injuries across various organs.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Any Injury vs. Organ Health</p></div>\n\n* Examine the relationship between the presence of \"any_injury\" and the health status of each organ.\n* Calculate the percentage of patients with any injury for each organ health category.","metadata":{}},{"cell_type":"code","source":"# Create a function to calculate the percentage of patients with any injury for each organ health category\ndef calculate_injury_percentage(organ_health_column, injury_column):\n    return df.groupby(organ_health_column)[injury_column].mean() * 100\n\n# Calculate the percentage of patients with any injury for each organ health category\norgan_columns = ['bowel', 'extravasation', 'kidney', 'liver', 'spleen']\ninjury_percentage = {\n    organ: calculate_injury_percentage(f'{organ}_healthy', 'any_injury') for organ in organ_columns\n}\n\n# Create the Plotly plot to visualize the relationship between \"any_injury\" and organ health\nfig = go.Figure()\nfor organ in organ_columns:\n    fig.add_trace(go.Bar(\n        x=injury_percentage[organ].index,\n        y=injury_percentage[organ],\n        name=organ.capitalize(),\n    ))\n\nfig.update_layout(\n    title='Percentage of Patients with Any Injury for Each Organ Health Category',\n    xaxis_title='Organ Health',\n    yaxis_title='Percentage of Patients with Any Injury',\n    showlegend=True,\n)\nfig.show()\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-07-28T07:12:17.940293Z","iopub.execute_input":"2023-07-28T07:12:17.940982Z","iopub.status.idle":"2023-07-28T07:12:17.969824Z","shell.execute_reply.started":"2023-07-28T07:12:17.940945Z","shell.execute_reply":"2023-07-28T07:12:17.968836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The bar plot above shows the percentage of patients with any injury for each organ health category (healthy, low health, and high health). Each bar represents an organ (bowel, extravasation, kidney, liver, spleen).\n\nObservations:\n\n* Bowel: 0% of patients with healthy bowels have any injury, while 100% of patients with low or high bowel health have injuries.\n* Extravasation: 100% of patients with healthy extravasation have injuries, while 0% of patients with low or high extravasation health have any injury.\n* Kidney: 100% of patients with healthy kidneys have injuries, while 0% of patients with low or high kidney health have any injury.\n* Liver: 0% of patients with healthy livers have any injury, while 100% of patients with low or high liver health have injuries.\n* Spleen: 100% of patients with healthy spleens have injuries, while 0% of patients with low or high spleen health have any injury.\n\nThe plot indicates a strong relationship between the presence of \"any_injury\" and the health status of each organ. For some organs, patients with healthy health status have no injuries, while patients with low or high health status have injuries. On the other hand, for other organs, patients with healthy health status have injuries, while patients with low or high health status do not have any injury.\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Patient Profiles</p></div>\n\n* Identify patient profiles based on their organ health and injury status.\n* Explore any patterns or common characteristics among patients with injuries.","metadata":{}},{"cell_type":"markdown","source":"To identify patient profiles based on their organ health and injury status, we will perform a clustering analysis. We'll use K-means clustering to group patients with similar organ health and injury patterns together. We'll then explore any patterns or common characteristics among patients with injuries using Plotly plots. Let's proceed with the analysis:","metadata":{}},{"cell_type":"code","source":"# Prepare the data for clustering\norgan_health_columns = ['bowel_healthy', 'extravasation_healthy', 'kidney_healthy', 'liver_healthy', 'spleen_healthy']\ninjury_columns = ['bowel_injury', 'extravasation_injury', 'kidney_high']  \npatient_data = df[organ_health_columns + injury_columns]\n\n# Standardize the data\nscaler = StandardScaler()\nscaled_data = scaler.fit_transform(patient_data)\n\n# Perform K-means clustering with 2 clusters (healthy and injured)\nkmeans = KMeans(n_clusters=2, random_state=42)\nclusters = kmeans.fit_predict(scaled_data)\n\n# Add the cluster labels to the DataFrame\ndf['cluster'] = clusters\n\n# Count the number of patients in each cluster\ncluster_counts = df['cluster'].value_counts()\n\n# Create the Plotly plot to visualize the patient profiles\nfig = go.Figure()\nfig.add_trace(go.Bar(x=['Healthy', 'Injured'], y=cluster_counts))\nfig.update_layout(\n    title='Patient Profiles Based on Organ Health and Injury Status',\n    xaxis_title='Patient Profile',\n    yaxis_title='Number of Patients',\n)\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:12:30.578227Z","iopub.execute_input":"2023-07-28T07:12:30.578632Z","iopub.status.idle":"2023-07-28T07:12:31.384961Z","shell.execute_reply.started":"2023-07-28T07:12:30.578602Z","shell.execute_reply":"2023-07-28T07:12:31.383702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The bar plot above shows the number of patients in each identified cluster (healthy and injured) based on their organ health and injury status.\n\nObservations:\n\n- **Cluster 0:** This cluster represents healthy patients with no injuries. In this small dataset, there is only one patient in this cluster.\n- **Cluster 1:** This cluster represents injured patients with at least one injury in their organs. In this small dataset, there are four patients in this cluster.\n\nSince the dataset is small, the clustering resulted in two clusters, one for healthy patients and the other for injured patients. In a larger dataset, more meaningful patient profiles may emerge, allowing for more in-depth exploration of common characteristics among patients with injuries.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Build Model and Prediction</p></div>\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport plotly.graph_objects as go\nfrom sklearn.datasets import make_classification\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVC\nfrom sklearn.metrics import accuracy_score, confusion_matrix\n","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:12:45.020064Z","iopub.execute_input":"2023-07-28T07:12:45.020535Z","iopub.status.idle":"2023-07-28T07:12:45.245506Z","shell.execute_reply.started":"2023-07-28T07:12:45.020504Z","shell.execute_reply":"2023-07-28T07:12:45.244147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a synthetic binary classification dataset\nn_samples = 500\nn_features = 2\nn_classes = 2\nn_clusters_per_class = 1\nrandom_state = 42\n\n# Adjust the values of n_informative, n_redundant, and n_repeated\nn_informative = 2\nn_redundant = 0\nn_repeated = 0\n\nX, y = make_classification(\n    n_samples=n_samples,\n    n_features=n_features,\n    n_informative=n_informative,\n    n_redundant=n_redundant,\n    n_repeated=n_repeated,\n    n_classes=n_classes,\n    n_clusters_per_class=n_clusters_per_class,\n    random_state=random_state\n)\n\n# Split the dataset into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=random_state)\n\n# Train a logistic regression model on the dataset\nmodel = LogisticRegression(random_state=random_state)\nmodel.fit(X_train, y_train)\n\n# Make predictions on the test set\ny_pred = model.predict(X_test)\n\n# Calculate accuracy and confusion matrix\naccuracy = accuracy_score(y_test, y_pred)\nconfusion_mat = confusion_matrix(y_test, y_pred)\n\nprint(f\"Accuracy: {accuracy}\")\nprint(\"Confusion Matrix:\")\nprint(confusion_mat)\n","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:12:48.993321Z","iopub.execute_input":"2023-07-28T07:12:48.993727Z","iopub.status.idle":"2023-07-28T07:12:49.019607Z","shell.execute_reply.started":"2023-07-28T07:12:48.993695Z","shell.execute_reply":"2023-07-28T07:12:49.018474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Random Forests model\nrf_model = RandomForestClassifier(random_state=42)\nrf_model.fit(X_train, y_train)\nrf_pred = rf_model.predict(X_test)\nrf_accuracy = accuracy_score(y_test, rf_pred)\n\n# SVM model\nsvm_model = SVC(kernel='linear', random_state=42)\nsvm_model.fit(X_train, y_train)\nsvm_pred = svm_model.predict(X_test)\nsvm_accuracy = accuracy_score(y_test, svm_pred)\n\n# Gradient Boosting model\ngb_model = GradientBoostingClassifier(random_state=42)\ngb_model.fit(X_train, y_train)\ngb_pred = gb_model.predict(X_test)\ngb_accuracy = accuracy_score(y_test, gb_pred)\n\n# Confusion matrix for each model\nrf_conf_matrix = confusion_matrix(y_test, rf_pred)\nsvm_conf_matrix = confusion_matrix(y_test, svm_pred)\ngb_conf_matrix = confusion_matrix(y_test, gb_pred)\n\n# Create a Plotly confusion matrix plot\ndef plot_confusion_matrix(matrix, title):\n    fig = go.Figure(data=go.Heatmap(\n        z=matrix,\n        x=['Predicted Negative', 'Predicted Positive'],\n        y=['True Negative', 'True Positive'],\n        colorscale='Viridis',\n    ))\n    fig.update_layout(title=title)\n    return fig\n\n# Plot confusion matrices\nrf_fig = plot_confusion_matrix(rf_conf_matrix, 'Random Forests Confusion Matrix')\nsvm_fig = plot_confusion_matrix(svm_conf_matrix, 'SVM Confusion Matrix')\ngb_fig = plot_confusion_matrix(gb_conf_matrix, 'Gradient Boosting Confusion Matrix')\n\n# Display model accuracies\nprint(f'Random Forests Accuracy: {rf_accuracy:.2f}')\nprint(f'SVM Accuracy: {svm_accuracy:.2f}')\nprint(f'Gradient Boosting Accuracy: {gb_accuracy:.2f}')\n\nrf_fig.show()\nsvm_fig.show()\ngb_fig.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:12:55.61466Z","iopub.execute_input":"2023-07-28T07:12:55.615806Z","iopub.status.idle":"2023-07-28T07:12:55.958165Z","shell.execute_reply.started":"2023-07-28T07:12:55.615765Z","shell.execute_reply":"2023-07-28T07:12:55.957096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Evaluation Metrics</p></div>\n","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import log_loss\n\n# Extract predicted probabilities for the 'any_injury' class\npredicted_injury_probabilities = df['bowel_injury']\n\n# Extract the actual true labels for the 'any_injury' class\nactual_injury_labels = df['any_injury']\n\n# Calculate the log loss for each patient\nlog_losses = log_loss(actual_injury_labels, predicted_injury_probabilities)\n\n# Add log losses to the DataFrame\ndf['log_loss'] = log_losses\n\nprint(df)","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:13:04.511076Z","iopub.execute_input":"2023-07-28T07:13:04.51465Z","iopub.status.idle":"2023-07-28T07:13:04.54097Z","shell.execute_reply.started":"2023-07-28T07:13:04.514605Z","shell.execute_reply":"2023-07-28T07:13:04.540069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Sort the DataFrame by log_loss in ascending order\ndf_sorted = df.sort_values(by='log_loss', ascending=True)\n\n# Create the Plotly bar plot\nfig = go.Figure()\n\nfig.add_trace(go.Bar(\n    x=df_sorted['patient_id'],\n    y=df_sorted['log_loss'],\n    marker_color='blue',\n    text=df_sorted['log_loss'].round(4),  # Display log losses with four decimal places as hover text\n    textposition='outside',  # Display the hover text outside the bars\n    name='Log Loss',  # Name for the legend\n))\n\n# Create the Plotly scatter plot\nfig.add_trace(go.Scatter(\n    x=df_sorted['patient_id'],\n    y=df_sorted['log_loss'],\n    mode='markers',\n    marker=dict(size=10, color='red'),\n    name='Scatter Plot',\n))\n\nfig.update_layout(\n    title='Log Loss for Each Patient',\n    xaxis_title='Patient ID',\n    yaxis_title='Log Loss',\n    showlegend=True,  # Show the legend on the plot\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:13:10.07352Z","iopub.execute_input":"2023-07-28T07:13:10.074363Z","iopub.status.idle":"2023-07-28T07:13:10.129258Z","shell.execute_reply.started":"2023-07-28T07:13:10.074317Z","shell.execute_reply":"2023-07-28T07:13:10.128168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/sample_submission.csv')\nsub.head().style.set_properties(**{'background-color':'royalblue','color':'white','border-color':'#8b8c8c'})","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:13:21.293441Z","iopub.execute_input":"2023-07-28T07:13:21.293858Z","iopub.status.idle":"2023-07-28T07:13:21.321705Z","shell.execute_reply.started":"2023-07-28T07:13:21.293825Z","shell.execute_reply":"2023-07-28T07:13:21.319653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the predicted probabilities for 'any_injury' using the log loss values\ndf['any_injury_predicted'] = 1 - log_losses\n\n#Extract the relevant columns for the submission dataset\nsub = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/sample_submission.csv')\n\n#Save the submission dataset to a file (e.g., CSV)\nsub.to_csv('submission.csv', index=False)\n\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-28T07:13:26.130187Z","iopub.execute_input":"2023-07-28T07:13:26.130899Z","iopub.status.idle":"2023-07-28T07:13:26.161425Z","shell.execute_reply.started":"2023-07-28T07:13:26.130865Z","shell.execute_reply":"2023-07-28T07:13:26.16019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\"> 📌 \"Hey there! Your positive feedback and support for my notebook mean the world to me! It motivates me to create more valuable content. If you can spare a moment to give it an upvote, it would help others discover and benefit from it too. Together, let's foster a vibrant community of knowledge-sharing and empowerment. Thank you for considering it, and continued success on your learning journey!\"😃</div>","metadata":{}}]}