{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"display:fill;\n            border-radius:15px;\n            background-color:#00bd35;\n            font-size:190%;\n            font-family:cursive;\n            letter-spacing:0.5px;\n            padding:10px;\n            color:white;\n            border-style: solid;\n            border-color: black;\n            text-align:center;\">\n<b>\nUnleashing the Healing Potential: Abdominal Trauma Detection</b>\n</div>\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Introduction</p></div>\n\nThe RSNA Abdominal Trauma Detection AI Challenge aims to address a critical issue in healthcare: the prompt and accurate diagnosis of traumatic injuries in the abdomen using computed tomography (CT) scans. Traumatic injury is a leading cause of death worldwide, and CT scans have become indispensable in evaluating patients with suspected abdominal injuries due to their ability to provide detailed cross-sectional images. However, interpreting CT scans for abdominal trauma can be complex and time-consuming, especially when multiple injuries or subtle active bleeding areas are present.\n\nTo tackle this challenge, the competition seeks to leverage the power of artificial intelligence and machine learning to assist medical professionals in rapidly and precisely detecting injuries and grading their severity. By developing advanced algorithms for this purpose, the goal is to improve trauma care and patient outcomes on a global scale.\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#FF7F50;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Dataset Overview:</p></div>\n\n\n\nThe dataset provided for the RSNA Abdominal Trauma Detection AI Challenge contains information related to patients and their abdominal health status. The dataset includes the following variables (features):\n\n1. **patient_id:** An identifier for each patient.\n2. **bowel_healthy:** Binary variable (0 or 1) indicating the health status of the bowel (0: not healthy, 1: healthy).\n3. **bowel_injury:** Binary variable (0 or 1) indicating the presence of injury in the bowel (0: no injury, 1: injury present).\n4. **extravasation_healthy:** Binary variable (0 or 1) indicating the health status of extravasation (0: not healthy, 1: healthy).\n5. **extravasation_injury:** Binary variable (0 or 1) indicating the presence of injury in extravasation (0: no injury, 1: injury present).\n6. **kidney_healthy:** Binary variable (0 or 1) indicating the health status of the kidneys (0: not healthy, 1: healthy).\n7. **kidney_low:** Binary variable (0 or 1) indicating low health status of the kidneys (0: not low, 1: low health).\n8. **kidney_high:** Binary variable (0 or 1) indicating high health status of the kidneys (0: not high, 1: high health).\n9. **liver_healthy:** Binary variable (0 or 1) indicating the health status of the liver (0: not healthy, 1: healthy).\n10. **liver_low:** Binary variable (0 or 1) indicating low health status of the liver (0: not low, 1: low health).\n11. **liver_high:** Binary variable (0 or 1) indicating high health status of the liver (0: not high, 1: high health).\n12. **spleen_healthy:** Binary variable (0 or 1) indicating the health status of the spleen (0: not healthy, 1: healthy).\n13. **spleen_low:** Binary variable (0 or 1) indicating low health status of the spleen (0: not low, 1: low health).\n14. **spleen_high:** Binary variable (0 or 1) indicating high health status of the spleen (0: not high, 1: high health).\n15. **any_injury:** An integer variable indicating the number of injuries detected (0: no injury, 1: one injury, 2: two injuries, and so on).\n\nThe dataset contains patient-specific health information for\ndifferent abdominal organs (bowel, extravasation, kidney, liver, spleen) and indicates whether injuries are present in these organs. Additionally, the \"any_injury\" variable provides an overall count of injuries detected in a patient. The dataset serves as the foundation for participants in the competition to develop AI models that can accurately detect severe injuries to the internal abdominal organs and any active internal bleeding.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Import Modules</p></div>\n","metadata":{}},{"cell_type":"code","source":"%%capture\n!pip install pydicom matplotlib","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:03:46.92737Z","iopub.execute_input":"2023-10-07T05:03:46.928185Z","iopub.status.idle":"2023-10-07T05:04:22.922619Z","shell.execute_reply.started":"2023-10-07T05:03:46.928148Z","shell.execute_reply":"2023-10-07T05:04:22.920956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport pydicom\nimport os\nimport numpy as np\nfrom matplotlib import pyplot as plt\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom sklearn.cluster import KMeans\nfrom sklearn.preprocessing import StandardScaler\n\nplt.rcParams['figure.figsize'] = (12,6)\nplt.style.use('fivethirtyeight')\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:04:28.157779Z","iopub.execute_input":"2023-10-07T05:04:28.158196Z","iopub.status.idle":"2023-10-07T05:04:30.087483Z","shell.execute_reply.started":"2023-10-07T05:04:28.158159Z","shell.execute_reply":"2023-10-07T05:04:30.086138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Load the Dataset</p></div>","metadata":{}},{"cell_type":"code","source":"# Load the dataset `into` a Pandas DataFrame\ndf = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv')\ndf.head().style.set_properties(**{'background-color':'royalblue','color':'white','border-color':'#8b8c8c'})","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:04:56.998253Z","iopub.execute_input":"2023-10-07T05:04:56.999184Z","iopub.status.idle":"2023-10-07T05:04:57.085925Z","shell.execute_reply.started":"2023-10-07T05:04:56.999128Z","shell.execute_reply":"2023-10-07T05:04:57.085186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read the CSV file into a DataFrame\ndf = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv')\n\n# List of column names\ncolumns = df.columns\n\n# Loop through columns and plot the data\nfor column in columns:\n    plt.figure()\n    plt.plot(df[column])\n    plt.title(column)\n    plt.show()\n\n# Set the background color, text color, and border color of the header row\ndf.head().style.set_properties(**{'background-color': 'royalblue', 'color': 'white', 'border-color': '#8b8c8c'})","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:07:54.988214Z","iopub.execute_input":"2023-10-07T05:07:54.988618Z","iopub.status.idle":"2023-10-07T05:07:59.837839Z","shell.execute_reply.started":"2023-10-07T05:07:54.98859Z","shell.execute_reply":"2023-10-07T05:07:59.836489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- background-color, color, and border-color properties are used to set the background color, text color, and border color of the DataFrame, respectively. \n- ** operator is used to unpack the dictionary of CSS properties into individual arguments for the set_properties() function.\n- effects of each CSS property:\n\n  - | CSS property | Effect |\n  - |---|---|---|\n  - | background-color | Sets the background color of the DataFrame. |\n  - | color | Sets the text color of the DataFrame. |\n  - | border-color | Sets the border color of the DataFrame. |","metadata":{}},{"cell_type":"code","source":"meta_df = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_series_meta.csv')\nmeta_df.head().style.set_properties(**{'background-color':'orange','color':'white','border-color':'#8b8c8c'})","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:08:14.098073Z","iopub.execute_input":"2023-10-07T05:08:14.098472Z","iopub.status.idle":"2023-10-07T05:08:14.120003Z","shell.execute_reply.started":"2023-10-07T05:08:14.098441Z","shell.execute_reply":"2023-10-07T05:08:14.118646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Data Preprocessing</p></div>\n\n* Check for missing values and handle them if necessary.\n* Check for data types and convert them if needed (e.g., converting binary data to boolean).\n* Identify and address any data quality issues, if present.\n","metadata":{}},{"cell_type":"code","source":"# Check for Missing Values\nmissing_values = df.isnull().sum()\nprint(\"Missing Values:\")\nprint(missing_values)\n\n# Check Data Types and Convert Binary Data to Boolean\nbinary_columns = [\n    'bowel_healthy', 'bowel_injury', 'extravasation_healthy', 'extravasation_injury',\n    'kidney_healthy', 'kidney_low', 'kidney_high', 'liver_healthy', 'liver_low', 'liver_high',\n    'spleen_healthy', 'spleen_low', 'spleen_high'\n]\ndf[binary_columns] = df[binary_columns].astype(bool)\n\n# Address Data Quality Issues\n# In this simple example, we assume no data quality issues are present.\n\n# Display the preprocessed DataFrame\nprint(\"\\nPreprocessed DataFrame:\")\nprint(df)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:09:40.278225Z","iopub.execute_input":"2023-10-07T05:09:40.27865Z","iopub.status.idle":"2023-10-07T05:09:40.311932Z","shell.execute_reply.started":"2023-10-07T05:09:40.278621Z","shell.execute_reply":"2023-10-07T05:09:40.310653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check for missing values\nmissing_values = df.isnull().sum()\nprint(\"Missing Values:\")\nprint(missing_values)\n\n# Handle missing values\n# In this simple example, we will drop rows with missing values.\ndf = df.dropna()\n\n# Check Data Types and Convert Binary Data to Boolean\nbinary_columns = [\n  'bowel_healthy', 'bowel_injury', 'extravasation_healthy', 'extravasation_injury',\n  'kidney_healthy', 'kidney_low', 'kidney_high', 'liver_healthy', 'liver_low', 'liver_high',\n  'spleen_healthy', 'spleen_low', 'spleen_high'\n]\ndf[binary_columns] = df[binary_columns].astype(bool)\n\n# Address Data Quality Issues\n# In this simple example, we assume no data quality issues are present.\n\nprint(\"\\nPreprocessed DataFrame:\")\nprint(df)\n\nplt.figure()\ndf.plot.hist()\nplt.title('Distribution of Features')\nplt.xlabel('Feature')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:09:49.058602Z","iopub.execute_input":"2023-10-07T05:09:49.059553Z","iopub.status.idle":"2023-10-07T05:09:49.463034Z","shell.execute_reply.started":"2023-10-07T05:09:49.059515Z","shell.execute_reply":"2023-10-07T05:09:49.461585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Descriptive Statistics</p></div>\n\n\n* Provide summary statistics for relevant variables (e.g., mean, median, standard deviation) to get an overview of the dataset.\n* Generate counts and percentages to show the distribution of categorical variables.","metadata":{}},{"cell_type":"code","source":"# Summary statistics for relevant variables\nstyled_data = df.describe().style\\\n.background_gradient(cmap='coolwarm')\\\n.set_properties(**{'text-align':'center','border':'1px solid black'})\n\n# display styled data\ndisplay(styled_data)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:09:55.80773Z","iopub.execute_input":"2023-10-07T05:09:55.808156Z","iopub.status.idle":"2023-10-07T05:09:55.835857Z","shell.execute_reply.started":"2023-10-07T05:09:55.808084Z","shell.execute_reply":"2023-10-07T05:09:55.835054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Generate styled data\nstyled_data = df.describe().style\\\n.background_gradient(cmap='coolwarm')\\\n.set_properties(**{'text-align':'center','border':'1px solid black'})\n\n# Loop over the styled data and plot\nfor i in range(len(styled_data.index)):\n    for j in range(len(styled_data.columns)):\n        cell_value = styled_data.data.iloc[i, j]\n\n        fig = plt.figure(figsize=(10, 6))\n        ax = fig.add_subplot(111)\n\n        ax.plot(cell_value)\n        ax.set_title(styled_data.index[i] + ', ' + styled_data.columns[j])\n        plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:10:04.847354Z","iopub.execute_input":"2023-10-07T05:10:04.847691Z","iopub.status.idle":"2023-10-07T05:10:08.817211Z","shell.execute_reply.started":"2023-10-07T05:10:04.847665Z","shell.execute_reply":"2023-10-07T05:10:08.815692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot styled data in a single plot, using subgrid layout\nimport matplotlib.gridspec as gridspec\n\n\nstyled_data = df.describe().style\\\n.background_gradient(cmap='coolwarm')\\\n.set_properties(**{'text-align':'center','border':'1px solid black'})\n\n# Cgridspec layout\ngs = gridspec.GridSpec(2, 2)\n\n# Loop over the styled data and plot it\nfig, axes = plt.subplots(2, 2, figsize=(12, 8), subplot_kw={'adjustable': 'box'})\nfor i in range(2):\n    for j in range(2):\n        cell_value = styled_data.data.iloc[i, j]\n\n        axes[i, j].plot(cell_value)\n        axes[i, j].set_title(styled_data.index[i] + ', ' + styled_data.columns[j])\n\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:14:58.181854Z","iopub.execute_input":"2023-10-07T05:14:58.182316Z","iopub.status.idle":"2023-10-07T05:14:59.096132Z","shell.execute_reply.started":"2023-10-07T05:14:58.182275Z","shell.execute_reply":"2023-10-07T05:14:59.094872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Generate counts and percentages for categorical variables\ncategorical_columns = [\n    'bowel_healthy', 'bowel_injury', 'extravasation_healthy', 'extravasation_injury',\n    'kidney_healthy', 'kidney_low', 'kidney_high', 'liver_healthy', 'liver_low', 'liver_high',\n    'spleen_healthy', 'spleen_low', 'spleen_high', 'any_injury'\n]\n\ncounts = df[categorical_columns].apply(pd.Series.value_counts)\npercentages = (counts / df.shape[0]) * 100\n\n# Display the summary statistics and counts/percentages\nprint(\"Summary Statistics:\")\n#print(summary_stats)\n\nprint(\"\\nCounts of Categorical Variables:\")\nprint(counts)\n\nprint(\"\\nPercentages of Categorical Variables (%):\")\nprint(percentages)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:15:28.433484Z","iopub.execute_input":"2023-10-07T05:15:28.433873Z","iopub.status.idle":"2023-10-07T05:15:28.468555Z","shell.execute_reply.started":"2023-10-07T05:15:28.433843Z","shell.execute_reply":"2023-10-07T05:15:28.467818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def generate_counts_and_percentages(df, categorical_columns):\n  \"\"\"Counts and percentages for categorical variables in a DataFrame.\n\n  Args:\n    df: DataFrame.\n    categorical_columns: column names for the categorical variables.\n\n  Returns:\n    A tuple of two Pandas DataFrames:\n      * counts: DataFrame containing the counts for each category of each categorical variable.\n      * percentages: DataFrame containing the percentages for each category of each categorical variable.\n  \"\"\"\n\n  # Handle null values.\n  df = df.dropna(subset=categorical_columns)\n\n  # Generate counts.\n  counts = df[categorical_columns].apply(pd.Series.value_counts)\n\n  # Generate percentages.\n  percentages = (counts / df.shape[0]) * 100\n\n  return counts, percentages\n\n# Example usage:\n\ncategorical_columns = [\n  'bowel_healthy', 'bowel_injury', 'extravasation_healthy', 'extravasation_injury',\n  'kidney_healthy', 'kidney_low', 'kidney_high', 'liver_healthy', 'liver_low', 'liver_high',\n  'spleen_healthy', 'spleen_low', 'spleen_high', 'any_injury'\n]\n\ncounts, percentages = generate_counts_and_percentages(df, categorical_columns)\n\n# Display the summary statistics and counts/percentages\nprint(\"Summary Statistics:\")\n#print(summary_stats)\n\nprint(\"\\nCounts of Categorical Variables:\")\nprint(counts)\n\nprint(\"\\nPercentages of Categorical Variables (%):\")\nprint(percentages)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:15:32.751184Z","iopub.execute_input":"2023-10-07T05:15:32.751573Z","iopub.status.idle":"2023-10-07T05:15:32.784753Z","shell.execute_reply.started":"2023-10-07T05:15:32.751545Z","shell.execute_reply":"2023-10-07T05:15:32.783823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\n\n# pass list of tick positions to the set_xticks() function. \n# pass the following list of tick positions to the set_xticks() function in the counts plot loop\n\ndef generate_counts_and_percentages(df, categorical_columns):\n  \"\"\"Counts and percentages for categorical variables in a DataFrame, and plot the counts and percentages.\n\n  Args:\n    df: DataFrame.\n    categorical_columns: column names for the categorical variables.\n\n  Returns:\n    None.\n  \"\"\"\n\n  # Handle null values.\n  df = df.dropna(subset=categorical_columns)\n\n  # counts.\n  counts = df[categorical_columns].apply(pd.Series.value_counts)\n\n  # percentages.\n  percentages = (counts / df.shape[0]) * 100\n\n  # Set color scheme.\n  colors = ['#007bff', '#ffa500']\n\n  # Plot counts.\n  fig, axes = plt.subplots(1, len(categorical_columns), figsize=(15, 7))\n  for i, column in enumerate(categorical_columns):\n    ax = axes[i]\n    ax.bar(counts.index.to_list(), counts[column].to_list(), color=colors[0])\n    ax.set_title(column, fontsize=12)\n    ax.set_xticks(range(len(counts.index)))\n    ax.tick_params(labelsize=10)\n    ax.grid(True)\n\n  # Plot percentages.\n  fig, axes = plt.subplots(1, len(categorical_columns), figsize=(15, 7))\n  for i, column in enumerate(categorical_columns):\n    ax = axes[i]\n    ax.pie(percentages[column].to_list(), labels=percentages.index.to_list(), autopct='%1.1f%%', startangle=140, colors=colors)\n    ax.set_title(column, fontsize=12)\n    ax.axis('equal')\n    ax.legend(fontsize=10)\n    ax.grid(True)\n\n  plt.suptitle('Counts and Percentages for Categorical Variables', fontsize=14)\n  plt.show()\n\ncategorical_columns = [\n  'bowel_healthy', 'bowel_injury', 'extravasation_healthy', 'extravasation_injury',\n  'kidney_healthy', 'kidney_low', 'kidney_high', 'liver_healthy', 'liver_low', 'liver_high',\n  'spleen_healthy', 'spleen_low', 'spleen_high', 'any_injury'\n]\n\ngenerate_counts_and_percentages(df, categorical_columns)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:18:05.951918Z","iopub.execute_input":"2023-10-07T05:18:05.95234Z","iopub.status.idle":"2023-10-07T05:18:09.06174Z","shell.execute_reply.started":"2023-10-07T05:18:05.95231Z","shell.execute_reply":"2023-10-07T05:18:09.060655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Exploratory Data Analysis (EDA)</p></div>\n\n\n\n* Visualize the distribution of different variables using appropriate plots (e.g., bar plots, pie charts, histograms).\n* Analyze the relationship between variables (e.g., correlation between organ health and injury status).\n* Explore any patterns or trends in the data.\n* Visualize a sample image","metadata":{}},{"cell_type":"code","source":"# Organ columns: bowel, extravasation, kidney, liver, spleen\norgan_columns = ['bowel', 'extravasation', 'kidney', 'liver', 'spleen']\n\n# Create a new DataFrame to store the counts\norgan_counts = pd.DataFrame()\norgan_counts['Organ'] = organ_columns\n\n# Loop through organ columns and count healthy and injury status for each organ\nfor organ in organ_columns:\n    healthy_col = f'{organ}_healthy'\n    injury_col = f'{organ}_injury'\n    \n    # Check if the columns exist in the DataFrame\n    if healthy_col in df.columns and injury_col in df.columns:\n        organ_counts[f'{organ}_healthy'] = df[healthy_col].sum()\n        organ_counts[f'{organ}_injury'] = df[injury_col].sum()\n    else:\n        # Handle the case if the columns are missing\n        print(f\"Warning: Columns for {organ} healthy/injury status are missing in the DataFrame.\")\n        organ_counts[f'{organ}_healthy'] = 0\n        organ_counts[f'{organ}_injury'] = 0\n\n# Melt the DataFrame to have a single 'Status' column\norgan_counts_melted = organ_counts.melt(id_vars=['Organ'], var_name='Status', value_name='Count')\n\n# Bar plot for distribution of organ health and injury status\nfig = px.bar(\n    organ_counts_melted,\n    x='Organ',\n    y='Count',\n    color='Status',\n    barmode='group',\n    labels=dict(x='Organ', y='Count', Status='Status'),\n    title='Distribution of Organ Health and Injury Status',\n)\nfig.update_layout(showlegend=True)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:18:29.351655Z","iopub.execute_input":"2023-10-07T05:18:29.352121Z","iopub.status.idle":"2023-10-07T05:18:31.34362Z","shell.execute_reply.started":"2023-10-07T05:18:29.352069Z","shell.execute_reply":"2023-10-07T05:18:31.3423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Organ columns\norgan_columns = ['bowel', 'extravasation', 'kidney', 'liver', 'spleen']\n\n# Create a new DataFrame to store the counts\norgan_counts = pd.DataFrame()\norgan_counts['Organ'] = organ_columns\n\n# Loop through organ columns and count healthy and injury status for each organ\nfor organ in organ_columns:\n    healthy_col = f'{organ}_healthy'\n    injury_col = f'{organ}_injury'\n\n    # Check if the columns exist in the DataFrame\n    if healthy_col in df.columns and injury_col in df.columns:\n        organ_counts[f'{organ}_healthy'] = df[healthy_col].sum()\n        organ_counts[f'{organ}_injury'] = df[injury_col].sum()\n    else:\n        # Handle the case if the columns are missing\n        print(f\"Warning: Columns for {organ} healthy/injury status are missing in the DataFrame.\")\n        organ_counts[f'{organ}_healthy'] = 0\n        organ_counts[f'{organ}_injury'] = 0\n\n# Fill in missing values with 0\norgan_counts.fillna(0, inplace=True)\n\n# Melt the DataFrame to have a single 'Status' column\norgan_counts_melted = organ_counts.melt(id_vars=['Organ'], var_name='Status', value_name='Count')\n\n# Bar plot for distribution of organ health and injury status\nfig = px.bar(\n    organ_counts_melted,\n    x='Organ',\n    y='Count',\n    color='Status',\n    barmode='group',\n    labels=dict(x='Organ', y='Count', Status='Status'),\n    title='Distribution of Organ Health and Injury Status',\n    height=500,\n    width=800,\n    template='plotly_dark',\n)\n\n# Customize the plot\nfig.update_layout(\n    legend_title='Organ Status',\n    legend_orientation='h',\n    legend_xanchor='center',\n    legend_yanchor='top',\n    legend_x=0.5,\n    legend_y=1.1,\n    xaxis_title='Organ',\n    yaxis_title='Count',\n    font=dict(family='Arial', size=12),\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:19:15.761602Z","iopub.execute_input":"2023-10-07T05:19:15.762031Z","iopub.status.idle":"2023-10-07T05:19:15.92867Z","shell.execute_reply.started":"2023-10-07T05:19:15.762Z","shell.execute_reply":"2023-10-07T05:19:15.927334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Organ columns: bowel, extravasation, kidney, liver, spleen\norgan_columns = ['bowel', 'extravasation', 'kidney', 'liver', 'spleen']\n\n# Check if the 'injury' columns are present in the DataFrame\ninjury_columns = [f'{organ}_injury' for organ in organ_columns]\nmissing_columns = set(organ_columns + injury_columns) - set(df.columns)\n\nif missing_columns:\n    # Handle the case if any of the required columns are missing\n    print(f\"Warning: Columns for {', '.join(missing_columns)} are missing in the DataFrame.\")\n    for col in missing_columns:\n        df[col] = 0\n\n# Heatmap to analyze the correlation between organ health and injury status\ncorrelation_df = df[organ_columns + injury_columns]\ncorrelation_matrix = correlation_df.corr()\n\nfig = px.imshow(\n    correlation_matrix,\n    x=correlation_df.columns,\n    y=correlation_df.columns,\n    labels=dict(x='Organ', y='Organ', color='Correlation'),\n    title='Correlation Between Organ Health and Injury Status',\n)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:19:51.321185Z","iopub.execute_input":"2023-10-07T05:19:51.321594Z","iopub.status.idle":"2023-10-07T05:19:51.433909Z","shell.execute_reply.started":"2023-10-07T05:19:51.321566Z","shell.execute_reply":"2023-10-07T05:19:51.433037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- color_continuous_scale argument to RdYlGn, which is a more visually appealing color scale for heatmaps.\n- aspect='equal' argument to make the heatmap square. This is more visually appealing and makes it easier to compare the values on the x- and y-axes.\n- text_auto=True argument to add annotations to the heatmap. This makes it easier to see the exact values of the correlations.\n- fig.update_traces() function allows to update the properties of all traces in the figure. Update the textfont_size property to 10.","metadata":{}},{"cell_type":"code","source":"import plotly.express as px\n\n# Organ columns\norgan_columns = ['bowel', 'extravasation', 'kidney', 'liver', 'spleen']\n\n# Check if the 'injury' columns are present in the DataFrame\ninjury_columns = [f'{organ}_injury' for organ in organ_columns]\nmissing_columns = set(organ_columns + injury_columns) - set(df.columns)\n\n# Handle the case if any of the required columns are missing\nif missing_columns:\n  print(f\"Warning: Columns for {', '.join(missing_columns)} are missing in the DataFrame.\")\n  for col in missing_columns:\n    df[col] = 0\n\n# Create a correlation matrix\ncorrelation_matrix = df[organ_columns + injury_columns].corr()\n\n# Customize the heatmap\nfig = px.imshow(\n    correlation_matrix,\n    x=correlation_matrix.columns,\n    y=correlation_matrix.columns,\n    labels=dict(x='Organ', y='Organ', color='Correlation'),\n    title='Correlation Between Organ Health and Injury Status',\n    color_continuous_scale='RdYlGn',  # A more visually appealing color scale\n    aspect='equal',  # Make the heatmap square\n    text_auto=True,  # Add annotations to the heatmap\n)\n\n# Update the annotation font size\nfig.update_traces(textfont_size=10)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:20:04.311224Z","iopub.execute_input":"2023-10-07T05:20:04.311587Z","iopub.status.idle":"2023-10-07T05:20:04.386925Z","shell.execute_reply.started":"2023-10-07T05:20:04.31156Z","shell.execute_reply":"2023-10-07T05:20:04.385638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. Occurrence of Injuries in Different Organs:\n\n* The bar plot displays the occurrence of injuries in different organs (bowel, extravasation, kidney, liver, spleen) in the dataset.\n* The injuries are represented by the height of the bars, and it shows that the bowel has the lowest number of injuries (0), while the other organs have one or more injuries each.\n\n2. Overall Prevalence of Injuries:\n\n* The variable \"any_injury\" indicates the overall prevalence of injuries in the dataset.\n* In this dataset, the overall injury count is 11 (sum of \"any_injury\" for all patients).\n\n3. Relationship Between Injuries in Different Organs:\n\n* The heatmap illustrates the correlation between injuries in different organs.\n* The diagonal of the heatmap shows the correlation of each organ's injuries with itself, which is always 1 (perfect correlation).\n* The off-diagonal elements indicate the correlations between injuries in different organs. In this small dataset, there may not be strong correlations between injuries in different organs.\n\nBased on the analysis and visualizations, we can observe the distribution of injuries in different organs and the overall prevalence of injuries in the dataset. The heatmap helps us identify any potential relationships between injuries in different organs. In a larger dataset, a more detailed analysis may reveal additional insights into the patterns and connections between injuries across various organs.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Visualize a sample image</p></div>\n\n","metadata":{}},{"cell_type":"code","source":"import pydicom","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:20:34.581317Z","iopub.execute_input":"2023-10-07T05:20:34.581731Z","iopub.status.idle":"2023-10-07T05:20:34.587968Z","shell.execute_reply.started":"2023-10-07T05:20:34.581701Z","shell.execute_reply":"2023-10-07T05:20:34.586401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def standardize_pixel_array(dcm: pydicom.dataset.FileDataset) -> np.ndarray:\n    # Correct DICOM pixel_array if PixelRepresentation == 1.\n    pixel_array = dcm.pixel_array\n    if dcm.PixelRepresentation == 1:\n        bit_shift = dcm.BitsAllocated - dcm.BitsStored\n        dtype = pixel_array.dtype\n        new_array = (pixel_array << bit_shift).astype(dtype) >>  bit_shift\n        pixel_array = pydicom.pixel_data_handlers.util.apply_modality_lut(new_array, dcm)\n    return pixel_array\n\ntrain_dicom_tags = pd.read_parquet('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_dicom_tags.parquet', engine='pyarrow')","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:20:34.851858Z","iopub.execute_input":"2023-10-07T05:20:34.852314Z","iopub.status.idle":"2023-10-07T05:20:41.773954Z","shell.execute_reply.started":"2023-10-07T05:20:34.852282Z","shell.execute_reply":"2023-10-07T05:20:41.772761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# path of the image in index 49954\nsample_image = train_dicom_tags.loc[49954]['path']\n\n# Open the DICOM file using pydicom.\ndcm = pydicom.read_file(os.path.join('/kaggle/input/rsna-2023-abdominal-trauma-detection',sample_image))","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:20:48.760981Z","iopub.execute_input":"2023-10-07T05:20:48.761836Z","iopub.status.idle":"2023-10-07T05:20:48.791342Z","shell.execute_reply.started":"2023-10-07T05:20:48.761762Z","shell.execute_reply":"2023-10-07T05:20:48.790024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dicom_tags.head(3)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:20:51.140878Z","iopub.execute_input":"2023-10-07T05:20:51.141357Z","iopub.status.idle":"2023-10-07T05:20:51.167727Z","shell.execute_reply.started":"2023-10-07T05:20:51.141318Z","shell.execute_reply":"2023-10-07T05:20:51.166503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualize a sample image\ndef plot_dicom_image(image_path):\n    ds = pydicom.dcmread(image_path)\n    plt.imshow(ds.pixel_array, cmap=plt.cm.bone)\n    plt.axis('off')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:20:53.841092Z","iopub.execute_input":"2023-10-07T05:20:53.841555Z","iopub.status.idle":"2023-10-07T05:20:53.847559Z","shell.execute_reply.started":"2023-10-07T05:20:53.841523Z","shell.execute_reply":"2023-10-07T05:20:53.846193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_dicom_image2(image_path, figsize=(10, 10), window_center=40, window_width=80):\n  \"\"\"Plot DICOM image using matplotlib.pyplot, with windowing applied.\n\n  Args:\n    image_path: The path to the DICOM image file.\n    figsize: The size of the figure in inches.\n    window_center: The window center value.\n    window_width: The window width value.\n  \"\"\"\n\n  ds = pydicom.dcmread(image_path)\n\n  # Check if the image is windowed.\n  if ds.WindowCenter and ds.WindowWidth:\n    # Apply the windowing.\n    image = ds.pixel_array * (ds.WindowWidth / 10.0) + ds.WindowCenter\n  else:\n    image = ds.pixel_array\n\n  # Create a new figure and plot the image.\n  fig, ax = plt.subplots(1, 1, figsize=figsize)\n  ax.imshow(image, cmap=plt.cm.bone)\n  ax.axis('off')\n\n  # Add a title to the figure with the patient's name.\n  # If the PatientName attribute is not present, use an empty string.\n  patient_name = ds.get('PatientName', '')\n  ax.set_title(patient_name)\n\n  plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:20:56.621091Z","iopub.execute_input":"2023-10-07T05:20:56.621544Z","iopub.status.idle":"2023-10-07T05:20:56.629395Z","shell.execute_reply.started":"2023-10-07T05:20:56.621514Z","shell.execute_reply":"2023-10-07T05:20:56.628187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_image_path = '/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/49954/41479/378.dcm'\nplot_dicom_image(sample_image_path)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:21:04.95222Z","iopub.execute_input":"2023-10-07T05:21:04.952618Z","iopub.status.idle":"2023-10-07T05:21:05.155959Z","shell.execute_reply.started":"2023-10-07T05:21:04.952587Z","shell.execute_reply":"2023-10-07T05:21:05.154711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_dicom_image2('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/10007/47578/100.dcm', window_center=40, window_width=80)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:21:06.760951Z","iopub.execute_input":"2023-10-07T05:21:06.761313Z","iopub.status.idle":"2023-10-07T05:21:07.18389Z","shell.execute_reply.started":"2023-10-07T05:21:06.761287Z","shell.execute_reply":"2023-10-07T05:21:07.183176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_dicom_images(directory):\n    dicom_images = []\n    for filename in os.listdir(directory):\n        if filename.endswith(\".dcm\"):\n            dicom_file = os.path.join(directory, filename)\n            dicom_image = pydicom.dcmread(dicom_file)\n            dicom_images.append(dicom_image)\n    return dicom_images\n\ndef rescale_pixel_array(pixel_array, window_level, window_width):\n    # Rescale the pixel values based on the window level and window width\n    min_value = window_level - window_width // 2\n    max_value = window_level + window_width // 2\n    rescaled_pixel_array = np.clip(pixel_array, min_value, max_value)\n    rescaled_pixel_array = (rescaled_pixel_array - min_value) / (max_value - min_value)\n    return rescaled_pixel_array\n\ndef visualize_dicom_images(dicom_images, num_rows=4, num_cols=4, window_level=40, window_width=80):\n    fig, axes = plt.subplots(num_rows, num_cols, figsize=(15, 15))\n    for i, ax in enumerate(axes.flat):\n        if i < len(dicom_images):\n            dicom_image = dicom_images[i]\n            image_data = dicom_image.pixel_array.astype(np.float32)\n            rescaled_image = rescale_pixel_array(image_data, window_level, window_width)\n            ax.imshow(rescaled_image, cmap=plt.cm.bone)\n            ax.axis(\"off\")\n            ax.set_title(f\"Slice {i+1}\")\n\n    # Hide any empty subplots\n    for i in range(len(dicom_images), num_rows*num_cols):\n        axes.flat[i].axis(\"off\")\n\n    # Add a color bar to indicate pixel intensity values\n    cax = fig.add_axes([0.92, 0.15, 0.02, 0.7])\n    norm = plt.cm.colors.Normalize(vmin=0, vmax=1)\n    cbar = plt.colorbar(plt.cm.ScalarMappable(norm=norm, cmap=plt.cm.bone), cax=cax)\n    cbar.ax.set_ylabel(\"Pixel Intensity\")\n\n    plt.tight_layout()\n    plt.show()\n\nif __name__ == \"__main__\":\n    # Replace 'path_to_directory' with the actual path where your DICOM images are located\n    path_to_directory = \"/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/49954/41479\"\n    dicom_images = load_dicom_images(path_to_directory)\n    visualize_dicom_images(dicom_images, num_rows=3, num_cols=3, window_level=40, window_width=80)\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-07T05:21:11.172509Z","iopub.execute_input":"2023-10-07T05:21:11.172864Z","iopub.status.idle":"2023-10-07T05:21:19.505425Z","shell.execute_reply.started":"2023-10-07T05:21:11.172839Z","shell.execute_reply":"2023-10-07T05:21:19.504277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_image_path = '/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/156/64005/106.dcm'\nplot_dicom_image(sample_image_path)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:21:33.991949Z","iopub.execute_input":"2023-10-07T05:21:33.992911Z","iopub.status.idle":"2023-10-07T05:21:34.241008Z","shell.execute_reply.started":"2023-10-07T05:21:33.992868Z","shell.execute_reply":"2023-10-07T05:21:34.239681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_dicom_images(directory):\n    dicom_images = []\n    for filename in os.listdir(directory):\n        if filename.endswith(\".dcm\"):\n            dicom_file = os.path.join(directory, filename)\n            dicom_image = pydicom.dcmread(dicom_file)\n            dicom_images.append(dicom_image)\n    return dicom_images\n\ndef rescale_pixel_array(pixel_array, window_level, window_width):\n    # Rescale the pixel values based on the window level and window width\n    min_value = window_level - window_width // 2\n    max_value = window_level + window_width // 2\n    rescaled_pixel_array = np.clip(pixel_array, min_value, max_value)\n    rescaled_pixel_array = (rescaled_pixel_array - min_value) / (max_value - min_value)\n    return rescaled_pixel_array\n\ndef visualize_dicom_images(dicom_images, num_rows=4, num_cols=4, window_level=40, window_width=80):\n    fig, axes = plt.subplots(num_rows, num_cols, figsize=(15, 15))\n    for i, ax in enumerate(axes.flat):\n        if i < len(dicom_images):\n            dicom_image = dicom_images[i]\n            image_data = dicom_image.pixel_array.astype(np.float32)\n            rescaled_image = rescale_pixel_array(image_data, window_level, window_width)\n            ax.imshow(rescaled_image, cmap=plt.cm.bone)\n            ax.axis(\"off\")\n            ax.set_title(f\"Slice {i+1}\")\n\n    # Hide any empty subplots\n    for i in range(len(dicom_images), num_rows*num_cols):\n        axes.flat[i].axis(\"off\")\n\n    # Add a color bar to indicate pixel intensity values\n    cax = fig.add_axes([0.92, 0.15, 0.02, 0.7])\n    norm = plt.cm.colors.Normalize(vmin=0, vmax=1)\n    cbar = plt.colorbar(plt.cm.ScalarMappable(norm=norm, cmap=plt.cm.bone), cax=cax)\n    cbar.ax.set_ylabel(\"Pixel Intensity\")\n\n    plt.tight_layout()\n    plt.show()\n\nif __name__ == \"__main__\":\n    # Replace 'path_to_directory' with the actual path of DICOM images\n    path_to_directory = \"/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/156/64005\"\n    dicom_images = load_dicom_images(path_to_directory)\n    visualize_dicom_images(dicom_images, num_rows=3, num_cols=3, window_level=40, window_width=80)\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-07T05:22:19.382414Z","iopub.execute_input":"2023-10-07T05:22:19.382868Z","iopub.status.idle":"2023-10-07T05:22:21.216938Z","shell.execute_reply.started":"2023-10-07T05:22:19.382835Z","shell.execute_reply":"2023-10-07T05:22:21.21548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Multi-Planar Reconstruction (MPR):\n\nMPR involves displaying slices in different planes (e.g., axial, sagittal, and coronal) simultaneously. You can use libraries like matplotlib or pyvista to create interactive MPR visualizations.\n\nMPR using matplotlib:","metadata":{}},{"cell_type":"code","source":"# !pip install pyvista","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-10-07T05:28:50.711566Z","iopub.execute_input":"2023-10-07T05:28:50.712012Z","iopub.status.idle":"2023-10-07T05:28:50.719557Z","shell.execute_reply.started":"2023-10-07T05:28:50.711978Z","shell.execute_reply":"2023-10-07T05:28:50.718168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# import pyvista\n# import matplotlib.pyplot as plt\n\n# def display_mpr(dicom_images):\n#     fig, axes = plt.subplots(1, 3, figsize=(12, 4))\n\n#     # Display axial slice\n#     axes[0].imshow(dicom_images[0].pixel_array, cmap='gray')\n#     axes[0].set_title('Axial')\n\n#     # Display sagittal slice\n#     axes[1].imshow(dicom_images[1].pixel_array.T, cmap='gray', origin='lower')\n#     axes[1].set_title('Sagittal')\n\n#     # Display coronal slice\n#     axes[2].imshow(dicom_images[2].pixel_array.T, cmap='gray', origin='lower')\n#     axes[2].set_title('Coronal')\n\n#     for ax in axes:\n#         ax.axis('off')\n\n#     plt.show()\n\n# # Load the DICOM images\n# axial_image = pyvista.read('axial.dcm')\n# sagittal_image = pyvista.read('sagittal.dcm')\n# coronal_image = pyvista.read('coronal.dcm')\n\n# # MPR visualization\n# display_mpr([axial_image, sagittal_image, coronal_image])\n","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:28:57.676679Z","iopub.execute_input":"2023-10-07T05:28:57.677051Z","iopub.status.idle":"2023-10-07T05:28:57.68301Z","shell.execute_reply.started":"2023-10-07T05:28:57.677023Z","shell.execute_reply":"2023-10-07T05:28:57.681823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\ndef display_mpr(dicom_images):\n    fig, axes = plt.subplots(1, 3, figsize=(12, 4))\n\n    # Display axial slice\n    axes[0].imshow(dicom_images[0].pixel_array, cmap='gray')\n    axes[0].set_title('Axial')\n\n    # Display sagittal slice\n    axes[1].imshow(dicom_images[1].pixel_array.T, cmap='gray', origin='lower')\n    axes[1].set_title('Sagittal')\n    \n\n    # Display coronal slice\n    axes[2].imshow(dicom_images[2].pixel_array.T, cmap='gray', origin='lower')\n    axes[2].set_title('Coronal')\n\n    for ax in axes:\n        ax.axis('off')\n\n    plt.show()\n\n# Assuming you have 3 DICOM images for axial, sagittal, and coronal planes, respectively\n# Replace 'dicom_images' with your actual DICOM image data\ndisplay_mpr(dicom_images)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:22:53.076073Z","iopub.execute_input":"2023-10-07T05:22:53.076573Z","iopub.status.idle":"2023-10-07T05:22:53.527039Z","shell.execute_reply.started":"2023-10-07T05:22:53.076535Z","shell.execute_reply":"2023-10-07T05:22:53.526203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Analysis of Organ Health</p></div>\n\n* Compare the prevalence of healthy and low/high health conditions for each organ (bowel, extravasation, kidney, liver, spleen).\n* Identify the organs with the highest and lowest health status.\n","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv')\n\n# List of organ columns\norgan_columns = ['bowel', 'extravasation', 'kidney', 'liver', 'spleen']\n\n# Initialize lists to store counts\nhealthy_count = [df[f'{organ}_healthy'].sum() for organ in organ_columns]\nlow_count = []\nhigh_count = []\n\n# Check if '_low' and '_high' columns exist before accessing them\nfor organ in organ_columns:\n    if f'{organ}_low' in df.columns:\n        low_count.append(df[f'{organ}_low'].sum())\n    else:\n        low_count.append(0)\n\n    if f'{organ}_high' in df.columns:\n        high_count.append(df[f'{organ}_high'].sum())\n    else:\n        high_count.append(0)\n        \n# Create the Plotly plot to compare organ health prevalence\nfig = go.Figure()\nfig.add_trace(go.Bar(x=organ_columns, y=healthy_count, name='Healthy'))\nfig.add_trace(go.Bar(x=organ_columns, y=low_count, name='Low Health'))\nfig.add_trace(go.Bar(x=organ_columns, y=high_count, name='High Health'))\n\nfig.update_layout(\n    title='Comparison of Organ Health Prevalence',\n    xaxis_title='Organ',\n    yaxis_title='Count',\n    barmode='stack',\n    showlegend=True,\n)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:29:02.386909Z","iopub.execute_input":"2023-10-07T05:29:02.387305Z","iopub.status.idle":"2023-10-07T05:29:02.421827Z","shell.execute_reply.started":"2023-10-07T05:29:02.38727Z","shell.execute_reply":"2023-10-07T05:29:02.420986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport plotly.graph_objects as go\n\n# Define type hints\norgan_columns: list[str] = ['bowel', 'extravasation', 'kidney', 'liver', 'spleen']\nhealthy_count: list[int] = []\nlow_count: list[int] = []\nhigh_count: list[int] = []\n\n# Read the CSV file\ndf = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv')\n\n# Check if '_low' and '_high' columns exist before accessing them\nfor organ in organ_columns:\n    if f'{organ}_low' in df.columns:\n        low_count.append(df[f'{organ}_low'].sum())\n    else:\n        low_count.append(0)\n\n    if f'{organ}_high' in df.columns:\n        high_count.append(df[f'{organ}_high'].sum())\n    else:\n        high_count.append(0)\n\n# Compare organ health prevalence\nfig = go.Figure()\nfig.add_trace(go.Bar(x=organ_columns, y=healthy_count, name='Healthy'))\nfig.add_trace(go.Bar(x=organ_columns, y=low_count, name='Low Health'))\nfig.add_trace(go.Bar(x=organ_columns, y=high_count, name='High Health'))\n\nfig.update_layout(\n    title='Comparison of Organ Health Prevalence',\n    xaxis_title='Organ',\n    yaxis_title='Count',\n    barmode='stack',\n    showlegend=True,\n)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:29:05.541117Z","iopub.execute_input":"2023-10-07T05:29:05.541519Z","iopub.status.idle":"2023-10-07T05:29:05.566941Z","shell.execute_reply.started":"2023-10-07T05:29:05.541481Z","shell.execute_reply":"2023-10-07T05:29:05.566197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The plot above compares the prevalence of healthy, low health, and high health conditions for each organ (bowel, extravasation, kidney, liver, spleen) in the dataset. The bars represent the count of occurrences for each health status.\n\n**Observations:**\n\n- **Bowel:** All patients have healthy bowels (represented by the \"Healthy\" bar).\n- **Extravasation:** The majority of patients have healthy extravasation, with only one patient showing an injury (represented by the \"Healthy\" and \"Low Health\" bars).\n- **Kidney:** All patients have healthy kidneys (represented by the \"Healthy\" bar), and none have low or high health status.\n- **Liver:** Most patients have healthy livers, while one patient has both low and high health status (represented by the \"Healthy,\" \"Low Health,\" and \"High Health\" bars).\n- **Spleen:** The majority of patients have healthy spleens, with one patient having both low and high health status (represented by the \"Healthy,\" \"Low Health,\" and \"High Health\" bars).\n\nThe plot highlights that overall, the dataset predominantly consists of healthy organ conditions, with only a few instances of low or high health status. This may indicate that the dataset is biased toward healthy patients or that the study focuses on detecting injuries in predominantly healthy patients.\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Analysis of Injuries</p></div>\n\n* Analyze the occurrence of injuries in different organs.\n* Determine the overall prevalence of injuries in the dataset.\n* Investigate if there is any relationship between injuries in different organs.","metadata":{}},{"cell_type":"code","source":"df.columns","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:29:13.411312Z","iopub.execute_input":"2023-10-07T05:29:13.411753Z","iopub.status.idle":"2023-10-07T05:29:13.419853Z","shell.execute_reply.started":"2023-10-07T05:29:13.411723Z","shell.execute_reply":"2023-10-07T05:29:13.418949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Analyze the occurrence of injuries in different organs\norgan_columns = ['bowel_injury', 'extravasation_injury', 'any_injury']  # Include the '_injury' suffix\ninjury_counts = [df[column].sum() for column in organ_columns]\n\n# Determine the overall prevalence of injuries in the dataset\noverall_injury_count = df['any_injury'].sum()\n\n# Create the Plotly plot to visualize the occurrence of injuries in different organs\nfig = go.Figure()\nfig.add_trace(go.Bar(x=organ_columns, y=injury_counts))\nfig.update_layout(title='Occurrence of Injuries in Different Organs', xaxis_title='Organ', yaxis_title='Injury Count')\nfig.show()\n\n# Investigate the relationship between injuries in different organs\ninjury_relationship_df = df[organ_columns]  # Use only the 'organ_columns' for injury_relationship_df\ninjury_correlation_matrix = injury_relationship_df.corr()\n\n# Create the Plotly heatmap to visualize the correlation between injuries in different organs\nfig = go.Figure(data=go.Heatmap(\n    z=injury_correlation_matrix.values,\n    x=injury_correlation_matrix.columns,\n    y=injury_correlation_matrix.columns,\n    colorscale='Viridis',\n))\nfig.update_layout(title='Correlation Between Injuries in Different Organs')\nfig.show()\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-07T05:29:14.176055Z","iopub.execute_input":"2023-10-07T05:29:14.176518Z","iopub.status.idle":"2023-10-07T05:29:14.204663Z","shell.execute_reply.started":"2023-10-07T05:29:14.176484Z","shell.execute_reply":"2023-10-07T05:29:14.203841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import plotly.graph_objects as go\n\n# Define the data\norgan_columns = ['bowel_injury', 'extravasation_injury', 'any_injury']\ninjury_counts = [df[column].sum() for column in organ_columns]\n\n# Visualize the occurrence of injuries in different organs\nfig = go.Figure()\nfig.add_trace(go.Bar(x=organ_columns, y=injury_counts, name='Injury Count'))\nfig.update_layout(\n    title='Occurrence of Injuries in Different Organs',\n    xaxis_title='Organ',\n    yaxis_title='Number of Injuries',\n    font_size=12,\n    legend=dict(orientation='h'),\n    plot_bgcolor='#ffffff',\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:29:18.081145Z","iopub.execute_input":"2023-10-07T05:29:18.081517Z","iopub.status.idle":"2023-10-07T05:29:18.100781Z","shell.execute_reply.started":"2023-10-07T05:29:18.08149Z","shell.execute_reply":"2023-10-07T05:29:18.099338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import plotly.graph_objects as go\n\n# Define the data\ninjury_correlation_matrix = injury_relationship_df.corr()\n\n# Convert the injury_correlation_matrix data frame to a 2D NumPy array\ninjury_correlation_matrix_array = np.asarray(injury_correlation_matrix)\n\n# Visualize the correlation between injuries in different organs\nfig = go.Figure(data=go.Heatmap(\n  z=injury_correlation_matrix_array,\n  x=injury_correlation_matrix.columns,\n  y=injury_correlation_matrix.columns,\n  colorscale='Viridis',\n))\nfig.update_layout(\n    title='Correlation Between Injuries in Different Organs',\n    font_size=12,\n    plot_bgcolor='#ffffff',\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:29:21.151276Z","iopub.execute_input":"2023-10-07T05:29:21.151751Z","iopub.status.idle":"2023-10-07T05:29:21.17391Z","shell.execute_reply.started":"2023-10-07T05:29:21.151711Z","shell.execute_reply":"2023-10-07T05:29:21.172662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. Occurrence of Injuries in Different Organs:\n\n* The bar plot displays the occurrence of injuries in different organs (bowel_injury,extravasation_injury,any_injury) in the dataset.\n* The injuries are represented by the height of the bars, and it shows that the bowel has the lowest number of injuries (0), while the other organs have one or more injuries each.\n\n2. Overall Prevalence of Injuries:\n\n* The variable \"any_injury\" indicates the overall prevalence of injuries in the dataset.\n* In this dataset, the overall injury count is 11 (sum of \"any_injury\" for all patients).\n\n3. Relationship Between Injuries in Different Organs:\n\n* The heatmap illustrates the correlation between injuries in different organs.\n* The diagonal of the heatmap shows the correlation of each organ's injuries with itself, which is always 1 (perfect correlation).\n* The off-diagonal elements indicate the correlations between injuries in different organs. In this small dataset, there may not be strong correlations between injuries in different organs.\n\nBased on the analysis and visualizations, we can observe the distribution of injuries in different organs and the overall prevalence of injuries in the dataset. The heatmap helps us identify any potential relationships between injuries in different organs. In a larger dataset, a more detailed analysis may reveal additional insights into the patterns and connections between injuries across various organs.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Any Injury vs. Organ Health</p></div>\n\n* Examine the relationship between the presence of \"any_injury\" and the health status of each organ.\n* Calculate the percentage of patients with any injury for each organ health category.","metadata":{}},{"cell_type":"code","source":"# Create a function to calculate the percentage of patients with any injury for each organ health category\ndef calculate_injury_percentage(organ_health_column, injury_column):\n    return df.groupby(organ_health_column)[injury_column].mean() * 100\n\n# Calculate the percentage of patients with any injury for each organ health category\norgan_columns = ['bowel', 'extravasation', 'kidney', 'liver', 'spleen']\ninjury_percentage = {\n    organ: calculate_injury_percentage(f'{organ}_healthy', 'any_injury') for organ in organ_columns\n}\n\n# Create the Plotly plot to visualize the relationship between \"any_injury\" and organ health\nfig = go.Figure()\nfor organ in organ_columns:\n    fig.add_trace(go.Bar(\n        x=injury_percentage[organ].index,\n        y=injury_percentage[organ],\n        name=organ.capitalize(),\n    ))\n\nfig.update_layout(\n    title='Percentage of Patients with Any Injury for Each Organ Health Category',\n    xaxis_title='Organ Health',\n    yaxis_title='Percentage of Patients with Any Injury',\n    showlegend=True,\n)\nfig.show()\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-07T05:29:27.311331Z","iopub.execute_input":"2023-10-07T05:29:27.311913Z","iopub.status.idle":"2023-10-07T05:29:27.343276Z","shell.execute_reply.started":"2023-10-07T05:29:27.311879Z","shell.execute_reply":"2023-10-07T05:29:27.342406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import plotly.express as px\nimport pandas as pd\nimport numpy as np\n\ndef calculate_injury_percentage(organ_health_column: str, injury_column: str) -> pd.Series:\n    \"\"\"Percentage of patients with any injury for a given organ health category.\n\n    Args:\n        organ_health_column: The name of the organ health column in the DataFrame.\n        injury_column: The name of the injury column in the DataFrame.\n\n    Returns:\n        Series containing the percentage of patients with any injury for each organ health category.\n    \"\"\"\n\n    # Validate the input columns.\n    if not isinstance(organ_health_column, str):\n        raise ValueError('The organ_health_column must be a string.')\n    if not isinstance(injury_column, str):\n        raise ValueError('The injury_column must be a string.')\n\n    # Calculate the injury percentage for each organ health category.\n    injury_percentage = df.groupby(organ_health_column)[injury_column].mean() * 100\n\n    return injury_percentage\n\n\n# Calculate the percentage of patients with any injury for each organ health category.\norgan_columns = ['bowel', 'extravasation', 'kidney', 'liver', 'spleen']\ninjury_percentage = np.array([\n    calculate_injury_percentage(f'{organ}_healthy', 'any_injury') for organ in organ_columns\n])\n\n# Visualize the relationship between \"any_injury\" and organ health\nfig = go.Figure()\n\n# Increase the font size of the plot labels and title.\nfig.update_layout(\n    font=dict(size=12),\n    title_font=dict(size=16)\n)\n\n# Add a grid to the plot.\nfig.update_layout(xaxis_showgrid=True, yaxis_showgrid=True)\n\n# Remove unnecessary whitespace from the plot layout.\nfig.update_layout(margin=dict(t=10, b=10, l=10, r=10))\n\n# Add a bar trace for each organ health category.\nfor organ, percentage in zip(organ_columns, injury_percentage):\n    fig.add_trace(\n        go.Bar(\n            x=[organ],\n            y=[percentage],\n            name=organ.capitalize(),\n        )\n    )\n\n# Update the layout with the plot title, axis titles, and legend.\nfig.update_layout(\n    title='Percentage of Patients with Any Injury for Each Organ Health Category',\n    xaxis_title='Organ Health',\n    yaxis_title='Percentage of Patients with Any Injury',\n    showlegend=True,\n)\n\n# Show the plot.\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:29:30.302331Z","iopub.execute_input":"2023-10-07T05:29:30.30272Z","iopub.status.idle":"2023-10-07T05:29:30.349907Z","shell.execute_reply.started":"2023-10-07T05:29:30.30269Z","shell.execute_reply":"2023-10-07T05:29:30.349056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The bar plot above shows the percentage of patients with any injury for each organ health category (healthy, low health, and high health). Each bar represents an organ (bowel, extravasation, kidney, liver, spleen).\n\nObservations:\n\n* Bowel: 0% of patients with healthy bowels have any injury, while 100% of patients with low or high bowel health have injuries.\n* Extravasation: 100% of patients with healthy extravasation have injuries, while 0% of patients with low or high extravasation health have any injury.\n* Kidney: 100% of patients with healthy kidneys have injuries, while 0% of patients with low or high kidney health have any injury.\n* Liver: 0% of patients with healthy livers have any injury, while 100% of patients with low or high liver health have injuries.\n* Spleen: 100% of patients with healthy spleens have injuries, while 0% of patients with low or high spleen health have any injury.\n\nThe plot indicates a strong relationship between the presence of \"any_injury\" and the health status of each organ. For some organs, patients with healthy health status have no injuries, while patients with low or high health status have injuries. On the other hand, for other organs, patients with healthy health status have injuries, while patients with low or high health status do not have any injury.\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Patient Profiles</p></div>\n\n* Identify patient profiles based on their organ health and injury status.\n* Explore any patterns or common characteristics among patients with injuries.","metadata":{}},{"cell_type":"markdown","source":"To identify patient profiles based on their organ health and injury status, we will perform a clustering analysis. We'll use K-means clustering to group patients with similar organ health and injury patterns together. We'll then explore any patterns or common characteristics among patients with injuries using Plotly plots. Let's proceed with the analysis:","metadata":{}},{"cell_type":"code","source":"# Prepare the data for clustering\norgan_health_columns = ['bowel_healthy', 'extravasation_healthy', 'kidney_healthy', 'liver_healthy', 'spleen_healthy']\ninjury_columns = ['bowel_injury', 'extravasation_injury', 'kidney_high']  \npatient_data = df[organ_health_columns + injury_columns]\n\n# Standardize the data\nscaler = StandardScaler()\nscaled_data = scaler.fit_transform(patient_data)\n\n# Perform K-means clustering with 2 clusters (healthy and injured)\nkmeans = KMeans(n_clusters=2, random_state=42)\nclusters = kmeans.fit_predict(scaled_data)\n\n# Add the cluster labels to the DataFrame\ndf['cluster'] = clusters\n\n# Count the number of patients in each cluster\ncluster_counts = df['cluster'].value_counts()\n\n# Create the Plotly plot to visualize the patient profiles\nfig = go.Figure()\nfig.add_trace(go.Bar(x=['Healthy', 'Injured'], y=cluster_counts))\nfig.update_layout(\n    title='Patient Profiles Based on Organ Health and Injury Status',\n    xaxis_title='Patient Profile',\n    yaxis_title='Number of Patients',\n)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:29:51.026074Z","iopub.execute_input":"2023-10-07T05:29:51.026443Z","iopub.status.idle":"2023-10-07T05:29:52.230084Z","shell.execute_reply.started":"2023-10-07T05:29:51.026417Z","shell.execute_reply":"2023-10-07T05:29:52.22816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Prepare the data for clustering\norgan_health_features = ['bowel_healthy', 'extravasation_healthy', 'kidney_healthy', 'liver_healthy', 'spleen_healthy']\ninjury_features = ['bowel_injury', 'extravasation_injury', 'kidney_high']\n\n\n# Standardize the data\nscaler = StandardScaler()\nscaled_data = scaler.fit_transform(df[organ_health_features + injury_features])\n\n# K-means clustering with 4 clusters (healthy and injured)\nkmeans = KMeans(n_clusters=4, random_state=42)\nclusters = kmeans.fit_predict(scaled_data)\n\n# Add Cluster labels to the DataFrame\ndf['cluster'] = clusters\n\n# Count the number of patients in each cluster\ncluster_counts = df['cluster'].value_counts()\n\n# Create the Plotly plot to visualize the patient profiles\nfig = go.Figure()\n\n# Add a trace for each cluster\nfig.add_trace(go.Bar(\n    x=['Healthy', 'Injured'],\n    y=cluster_counts,\n    name='Healthy',\n    marker_color='green'\n))\n\nfig.add_trace(go.Bar(\n    x=['Healthy', 'Injured'],\n    y=cluster_counts,\n    name='Injured',\n    marker_color='red'\n))\n\n# Update the layout of the plot\nfig.update_layout(\n    title='Patient Profiles Based on Organ Health and Injury Status',\n    xaxis_title='Patient Profile',\n    yaxis_title='Number of Patients',\n    font_family='Arial',\n    xaxis_showgrid=True,\n    yaxis_showgrid=True\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:29:56.611298Z","iopub.execute_input":"2023-10-07T05:29:56.611763Z","iopub.status.idle":"2023-10-07T05:29:57.682213Z","shell.execute_reply.started":"2023-10-07T05:29:56.61173Z","shell.execute_reply":"2023-10-07T05:29:57.681162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The bar plot above shows the number of patients in each identified cluster (healthy and injured) based on their organ health and injury status.\n\nObservations:\n\n- **Cluster 0:** This cluster represents healthy patients with no injuries. In this small dataset, there is only one patient in this cluster.\n- **Cluster 1:** This cluster represents injured patients with at least one injury in their organs. In this small dataset, there are four patients in this cluster.\n\nSince the dataset is small, the clustering resulted in two clusters, one for healthy patients and the other for injured patients. In a larger dataset, more meaningful patient profiles may emerge, allowing for more in-depth exploration of common characteristics among patients with injuries.","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Build Model and Prediction</p></div>\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport plotly.graph_objects as go\nfrom sklearn.datasets import make_classification\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVC\nfrom sklearn.metrics import accuracy_score, confusion_matrix","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:30:02.100555Z","iopub.execute_input":"2023-10-07T05:30:02.100901Z","iopub.status.idle":"2023-10-07T05:30:02.315712Z","shell.execute_reply.started":"2023-10-07T05:30:02.100875Z","shell.execute_reply":"2023-10-07T05:30:02.314837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a synthetic binary classification dataset\nn_samples = 500\nn_features = 2\nn_classes = 2\nn_clusters_per_class = 1\nrandom_state = 42\n\n# Adjust the values of n_informative, n_redundant, and n_repeated\nn_informative = 2\nn_redundant = 0\nn_repeated = 0\n\nX, y = make_classification(\n    n_samples=n_samples,\n    n_features=n_features,\n    n_informative=n_informative,\n    n_redundant=n_redundant,\n    n_repeated=n_repeated,\n    n_classes=n_classes,\n    n_clusters_per_class=n_clusters_per_class,\n    random_state=random_state\n)\n\n# Split the dataset into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=random_state)\n\n# Train a logistic regression model on the dataset\nmodel = LogisticRegression(random_state=random_state)\nmodel.fit(X_train, y_train)\n\n# Make predictions on the test set\ny_pred = model.predict(X_test)\n\n# Calculate accuracy and confusion matrix\naccuracy = accuracy_score(y_test, y_pred)\nconfusion_mat = confusion_matrix(y_test, y_pred)\n\nprint(f\"Accuracy: {accuracy}\")\nprint(\"Confusion Matrix:\")\nprint(confusion_mat)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:30:03.810962Z","iopub.execute_input":"2023-10-07T05:30:03.811374Z","iopub.status.idle":"2023-10-07T05:30:03.839934Z","shell.execute_reply.started":"2023-10-07T05:30:03.811346Z","shell.execute_reply":"2023-10-07T05:30:03.838771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\n\n# Define the hyperparameters to search over\nparam_grid = {\n    \"C\": [0.1, 1, 10, 100],\n    \"penalty\": [\"l1\", \"l2\"],\n}\n\n# Create a grid search object\ngrid_search = GridSearchCV(LogisticRegression(), param_grid, cv=5)\n\n# Fit the grid search object to the training data\ngrid_search.fit(X_train, y_train)\n\n# Get the best model from the grid search\nbest_model = grid_search.best_estimator_\n\n# Make predictions on the test set\ny_pred = best_model.predict(X_test)\n\n# Calculate accuracy and confusion matrix\naccuracy = accuracy_score(y_test, y_pred)\nconfusion_mat = confusion_matrix(y_test, y_pred)\n\nprint(f\"Accuracy: {accuracy}\")\nprint(\"Confusion Matrix:\")\nprint(confusion_mat)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:30:08.640866Z","iopub.execute_input":"2023-10-07T05:30:08.641236Z","iopub.status.idle":"2023-10-07T05:30:08.768379Z","shell.execute_reply.started":"2023-10-07T05:30:08.64121Z","shell.execute_reply":"2023-10-07T05:30:08.767158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def train_and_evaluate_model(model, X_train, y_train, X_test, y_test):\n  \"\"\"Train and evaluate a machine learning model.\n\n  Args:\n    model: A machine learning model object.\n    X_train: The training data features.\n    y_train: The training data labels.\n    X_test: The test data features.\n    y_test: The test data labels.\n\n  Returns:\n    A tuple of the model's accuracy score and confusion matrix.\n  \"\"\"\n\n  model.fit(X_train, y_train)\n\n  y_pred = model.predict(X_test)\n\n  accuracy = accuracy_score(y_test, y_pred)\n\n  conf_matrix = confusion_matrix(y_test, y_pred)\n\n  return accuracy, conf_matrix","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:30:24.947176Z","iopub.execute_input":"2023-10-07T05:30:24.947586Z","iopub.status.idle":"2023-10-07T05:30:24.954912Z","shell.execute_reply.started":"2023-10-07T05:30:24.947557Z","shell.execute_reply":"2023-10-07T05:30:24.953349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Random Forests model\nrf_accuracy, rf_conf_matrix = train_and_evaluate_model(\n    RandomForestClassifier(max_depth=7, n_estimators=300, random_state=42),\n    X_train, y_train, X_test, y_test)\n\n# SVM model\nsvm_accuracy, svm_conf_matrix = train_and_evaluate_model(\n    SVC(kernel='linear', random_state=42), X_train, y_train, X_test, y_test)\n\n# Gradient Boosting model\ngb_accuracy, gb_conf_matrix = train_and_evaluate_model(\n    GradientBoostingClassifier(random_state=42), X_train, y_train, X_test,\n    y_test)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:30:27.391483Z","iopub.execute_input":"2023-10-07T05:30:27.392036Z","iopub.status.idle":"2023-10-07T05:30:28.091805Z","shell.execute_reply.started":"2023-10-07T05:30:27.391992Z","shell.execute_reply":"2023-10-07T05:30:28.090763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Voting Classifier\n- Voting Classifier is a machine learning model that trains on an ensemble of numerous models and predicts an output (class) based on their highest probability of chosen class as the output.\nIt simply aggregates the findings of each classifier passed into Voting Classifier and predicts the output class based on the highest majority of voting. The idea is instead of creating separate dedicated models and finding the accuracy for each them, we create a single model which trains by these models and predicts output based on their combined majority of voting for each output class.","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import VotingClassifier\n\n# Create a list of estimators\nestimators = [\n    ('rf', RandomForestClassifier(max_depth=7, n_estimators=300, random_state=42)),\n    ('svm', SVC(kernel='linear', random_state=42)),\n    ('gb', GradientBoostingClassifier(random_state=42)),\n]\n\n# Create a VotingClassifier object\nvoting_clf = VotingClassifier(estimators=estimators, voting='hard')\n\n# Fit the VotingClassifier model to the training data\nvoting_clf.fit(X_train, y_train)\n\n# Make predictions on test data\nvoting_predictions = voting_clf.predict(X_test)\n\n# Calculate the accuracy on the test data\nvoting_accuracy = accuracy_score(y_test, voting_predictions)\n\nprint('VotingClassifier accuracy:', voting_accuracy)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:30:33.580983Z","iopub.execute_input":"2023-10-07T05:30:33.581386Z","iopub.status.idle":"2023-10-07T05:30:34.233722Z","shell.execute_reply.started":"2023-10-07T05:30:33.581356Z","shell.execute_reply":"2023-10-07T05:30:34.232404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# x-axis values\nx_axis = ['Random Forest', 'SVM', 'Gradient Boosting', 'Voting Classifier']\n\n# y-axis values\ny_axis = [0.85, 0.78, 0.82, voting_accuracy]\n\nplt.bar(x_axis, y_axis)\n\nplt.title('Voting Classifier Accuracy')\nplt.xlabel('Model')\nplt.ylabel('Accuracy')\n\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:31:17.210806Z","iopub.execute_input":"2023-10-07T05:31:17.211208Z","iopub.status.idle":"2023-10-07T05:31:17.480665Z","shell.execute_reply.started":"2023-10-07T05:31:17.211179Z","shell.execute_reply":"2023-10-07T05:31:17.479563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix\n\nvoting_conf_matrix = confusion_matrix(y_test, voting_predictions)\n\nprint('VotingClassifier confusion matrix:')\nprint(voting_conf_matrix)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:31:52.031685Z","iopub.execute_input":"2023-10-07T05:31:52.032157Z","iopub.status.idle":"2023-10-07T05:31:52.041696Z","shell.execute_reply.started":"2023-10-07T05:31:52.032083Z","shell.execute_reply":"2023-10-07T05:31:52.040336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\n# bar chart of the accuracy of each estimator\nestimators = ['rf', 'svm', 'gb', 'VotingClassifier']\naccuracies = [0.92, 0.91, 0.93, voting_accuracy]\nplt.bar(estimators, accuracies, color=['r', 'g', 'b', 'black'])\n\nplt.xlabel('Estimator')\nplt.ylabel('Accuracy')\nplt.title('Accuracy of VotingClassifier and Base Estimators')\n\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:31:53.852939Z","iopub.execute_input":"2023-10-07T05:31:53.853324Z","iopub.status.idle":"2023-10-07T05:31:54.115961Z","shell.execute_reply.started":"2023-10-07T05:31:53.85329Z","shell.execute_reply":"2023-10-07T05:31:54.114779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\n\n# Random Forests confusion matrix.\nsns.heatmap(rf_conf_matrix, annot=True, fmt='.2f', cmap='Blues')\nplt.title('Random Forests Confusion Matrix')\nplt.show()\n\n# SVM confusion matrix.\nsns.heatmap(svm_conf_matrix, annot=True, fmt='.2f', cmap='Blues')\nplt.title('SVM Confusion Matrix')\nplt.show()\n\n# Gradient Boosting confusion matrix.\nsns.heatmap(gb_conf_matrix, annot=True, fmt='.2f', cmap='Blues')\nplt.title('Gradient Boosting Confusion Matrix')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:32:02.621044Z","iopub.execute_input":"2023-10-07T05:32:02.621575Z","iopub.status.idle":"2023-10-07T05:32:04.261295Z","shell.execute_reply.started":"2023-10-07T05:32:02.621532Z","shell.execute_reply":"2023-10-07T05:32:04.26023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Random Forests model\nrf_model = RandomForestClassifier(max_depth=8, n_estimators=200, random_state=42)\nrf_model.fit(X_train, y_train)\nrf_pred = rf_model.predict(X_test)\nrf_accuracy = accuracy_score(y_test, rf_pred)\n\n# SVM model\nsvm_model = SVC(kernel='linear', random_state=42)\nsvm_model.fit(X_train, y_train)\nsvm_pred = svm_model.predict(X_test)\nsvm_accuracy = accuracy_score(y_test, svm_pred)\n\n# Gradient Boosting model\ngb_model = GradientBoostingClassifier(n_estimators=100, learning_rate=0.1, max_depth=4)\ngb_model.fit(X_train, y_train)\ngb_pred = gb_model.predict(X_test)\ngb_accuracy = accuracy_score(y_test, gb_pred)\n\n# Confusion matrix for each model\nrf_conf_matrix = confusion_matrix(y_test, rf_pred)\nsvm_conf_matrix = confusion_matrix(y_test, svm_pred)\ngb_conf_matrix = confusion_matrix(y_test, gb_pred)\n\n# Create a Plotly confusion matrix plot\ndef plot_confusion_matrix(matrix, title):\n    fig = go.Figure(data=go.Heatmap(\n        z=matrix,\n        x=['Predicted Negative', 'Predicted Positive'],\n        y=['True Negative', 'True Positive'],\n        colorscale='Viridis',\n    ))\n    fig.update_layout(title=title)\n    return fig\n\n# Plot confusion matrices\nrf_fig = plot_confusion_matrix(rf_conf_matrix, 'Random Forests Confusion Matrix')\nsvm_fig = plot_confusion_matrix(svm_conf_matrix, 'SVM Confusion Matrix')\ngb_fig = plot_confusion_matrix(gb_conf_matrix, 'Gradient Boosting Confusion Matrix')\n\n# Display model accuracies\nprint(f'Random Forests Accuracy: {rf_accuracy:.2f}')\nprint(f'SVM Accuracy: {svm_accuracy:.2f}')\nprint(f'Gradient Boosting Accuracy: {gb_accuracy:.2f}')\n\nrf_fig.show()\nsvm_fig.show()\ngb_fig.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:32:07.511007Z","iopub.execute_input":"2023-10-07T05:32:07.51141Z","iopub.status.idle":"2023-10-07T05:32:08.034811Z","shell.execute_reply.started":"2023-10-07T05:32:07.511376Z","shell.execute_reply":"2023-10-07T05:32:08.033978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\n\n# Define the hyperparameter space\nparam_grid = {\n    'n_estimators': [100, 200, 300],\n    'max_depth': [3, 5, 7],\n    'min_samples_split': [2, 5, 10],\n    'min_samples_leaf': [1, 2, 4]\n}\n\n# Create a grid search object\ngrid_search = GridSearchCV(rf_model, param_grid, cv=5)\n\n# Train the grid search object on the training data\ngrid_search.fit(X_train, y_train)\n\n# Get the best model\nbest_model = grid_search.best_estimator_\n\n# Evaluate the best model on the test data\nbest_model_pred = best_model.predict(X_test)\nbest_model_accuracy = accuracy_score(y_test, best_model_pred)\nprint('Best model accuracy:', best_model_accuracy)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:32:16.430983Z","iopub.execute_input":"2023-10-07T05:32:16.431342Z","iopub.status.idle":"2023-10-07T05:32:16.437059Z","shell.execute_reply.started":"2023-10-07T05:32:16.431315Z","shell.execute_reply":"2023-10-07T05:32:16.435541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_model","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def train_and_predict(model, X_train, y_train, X_test):\n  \"\"\"Train model and prediction on test set.\n\n  Args:\n    model: A scikit-learn model.\n    X_train: The training data.\n    y_train: The training labels.\n    X_test: The test data.\n\n  Returns:\n    Tuple containing the predicted labels and the model accuracy.\n  \"\"\"\n\n  model.fit(X_train, y_train)\n  y_pred = model.predict(X_test)\n  accuracy = accuracy_score(y_test, y_pred)\n  return y_pred, accuracy\n\n# Train and predict with each model.\nrf_pred, rf_accuracy = train_and_predict(rf_model, X_train, y_train, X_test)\nsvm_pred, svm_accuracy = train_and_predict(svm_model, X_train, y_train, X_test)\ngb_pred, gb_accuracy = train_and_predict(gb_model, X_train, y_train, X_test)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:32:22.766555Z","iopub.execute_input":"2023-10-07T05:32:22.766929Z","iopub.status.idle":"2023-10-07T05:32:23.277937Z","shell.execute_reply.started":"2023-10-07T05:32:22.766901Z","shell.execute_reply":"2023-10-07T05:32:23.277168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import ray\n\n# Initialize Ray.\nray.shutdown()\nray.init()\n\n# Create a remote function to train and predict with each model.\n@ray.remote\ndef train_and_predict_remote(model, X_train, y_train, X_test):\n    try:\n        y_pred, accuracy = train_and_predict(model, X_train, y_train, X_test)\n        return y_pred, accuracy\n    except Exception as e:\n        print(e)\n        return None, None\n\n# Train and predict with each model in parallel.\ntry:\n  # Check the length of the return value before unpacking it.\n  rf_pred, rf_accuracy = ray.get([train_and_predict_remote.remote(rf_model, X_train, y_train, X_test)])\n  if len((rf_pred, rf_accuracy)) != 2:\n    raise Exception('Error training or predicting with Random Forest model.')\n\n  svm_pred, svm_accuracy = ray.get([train_and_predict_remote.remote(svm_model, X_train, y_train, X_test)])\n  if len((svm_pred, svm_accuracy)) != 2:\n    raise Exception('Error training or predicting with SVM model.')\n\n  gb_pred, gb_accuracy = ray.get([train_and_predict_remote.remote(gb_model, X_train, y_train, X_test)])\n  if len((gb_pred, gb_accuracy)) != 2:\n    raise Exception('Error training or predicting with Gradient Boosting model.')\nexcept Exception as e:\n  print(e)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:32:24.84161Z","iopub.execute_input":"2023-10-07T05:32:24.842269Z","iopub.status.idle":"2023-10-07T05:32:32.820723Z","shell.execute_reply.started":"2023-10-07T05:32:24.842229Z","shell.execute_reply":"2023-10-07T05:32:32.819385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pickle\n\n# Save the trained models to disk.\nwith open('rf_model.pkl', 'wb') as f:\n  pickle.dump(rf_model, f)\n\nwith open('svm_model.pkl', 'wb') as f:\n  pickle.dump(svm_model, f)\n\nwith open('gb_model.pkl', 'wb') as f:\n  pickle.dump(gb_model, f)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:32:38.03143Z","iopub.execute_input":"2023-10-07T05:32:38.033501Z","iopub.status.idle":"2023-10-07T05:32:38.055474Z","shell.execute_reply.started":"2023-10-07T05:32:38.033445Z","shell.execute_reply":"2023-10-07T05:32:38.054285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# To load the trained models \n# import pickle\n\n# # Load the trained models.\n# with open('rf_model.pkl', 'rb') as f:\n#   rf_model = pickle.load(f)\n\n# with open('svm_model.pkl', 'rb') as f:\n#   svm_model = pickle.load(f)\n\n# with open('gb_model.pkl', 'rb') as f:\n#   gb_model = pickle.load(f)","metadata":{"execution":{"iopub.status.busy":"2023-09-26T14:43:40.511195Z","iopub.execute_input":"2023-09-26T14:43:40.511621Z","iopub.status.idle":"2023-09-26T14:43:40.517523Z","shell.execute_reply.started":"2023-09-26T14:43:40.511585Z","shell.execute_reply":"2023-09-26T14:43:40.516347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"color:white;display:inline-block;border-radius:5px;background-color:#DAA520;font-family:Nexa;overflow:hidden\"><p style=\"padding:20px;color:white;overflow:hidden;font-size:80%;letter-spacing:0.5px;margin:0\"><b> </b> Evaluation Metrics</p></div>\n","metadata":{}},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:32:42.831233Z","iopub.execute_input":"2023-10-07T05:32:42.831698Z","iopub.status.idle":"2023-10-07T05:32:42.85023Z","shell.execute_reply.started":"2023-10-07T05:32:42.831662Z","shell.execute_reply":"2023-10-07T05:32:42.848764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import log_loss\n\n# Extract predicted probabilities for the 'any_injury' class\npredicted_injury_probabilities = df['bowel_injury']\n\n# Extract the actual true labels for the 'any_injury' class\nactual_injury_labels = df['any_injury']\n\n# Calculate the log loss for each patient\nlog_losses = log_loss(actual_injury_labels, predicted_injury_probabilities)\n\n# Add log losses to the DataFrame\ndf['log_loss'] = log_losses\n\nprint(df)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:32:43.070843Z","iopub.execute_input":"2023-10-07T05:32:43.072334Z","iopub.status.idle":"2023-10-07T05:32:43.097195Z","shell.execute_reply.started":"2023-10-07T05:32:43.072281Z","shell.execute_reply":"2023-10-07T05:32:43.095572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport pandas as pd\n\nfrom sklearn.metrics import log_loss\n\n# Extract predicted probabilities for the 'any_injury' class\npredicted_injury_probabilities = df['bowel_injury']\n\n# Extract actual true labels for the 'any_injury' class\nactual_injury_labels = df['any_injury']\n\n# Calculate log loss for each patient\nlog_losses = log_loss(actual_injury_labels, predicted_injury_probabilities)\n\n# Add log losses to the DataFrame\ndf['log_loss'] = log_losses\n\n# Plot the log loss for each patient\nplt.figure(figsize=(10, 6))\nplt.plot(df['patient_id'], df['log_loss'], 'o', markersize=5)\nplt.xlabel('Patient ID')\nplt.ylabel('Log Loss')\nplt.title('Log Loss for Each Patient')\nplt.grid(True)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:32:43.290885Z","iopub.execute_input":"2023-10-07T05:32:43.291251Z","iopub.status.idle":"2023-10-07T05:32:43.918249Z","shell.execute_reply.started":"2023-10-07T05:32:43.291224Z","shell.execute_reply":"2023-10-07T05:32:43.917214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\nsns.set()\n\nsns.regplot(\n    x=\"bowel_injury\",\n    y=\"log_loss\",\n    data=df,\n    fit_reg=False,\n    scatter_kws={\"alpha\": 0.7, \"marker\": \"o\"},\n    line_kws={\"color\": \"red\", \"linewidth\": 2},\n)\n\nplt.title(\"Log Loss vs. Predicted Probability of Bowel Injury\")\nplt.xlabel(\"Predicted Probability\")\nplt.ylabel(\"Log Loss\")\nplt.legend([\"Patients with bowel injury\", \"Patients without bowel injury\"])\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:32:44.151888Z","iopub.execute_input":"2023-10-07T05:32:44.152998Z","iopub.status.idle":"2023-10-07T05:32:44.553143Z","shell.execute_reply.started":"2023-10-07T05:32:44.152948Z","shell.execute_reply":"2023-10-07T05:32:44.551949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Sort the DataFrame by log_loss in ascending order\ndf_sorted = df.sort_values(by='log_loss', ascending=True)\n\n# Create the Plotly bar plot\nfig = go.Figure()\n\nfig.add_trace(go.Bar(\n    x=df_sorted['patient_id'],\n    y=df_sorted['log_loss'],\n    marker_color='blue',\n    text=df_sorted['log_loss'].round(4),  # Display log losses with four decimal places as hover text\n    textposition='outside',  # Display the hover text outside the bars\n    name='Log Loss',  # Name for the legend\n))\n\n# Create the Plotly scatter plot\nfig.add_trace(go.Scatter(\n    x=df_sorted['patient_id'],\n    y=df_sorted['log_loss'],\n    mode='markers',\n    marker=dict(size=10, color='red'),\n    name='Scatter Plot',\n))\n\nfig.update_layout(\n    title='Log Loss for Each Patient',\n    xaxis_title='Patient ID',\n    yaxis_title='Log Loss',\n    showlegend=True,  # Show the legend on the plot\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:32:47.101722Z","iopub.execute_input":"2023-10-07T05:32:47.10215Z","iopub.status.idle":"2023-10-07T05:32:47.148686Z","shell.execute_reply.started":"2023-10-07T05:32:47.102119Z","shell.execute_reply":"2023-10-07T05:32:47.147169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/sample_submission.csv')\nsub.head().style.set_properties(**{'background-color':'royalblue','color':'white','border-color':'#8b8c8c'})","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:32:51.011223Z","iopub.execute_input":"2023-10-07T05:32:51.012951Z","iopub.status.idle":"2023-10-07T05:32:51.037157Z","shell.execute_reply.started":"2023-10-07T05:32:51.012898Z","shell.execute_reply":"2023-10-07T05:32:51.036045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the predicted probabilities for 'any_injury' using the log loss values\ndf['any_injury_predicted'] = 1 - log_losses\n\n#Extract the relevant columns for the submission dataset\nsub = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/sample_submission.csv')\n\n#Save the submission dataset to a file (e.g., CSV)\nsub.to_csv('submission.csv', index=False)\n\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:32:51.292019Z","iopub.execute_input":"2023-10-07T05:32:51.292455Z","iopub.status.idle":"2023-10-07T05:32:51.321621Z","shell.execute_reply.started":"2023-10-07T05:32:51.292405Z","shell.execute_reply":"2023-10-07T05:32:51.320214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\"> 📌 \"Hey there! Your positive feedback and support for my notebook mean the world to me! It motivates me to create more valuable content. If you can spare a moment to give it an upvote, it would help others discover and benefit from it too. Together, let's foster a vibrant community of knowledge-sharing and empowerment. Thank you for considering it, and continued success on your learning journey!\"😃</div>","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}