{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":99552,"databundleVersionId":13190393,"sourceType":"competition"}],"dockerImageVersionId":31090,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"border-left: 4px solid #4e4e4e; padding: 1.5rem; background-color: #1e1e1e; font-family: 'Segoe UI', sans-serif; color: #e0e0e0; line-height: 1.6; font-size: 16px; border-radius: 8px;\">\n\n\n\n  \n  <h3 style=\"color: #ffa07a;\">🩸 What is an Intracranial Aneurysm?</h3>\n  <div style=\"display: flex; justify-content: center; margin: 2rem 0;\">\n  <img src=\"https://upload.wikimedia.org/wikipedia/commons/8/80/Cerebral_aneurysm_NIH.jpg\" \n       alt=\"Brain Aneurysm Scan\" \n       style=\"width: 50%; border-radius: 10px;\" />\n</div>\n\n\n  <p>\n    An intracranial aneurysm is a bulging or ballooning of a blood vessel in the brain. If undetected, it can rupture — leading to life-threatening bleeding. Early detection is crucial. The challenge? These aneurysms are often subtle, hidden in complex anatomy, and differ widely between patients.\n\n\n\n  </p>\n  <p>\n    That’s where machine learning can step in — augmenting radiologist workflows by detecting patterns across thousands of image slices in a fraction of the time.\n  </p>\n\n  <h3 style=\"color: #ffa07a;\">📷 Imaging Modalities Used</h3>\n  <p>We’re working with 3D scans from the following modalities:</p>\n\n  <table style=\"width: 100%; border-collapse: collapse; margin-top: 1rem; font-size: 15px;\">\n    <thead>\n      <tr style=\"background-color: #2e2e2e;\">\n        <th style=\"text-align: left; padding: 8px;\">Modality</th>\n        <th style=\"text-align: left; padding: 8px;\">Full Form</th>\n        <th style=\"text-align: left; padding: 8px;\">What It Shows</th>\n      </tr>\n    </thead>\n    <tbody>\n      <tr>\n        <td style=\"padding: 8px;\">CTA</td>\n        <td style=\"padding: 8px;\">Computed Tomography Angiography</td>\n        <td style=\"padding: 8px;\">Blood vessels using contrast-enhanced CT</td>\n      </tr>\n      <tr style=\"background-color: #2b2b2b;\">\n        <td style=\"padding: 8px;\">MRA</td>\n        <td style=\"padding: 8px;\">Magnetic Resonance Angiography</td>\n        <td style=\"padding: 8px;\">Blood vessels using MRI</td>\n      </tr>\n      <tr>\n        <td style=\"padding: 8px;\">MRI</td>\n        <td style=\"padding: 8px;\">Magnetic Resonance Imaging</td>\n        <td style=\"padding: 8px;\">Structural brain tissue, no contrast</td>\n      </tr>\n    </tbody>\n  </table>\n\n  <img src=\"https://www.nature.com/articles/s41598-023-33182-1/figures/1\" alt=\"3D Scan Slices\" style=\"width: 100%; border-radius: 10px; margin: 1rem 0;\" />\n\n  <p>\n    Each scan is a 3D volume made up of 2D slices. Think of it as flipping through pages of a brain atlas.\n  </p>\n\n  <h3 style=\"color: #ffa07a;\">📌 What Are We Predicting?</h3>\n  <p>Two key tasks define this competition:</p>\n  <ul>\n    <li><strong>Classification</strong> — Does this scan contain an aneurysm?</li>\n    <li><strong>Localization</strong> — If yes, where is it located? (3D coordinates + 1 of 13 brain regions)</li>\n  </ul>\n\n  <h3 style=\"color: #ffa07a;\">🗃️ Dataset Overview</h3>\n  <ul>\n    <li><code>train.csv</code> — Aneurysm presence and location labels</li>\n    <li><code>train_localizers.csv</code> — 3D coordinates for aneurysms (subset)</li>\n    <li><code>DICOM folders</code> — Raw scan data (one folder per patient)</li>\n    <li><code>segmentations</code> — Optional vessel masks in <code>.nii.gz</code> format</li>\n  </ul>\n\n  <h3 style=\"color: #ffa07a;\">🌍 Why This Matters</h3>\n  <p>\n    In real-world hospitals, radiologists spend hours scanning hundreds of slices to find aneurysms. Our model has the potential to become a powerful clinical assistant flagging risk areas, reducing diagnostic time, and minimizing human error.\n  </p>\n  <p>\n    The stakes are high. So is the opportunity.\n  </p>\n</div>\n","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">Let’s Load The Data</h2>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ntrain_df = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train.csv')\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:02:32.752066Z","iopub.execute_input":"2025-07-29T13:02:32.75262Z","iopub.status.idle":"2025-07-29T13:02:32.782419Z","shell.execute_reply.started":"2025-07-29T13:02:32.752597Z","shell.execute_reply":"2025-07-29T13:02:32.781901Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">Aneurysm Counts per Location</h2>\n<H4>Where do aneurysms tend to occur?\nThis bar chart answers that — using the 13 location_* columns.</H4>","metadata":{}},{"cell_type":"code","source":"import plotly.express as px\n\nlocation_cols = [col for col in train_df.columns if col not in ['SeriesInstanceUID', 'PatientAge', 'PatientSex', 'Modality', 'Aneurysm Present']]\nlocation_counts = train_df[location_cols].sum().sort_values(ascending=False)\n\nfig = px.bar(\n    location_counts,\n    orientation='v',\n    title='📊 Aneurysm Count by Location',\n    labels={'value': 'Count', 'index': 'Location'},\n    color=location_counts.values,\n    color_continuous_scale=[[0, '#00BFC4'], [1, '#C77CFF']],\n    template='plotly_dark'\n)\nfig.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:02:32.783851Z","iopub.execute_input":"2025-07-29T13:02:32.78404Z","iopub.status.idle":"2025-07-29T13:02:34.998101Z","shell.execute_reply.started":"2025-07-29T13:02:32.784024Z","shell.execute_reply":"2025-07-29T13:02:34.99743Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">🧠 Imaging Modality Distribution</h2> <h4>Which imaging techniques are used? This donut chart shows the breakdown of modalities like CT and MRI in the dataset.</h4>","metadata":{}},{"cell_type":"code","source":"modality_counts = train_df['Modality'].value_counts()\n\nfig = px.pie(\n    names=modality_counts.index,\n    values=modality_counts.values,\n    title='🧠 Imaging Modality Distribution',\n    hole=0.4,\n    color_discrete_sequence=['#00BFC4', '#C77CFF'],\n    template='plotly_dark'\n)\nfig.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:02:34.998879Z","iopub.execute_input":"2025-07-29T13:02:34.999122Z","iopub.status.idle":"2025-07-29T13:02:35.049919Z","shell.execute_reply.started":"2025-07-29T13:02:34.9991Z","shell.execute_reply":"2025-07-29T13:02:35.049376Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">📊 Age Distribution by Aneurysm Presence</h2> <h4>Are aneurysms more common at certain ages? This histogram compares age trends for patients with and without aneurysms.</h4>","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\nplt.style.use('dark_background')\nplt.figure(figsize=(10,6))\nsns.histplot(\n    data=train_df,\n    x='PatientAge',\n    hue='Aneurysm Present',\n    bins=30,\n    kde=True,\n    palette={0: '#00BFC4', 1: '#C77CFF'}\n)\nplt.title(\"Age Distribution by Aneurysm Presence\", fontsize=14)\nplt.xlabel(\"Age\")\nplt.ylabel(\"Frequency\")\nplt.grid(alpha=0.2)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:02:35.050545Z","iopub.execute_input":"2025-07-29T13:02:35.05084Z","iopub.status.idle":"2025-07-29T13:02:36.22642Z","shell.execute_reply.started":"2025-07-29T13:02:35.050822Z","shell.execute_reply":"2025-07-29T13:02:36.225639Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pd.crosstab(train_df['PatientSex'], train_df['Aneurysm Present'], normalize='index') * 100\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:02:36.228318Z","iopub.execute_input":"2025-07-29T13:02:36.228631Z","iopub.status.idle":"2025-07-29T13:02:36.25204Z","shell.execute_reply.started":"2025-07-29T13:02:36.228613Z","shell.execute_reply":"2025-07-29T13:02:36.251555Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">⚖️ Class Imbalance: Any Aneurysm Present</h2> <h4>How balanced is the dataset? This chart shows how many cases have aneurysms vs. those that don’t — crucial for model performance.</h4>","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(\n    train_df,\n    x='Aneurysm Present',\n    title='⚖️ Class Imbalance: Any Aneurysm Present',\n    color='Aneurysm Present',\n    text_auto=True,\n    color_discrete_map={0: '#00BFC4', 1: '#C77CFF'},\n    template='plotly_dark'\n)\nfig.update_xaxes(type='category', tickvals=[0, 1], ticktext=[\"No Aneurysm\", \"Aneurysm\"])\nfig.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:02:36.252785Z","iopub.execute_input":"2025-07-29T13:02:36.253213Z","iopub.status.idle":"2025-07-29T13:02:36.325605Z","shell.execute_reply.started":"2025-07-29T13:02:36.253171Z","shell.execute_reply":"2025-07-29T13:02:36.325019Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">🧬 Aneurysm Risk by Gender</h2> <h4>What’s the gender split? This table shows the % of aneurysm cases within each gender — offering early insight into possible demographic risk factors.</h4>","metadata":{}},{"cell_type":"code","source":"# Simple % table (can be styled in notebook output)\ngender_ct = pd.crosstab(train_df['PatientSex'], train_df['Aneurysm Present'], normalize='index') * 100\ngender_ct = gender_ct.rename(columns={0: 'No Aneurysm (%)', 1: 'Aneurysm (%)'})\ngender_ct.style.background_gradient(cmap='crest').format(\"{:.1f}%\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:02:36.32629Z","iopub.execute_input":"2025-07-29T13:02:36.326509Z","iopub.status.idle":"2025-07-29T13:02:36.41216Z","shell.execute_reply.started":"2025-07-29T13:02:36.326493Z","shell.execute_reply":"2025-07-29T13:02:36.411569Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">Train Localizer Analysis </h2>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport ast\n\nlocalizers_df = pd.read_csv('/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv')\n\n# Convert coordinate strings to dicts\nlocalizers_df['coords'] = localizers_df['coordinates'].apply(ast.literal_eval)\nlocalizers_df['x'] = localizers_df['coords'].apply(lambda d: d['x'])\nlocalizers_df['y'] = localizers_df['coords'].apply(lambda d: d['y'])\n\nlocalizers_df.drop(columns=['coordinates', 'coords'], inplace=True)\nlocalizers_df.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:02:36.412944Z","iopub.execute_input":"2025-07-29T13:02:36.41334Z","iopub.status.idle":"2025-07-29T13:02:36.464993Z","shell.execute_reply.started":"2025-07-29T13:02:36.413322Z","shell.execute_reply":"2025-07-29T13:02:36.46445Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">📍 Aneurysm Localization Heatmap</h2> <h4>Where are aneurysms typically found in image space? This 2D heatmap visualizes the frequency of annotated aneurysm coordinates (`x`, `y`) across all scans.</h4>","metadata":{}},{"cell_type":"code","source":"fig = px.density_heatmap(\n    localizers_df,\n    x='x',\n    y='y',\n    nbinsx=50,\n    nbinsy=50,\n    title='🧠 Heatmap of Aneurysm Locations in Image Space',\n    color_continuous_scale='Turbo',\n    template='plotly_dark',\n)\nfig.update_yaxes(autorange=\"reversed\")\nfig.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:02:36.465597Z","iopub.execute_input":"2025-07-29T13:02:36.465783Z","iopub.status.idle":"2025-07-29T13:02:36.515716Z","shell.execute_reply.started":"2025-07-29T13:02:36.465769Z","shell.execute_reply":"2025-07-29T13:02:36.515035Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">🧭 Aneurysm Coordinates by Location</h2> <h4>Which parts of the brain are affected where? This 2D scatterplot maps each aneurysm's annotated `(x, y)` position and colors them by anatomical location.</h4>","metadata":{}},{"cell_type":"code","source":"fig = px.scatter(\n    localizers_df,\n    x='x',\n    y='y',\n    color='location',\n    title='🧠 2D Scatter of Aneurysm Coordinates by Location',\n    template='plotly_dark',\n    color_discrete_sequence=px.colors.qualitative.Dark24\n)\nfig.update_yaxes(autorange=\"reversed\")\nfig.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:02:36.516456Z","iopub.execute_input":"2025-07-29T13:02:36.516636Z","iopub.status.idle":"2025-07-29T13:02:36.590465Z","shell.execute_reply.started":"2025-07-29T13:02:36.516622Z","shell.execute_reply":"2025-07-29T13:02:36.589713Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">🔁 Aneurysm Co-occurrence Across Brain Regions</h2>\n\n<p style=\"color:#ccc\">\nEach patient in the dataset can have multiple aneurysms in different brain arteries. But are certain aneurysm sites more likely to occur together?\n\nTo investigate this, we treat the location columns as binary labels and compute a co-occurrence matrix — it tells us how often, for instance, an aneurysm in the <code>Left MCA</code> is also accompanied by one in the <code>ACoA</code>, and so on.\n</p>\n","metadata":{}},{"cell_type":"code","source":"df = train_df\n# Identify location columns correctly\nlocation_cols = df.columns[4:-1]  # skip UID, Age, Sex, Modality, and skip final label\nlocation_df = df[location_cols].astype(int)  # just in case they're still object type\n\n# Co-occurrence matrix\nco_matrix = location_df.T.dot(location_df)\n\n# Plot heatmap\nplt.figure(figsize=(12, 10))\nsns.heatmap(co_matrix, cmap=\"magma\", annot=True, fmt=\".0f\", linewidths=0.5)\nplt.title(\"🧠 Aneurysm Co-occurrence Matrix\", fontsize=16, color='white')\nplt.xticks(rotation=45, ha='right', fontsize=9)\nplt.yticks(rotation=0, fontsize=9)\nplt.gca().set_facecolor('black')\nplt.gcf().set_facecolor('#111111')\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:03:34.461402Z","iopub.execute_input":"2025-07-29T13:03:34.461665Z","iopub.status.idle":"2025-07-29T13:03:35.08289Z","shell.execute_reply.started":"2025-07-29T13:03:34.461643Z","shell.execute_reply":"2025-07-29T13:03:35.082109Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">🧠 Aneurysm Spread per Patient</h2> <h4>How many different brain regions are affected per case? This histogram shows the count of locations marked positive (1) for each patient in the dataset.</h4>","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# 1. Location Count Distribution (number of 1s per row)\nlocation_counts = location_df.sum(axis=1)\n\nplt.figure(figsize=(8, 5))\nsns.histplot(location_counts, bins=range(1, location_counts.max()+2), kde=False, color=\"teal\")\nplt.title(\"Distribution of Aneurysm Locations per Patient\")\nplt.xlabel(\"Number of Locations with Aneurysm\")\nplt.ylabel(\"Number of Patients\")\nplt.grid(True)\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:03:38.836367Z","iopub.execute_input":"2025-07-29T13:03:38.836744Z","iopub.status.idle":"2025-07-29T13:03:39.0427Z","shell.execute_reply.started":"2025-07-29T13:03:38.836722Z","shell.execute_reply":"2025-07-29T13:03:39.042033Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">🧬 Aneurysm Location Frequency by Sex</h2> <h4>Are certain aneurysm locations more common in one sex? This grouped bar chart breaks down aneurysm counts by brain region and patient sex.</h4>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Assuming 'df' is already loaded and preprocessed (i.e., binary columns are int)\n\n# Filter only rows where aneurysm is present\ndf_pos = df[df[\"Aneurysm Present\"] == 1]\n\n# List of all aneurysm location columns\nlocation_cols = df_pos.columns[4:-1]  # From 'Left Infraclinoid...' to 'Other Posterior Circulation'\n\n# Group by PatientSex and sum each location\nsex_location = df_pos.groupby(\"PatientSex\")[location_cols].sum().T\n\n# Reset index for plotting\nsex_location = sex_location.reset_index().melt(id_vars=\"index\", var_name=\"Sex\", value_name=\"Count\")\nsex_location = sex_location.rename(columns={\"index\": \"Location\"})\n\n# Plot\nplt.figure(figsize=(16, 6))\nsns.barplot(data=sex_location, x=\"Location\", y=\"Count\", hue=\"Sex\")\nplt.xticks(rotation=45, ha=\"right\")\nplt.title(\"🧬 Aneurysm Location Frequency by Sex\")\nplt.ylabel(\"Aneurysm Count\")\nplt.xlabel(\"Location\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:03:40.649785Z","iopub.execute_input":"2025-07-29T13:03:40.650036Z","iopub.status.idle":"2025-07-29T13:03:40.992175Z","shell.execute_reply.started":"2025-07-29T13:03:40.650016Z","shell.execute_reply":"2025-07-29T13:03:40.991539Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">🎂 Patient Age Distribution per Aneurysm Location</h2> <h4>Do certain aneurysm locations occur more frequently in specific age groups? This boxplot shows age ranges of patients for each brain region affected.</h4>","metadata":{}},{"cell_type":"code","source":"# List of location columns\nlocation_cols = df.columns[4:-1]\n\n# Filter only positive cases\ndf_pos = df[df[\"Aneurysm Present\"] == 1]\n\n# Prepare a DataFrame: for each location, collect patients with that location = 1\nage_location = []\n\nfor loc in location_cols:\n    subset = df_pos[df_pos[loc] == 1][[\"PatientAge\"]].copy()\n    subset[\"Location\"] = loc\n    age_location.append(subset)\n\n# Combine all into one DataFrame\nage_location_df = pd.concat(age_location)\n\n# Plot\nplt.figure(figsize=(16, 6))\nsns.boxplot(data=age_location_df, x=\"Location\", y=\"PatientAge\", palette=\"crest\")\nplt.xticks(rotation=45, ha=\"right\")\nplt.title(\"🎂 Patient Age Distribution per Aneurysm Location\")\nplt.ylabel(\"Age\")\nplt.xlabel(\"Location\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:03:42.216931Z","iopub.execute_input":"2025-07-29T13:03:42.217133Z","iopub.status.idle":"2025-07-29T13:03:42.546712Z","shell.execute_reply.started":"2025-07-29T13:03:42.217118Z","shell.execute_reply":"2025-07-29T13:03:42.546075Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">🎂 Patient Age Distribution per Aneurysm Location (Violin Plot)</h2> <h4>How does the age of patients vary across aneurysm locations? This violin plot reveals both distribution spread and central tendencies per region.</h4>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(16, 6))\nsns.violinplot(data=age_location_df, x=\"Location\", y=\"PatientAge\", palette=\"crest\")\nplt.xticks(rotation=45, ha=\"right\")\nplt.title(\"🎂 Patient Age Distribution per Aneurysm Location (Violin Plot)\")\nplt.ylabel(\"Age\")\nplt.xlabel(\"Location\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:03:43.878911Z","iopub.execute_input":"2025-07-29T13:03:43.879156Z","iopub.status.idle":"2025-07-29T13:03:44.346713Z","shell.execute_reply.started":"2025-07-29T13:03:43.879136Z","shell.execute_reply":"2025-07-29T13:03:44.346059Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">🧠 Modality Preference per Aneurysm Location</h2> <h4>Which imaging modality is preferred for each aneurysm location?</h4> This heatmap shows how often **CTA** (Computed Tomography Angiography) and **MRA** (Magnetic Resonance Angiography) are used for different aneurysm locations. It helps identify modality biases — for example, some regions might be more frequently scanned with CTA, while others lean toward MRA, possibly due to anatomical accessibility or diagnostic clarity.","metadata":{}},{"cell_type":"code","source":"# Filter rows with aneurysm present\ndf_pos = df[df[\"Aneurysm Present\"] == 1].copy()\n\n# Initialize a dictionary to store counts\nmodality_counts = {}\n\nfor loc in location_cols:\n    subset = df_pos[df_pos[loc] == 1]\n    modality_distribution = subset[\"Modality\"].value_counts()\n    modality_counts[loc] = modality_distribution\n\n# Convert to DataFrame and fill missing values with 0\nmodality_df = pd.DataFrame(modality_counts).T.fillna(0).astype(int)\nmodality_df = modality_df[[\"CTA\", \"MRA\"]] if \"CTA\" in modality_df.columns and \"MRA\" in modality_df.columns else modality_df\n\n# Plot heatmap\nplt.figure(figsize=(12, 6))\nsns.heatmap(modality_df.T, annot=True, cmap=\"crest\", fmt=\"d\")\nplt.title(\"🧠 Modality Preference per Aneurysm Location\")\nplt.xlabel(\"Location\")\nplt.ylabel(\"Modality\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:03:47.687832Z","iopub.execute_input":"2025-07-29T13:03:47.688483Z","iopub.status.idle":"2025-07-29T13:03:48.106244Z","shell.execute_reply.started":"2025-07-29T13:03:47.688462Z","shell.execute_reply":"2025-07-29T13:03:48.105605Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">🧠 Modality Usage per Aneurysm Location</h2> <h4>How many patients were imaged with each modality at each aneurysm site?</h4> This grouped bar chart compares **CTA** and **MRA** usage across aneurysm locations — offering a clear view of modality distribution patterns across anatomical regions.","metadata":{}},{"cell_type":"code","source":"modality_df_melted = modality_df.reset_index().melt(id_vars=\"index\", value_name=\"Count\", var_name=\"Modality\")\nmodality_df_melted.rename(columns={\"index\": \"Location\"}, inplace=True)\n\nplt.figure(figsize=(14, 6))\nsns.barplot(data=modality_df_melted, x=\"Location\", y=\"Count\", hue=\"Modality\", palette=\"crest\")\nplt.title(\"🧠 Modality Usage per Aneurysm Location\")\nplt.xticks(rotation=45, ha=\"right\")\nplt.xlabel(\"Location\")\nplt.ylabel(\"Patient Count\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:03:50.702087Z","iopub.execute_input":"2025-07-29T13:03:50.702752Z","iopub.status.idle":"2025-07-29T13:03:51.032172Z","shell.execute_reply.started":"2025-07-29T13:03:50.702729Z","shell.execute_reply":"2025-07-29T13:03:51.031623Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#D2B48C\">🧭 Spatial Heatmaps of Aneurysm Coordinates</h2> <h4>Do certain aneurysm types tend to cluster spatially?</h4> These KDE heatmaps highlight **where** aneurysms appear within scans for the top 5 locations — revealing anatomical clustering patterns based on `x` and `y` coordinates.","metadata":{}},{"cell_type":"code","source":"import ast\n\ntrain_localizations = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv\")\n# Parse coordinates\ntrain_localizations[\"coord_dict\"] = train_localizations[\"coordinates\"].apply(ast.literal_eval)\ntrain_localizations[\"x\"] = train_localizations[\"coord_dict\"].apply(lambda d: d[\"x\"])\ntrain_localizations[\"y\"] = train_localizations[\"coord_dict\"].apply(lambda d: d[\"y\"])\n\n# Top 4-5 most frequent locations\ntop_locations = train_localizations[\"location\"].value_counts().head(5).index\n\n# Plot heatmap per location\nfor loc in top_locations:\n    subset = train_localizations[train_localizations[\"location\"] == loc]\n    plt.figure(figsize=(6, 5))\n    sns.kdeplot(data=subset, x=\"x\", y=\"y\", fill=True, cmap=\"crest\")\n    plt.title(f\"🧭 Spatial Heatmap for {loc}\")\n    plt.xlabel(\"X Coordinate\")\n    plt.ylabel(\"Y Coordinate\")\n    plt.tight_layout()\n    plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:04:20.363047Z","iopub.execute_input":"2025-07-29T13:04:20.363641Z","iopub.status.idle":"2025-07-29T13:04:22.423265Z","shell.execute_reply.started":"2025-07-29T13:04:20.363618Z","shell.execute_reply":"2025-07-29T13:04:22.422564Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#C77CFF\">🔄 Comparing Aneurysm Frequencies Across Datasets</h2> <h4>How consistent are the labels and localized annotations?</h4> This side-by-side bar chart contrasts location frequencies between `train.csv` and `train_localizers.csv`. Minor mismatches may suggest annotation noise or underreported aneurysms.","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train.csv\")\n# Frequency from train.csv\nlabel_freq = train[location_cols].sum().sort_values(ascending=False).rename(\"train.csv\")\n\n# Frequency from localizations.csv\nlocalizer_freq = train_localizations[\"location\"].value_counts().rename(\"train_localizations.csv\")\n\n# Combine\nfreq_comparison = pd.concat([label_freq, localizer_freq], axis=1).fillna(0).astype(int)\n\n# Bar plot\nfreq_comparison.plot(kind=\"bar\", figsize=(14, 6), color=['#00BFC4',  '#C77CFF'])\nplt.title(\"🔄 Frequency Comparison of Aneurysm Locations\")\nplt.ylabel(\"Count\")\nplt.xticks(rotation=45, ha=\"right\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:04:50.972359Z","iopub.execute_input":"2025-07-29T13:04:50.97293Z","iopub.status.idle":"2025-07-29T13:04:51.304614Z","shell.execute_reply.started":"2025-07-29T13:04:50.972905Z","shell.execute_reply":"2025-07-29T13:04:51.303899Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#00BFC4\">🧮 CT Slice Count per Series</h2> \n<h4>How deep is each scan?</h4> \nThis summary shows how many DICOM slices are present per series. The average is <b>{229.798638}</b> slices, with a range from <b>{1.0000}</b> to <b>{{1441.000000}}</b>. Scans with fewer slices may lack anatomical detail, affecting 3D model accuracy.\n","metadata":{}},{"cell_type":"code","source":"import os\nimport pandas as pd\n\nseries_dir = '/kaggle/input/rsna-intracranial-aneurysm-detection/series'\nseries_folders = [f for f in os.listdir(series_dir) if os.path.isdir(os.path.join(series_dir, f))]\n\nseries_counts = []\nfor folder in series_folders:\n    dcm_files = os.listdir(os.path.join(series_dir, folder))\n    series_counts.append({'SeriesInstanceUID': folder, 'NumSlices': len(dcm_files)})\n\ndf_series = pd.DataFrame(series_counts)\ndf_series.describe()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:05:43.76049Z","iopub.execute_input":"2025-07-29T13:05:43.760758Z","iopub.status.idle":"2025-07-29T13:09:37.123276Z","shell.execute_reply.started":"2025-07-29T13:05:43.76074Z","shell.execute_reply":"2025-07-29T13:09:37.122651Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#F8766D\">🧭 DICOM Voxel Orientation & Spacing</h2> \n<h4>How is the 3D scan structured?</h4> \nThis snippet reveals the scan's physical properties: <b>voxel spacing</b> (pixel size in mm), <b>slice thickness</b>, and <b>patient orientation</b>. Understanding these helps convert 2D slices into accurate 3D volumes for model input.\n","metadata":{}},{"cell_type":"code","source":"import pydicom\n\nsample_path = os.path.join(series_dir, series_folders[0])\nsample_file = os.listdir(sample_path)[0]\ndcm = pydicom.dcmread(os.path.join(sample_path, sample_file))\n\nprint(f\"Orientation: {dcm.ImageOrientationPatient}\")\nprint(f\"Voxel spacing: {dcm.PixelSpacing}\")\nprint(f\"Slice Thickness: {dcm.SliceThickness}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:09:37.124457Z","iopub.execute_input":"2025-07-29T13:09:37.124717Z","iopub.status.idle":"2025-07-29T13:09:37.63045Z","shell.execute_reply.started":"2025-07-29T13:09:37.124699Z","shell.execute_reply":"2025-07-29T13:09:37.629827Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#7CAE00\">🖼️ Scroll Through DICOM Slices</h2>\n<h4>Peek inside a brain scan — one slice at a time.</h4>\nThis interactive viewer lets you scroll through axial slices of a CT/MR scan using a slider. Useful for verifying data quality, orientation, and anatomical features before modeling.\n","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport matplotlib.widgets as widgets\nimport numpy as np\n\ndef load_series(series_path):\n    files = sorted(os.listdir(series_path), key=lambda x: pydicom.dcmread(os.path.join(series_path, x)).InstanceNumber)\n    images = [pydicom.dcmread(os.path.join(series_path, f)).pixel_array for f in files]\n    return np.stack(images)\n\nvolume = load_series(os.path.join(series_dir, series_folders[0]))\n\n# Scrollable plot\nfrom ipywidgets import interact\n@interact(slice=(0, volume.shape[0]-1))\ndef show_slice(slice=0):\n    plt.imshow(volume[slice], cmap='gray')\n    plt.axis('off')\n    plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:09:37.630993Z","iopub.execute_input":"2025-07-29T13:09:37.631168Z","iopub.status.idle":"2025-07-29T13:09:39.744378Z","shell.execute_reply.started":"2025-07-29T13:09:37.631153Z","shell.execute_reply":"2025-07-29T13:09:39.743861Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#F8766D\">📐 Common DICOM Image Resolutions</h2>\n<h4>Do all brain scans come in the same shape and size?</h4>\nThis horizontal bar chart shows the most frequent image resolutions across series. Identifying common shapes helps standardize preprocessing pipelines like resizing or padding.\n","metadata":{}},{"cell_type":"code","source":"import os\nimport pydicom\n\n# Count slices and shape per series\ndicom_dir = '/kaggle/input/rsna-intracranial-aneurysm-detection/series'\nseries_stats = []\n\nfor series_id in os.listdir(dicom_dir)[:50]:  # limit for speed\n    series_path = os.path.join(dicom_dir, series_id)\n    files = os.listdir(series_path)\n    num_slices = len(files)\n    sample_dcm = pydicom.dcmread(os.path.join(series_path, files[0]))\n    shape = (sample_dcm.Rows, sample_dcm.Columns)\n    series_stats.append({\"SeriesInstanceUID\": series_id, \"Slices\": num_slices, \"Shape\": shape})\n\npd.DataFrame(series_stats).value_counts(\"Shape\").plot(kind=\"barh\", title=\"🖼 Common Image Resolutions\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:09:39.745459Z","iopub.execute_input":"2025-07-29T13:09:39.745673Z","iopub.status.idle":"2025-07-29T13:09:40.806167Z","shell.execute_reply.started":"2025-07-29T13:09:39.745656Z","shell.execute_reply":"2025-07-29T13:09:40.805501Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#00BFC4\">📏 Slice Thickness in Brain CT Scans</h2>\n<h4>How deep is each CT slice through the brain?</h4>\nThis histogram illustrates the variation in slice thickness across patient scans, crucial for accurate 3D reconstruction and model generalization.\n","metadata":{}},{"cell_type":"code","source":"voxel_data = []\n\nfor series_id in os.listdir(dicom_dir)[:50]:\n    slices = []\n    for f in sorted(os.listdir(os.path.join(dicom_dir, series_id))):\n        path = os.path.join(dicom_dir, series_id, f)\n        dcm = pydicom.dcmread(path)\n        slices.append(dcm)\n\n    try:\n        spacing = slices[0].PixelSpacing\n        thickness = float(slices[0].SliceThickness)\n        voxel_data.append({\n            \"SeriesInstanceUID\": series_id,\n            \"PixelSpacingX\": spacing[0],\n            \"PixelSpacingY\": spacing[1],\n            \"SliceThickness\": thickness,\n            \"NumSlices\": len(slices)\n        })\n    except:\n        continue\n\nvoxel_df = pd.DataFrame(voxel_data)\nsns.histplot(voxel_df[\"SliceThickness\"], bins=20)\nplt.title(\"📐 Slice Thickness Distribution\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:09:40.806877Z","iopub.execute_input":"2025-07-29T13:09:40.807047Z","iopub.status.idle":"2025-07-29T13:11:15.386881Z","shell.execute_reply.started":"2025-07-29T13:09:40.807033Z","shell.execute_reply":"2025-07-29T13:11:15.386094Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#C77CFF\">🧠 Vessel Segmentations: Coverage Check</h2>\n<h4>How many scans come with annotated vessels?</h4>\nOnly a small subset of CT series (~{seg_percent:.2f}%) include segmentation masks, which can limit supervised learning for vessel or aneurysm detection tasks.\n","metadata":{}},{"cell_type":"code","source":"seg_dir = \"/kaggle/input/rsna-intracranial-aneurysm-detection/segmentations\"\nseg_series = [f.replace(\".nii.gz\", \"\") for f in os.listdir(seg_dir)]\nprint(\"Segmented Series:\", len(seg_series))\n\n# % of series with segmentation\nseg_percent = len(seg_series) / len(os.listdir(dicom_dir)) * 100\nprint(f\"{seg_percent:.2f}% of series have vessel segmentation.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:11:15.387689Z","iopub.execute_input":"2025-07-29T13:11:15.388473Z","iopub.status.idle":"2025-07-29T13:11:15.416928Z","shell.execute_reply.started":"2025-07-29T13:11:15.388448Z","shell.execute_reply":"2025-07-29T13:11:15.416219Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h2 style=\"color:#00BFC4\">🧠 Baseline 3D CNN for Aneurysm Localization</h2>\n<h4>Compact and efficient model to kickstart voxel-wise aneurysm detection</h4>\nThis 3D CNN ingests normalized DICOM volumes and outputs multi-label predictions across 14 possible aneurysm locations plus presence. Designed for speed and flexibility.\n\n<h3 style=\"color:#F8766D\">🗂 DICOM Volume Preprocessing Pipeline</h3>\n<h4>From unordered slices to uniform, normalized 3D tensors</h4>\nVolumes are rescaled, clipped (1st–99th percentiles), normalized, and resized to 64³ using `ndimage.zoom`. Robust preprocessing ensures clean inputs for deep learning models.\n","metadata":{}},{"cell_type":"code","source":"import os\nimport shutil\nimport gc\nfrom collections import defaultdict\nfrom typing import Tuple, List\nimport warnings\nwarnings.filterwarnings('ignore')\n\nimport numpy as np\nimport pandas as pd\nimport polars as pl\nimport pydicom\nfrom scipy import ndimage\nfrom sklearn.preprocessing import StandardScaler\n\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.utils.data import Dataset, DataLoader\nimport torch.optim as optim\n\nimport kaggle_evaluation.rsna_inference_server\n\n# Competition constants\nID_COL = 'SeriesInstanceUID'\nLABEL_COLS = [\n    'Left Infraclinoid Internal Carotid Artery',\n    'Right Infraclinoid Internal Carotid Artery',\n    'Left Supraclinoid Internal Carotid Artery',\n    'Right Supraclinoid Internal Carotid Artery',\n    'Left Middle Cerebral Artery',\n    'Right Middle Cerebral Artery',\n    'Anterior Communicating Artery',\n    'Left Anterior Cerebral Artery',\n    'Right Anterior Cerebral Artery',\n    'Left Posterior Communicating Artery',\n    'Right Posterior Communicating Artery',\n    'Basilar Tip',\n    'Other Posterior Circulation',\n    'Aneurysm Present',\n]\n\nDICOM_TAG_ALLOWLIST = [\n    'BitsAllocated', 'BitsStored', 'Columns', 'FrameOfReferenceUID', 'HighBit',\n    'ImageOrientationPatient', 'ImagePositionPatient', 'InstanceNumber', 'Modality',\n    'PatientID', 'PhotometricInterpretation', 'PixelRepresentation', 'PixelSpacing',\n    'PlanarConfiguration', 'RescaleIntercept', 'RescaleSlope', 'RescaleType', 'Rows',\n    'SOPClassUID', 'SOPInstanceUID', 'SamplesPerPixel', 'SliceThickness',\n    'SpacingBetweenSlices', 'StudyInstanceUID', 'TransferSyntaxUID',\n]\n\n# Model configuration\nTARGET_SIZE = (64, 64, 64)  # Reduced size for memory efficiency\nDEVICE = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n\nclass DICOMProcessor:\n    \"\"\"Process DICOM series into normalized 3D volumes\"\"\"\n    \n    def __init__(self, target_size: Tuple[int, int, int] = TARGET_SIZE):\n        self.target_size = target_size\n        self.scaler = StandardScaler()\n    \n    def load_dicom_series(self, series_path: str) -> np.ndarray:\n        \"\"\"Load and process a DICOM series into a 3D volume\"\"\"\n        try:\n            # Get all DICOM files\n            dicom_files = []\n            for root, _, files in os.walk(series_path):\n                for file in files:\n                    if file.endswith('.dcm'):\n                        dicom_files.append(os.path.join(root, file))\n            \n            if not dicom_files:\n                raise ValueError(f\"No DICOM files found in {series_path}\")\n            \n            # Load DICOMs and sort by instance number\n            dicoms = []\n            for filepath in dicom_files:\n                try:\n                    ds = pydicom.dcmread(filepath, force=True)\n                    if hasattr(ds, 'PixelData'):\n                        dicoms.append((ds, filepath))\n                except Exception as e:\n                    print(f\"Error reading {filepath}: {e}\")\n                    continue\n            \n            if not dicoms:\n                raise ValueError(f\"No valid DICOM files with pixel data in {series_path}\")\n            \n            # Sort by instance number\n            dicoms.sort(key=lambda x: getattr(x[0], 'InstanceNumber', 0))\n            \n            # Extract volume\n            volume_slices = []\n            for ds, _ in dicoms:\n                try:\n                    # Get pixel array\n                    pixel_array = ds.pixel_array.astype(np.float32)\n                    \n                    # Apply rescale if available\n                    if hasattr(ds, 'RescaleSlope') and hasattr(ds, 'RescaleIntercept'):\n                        slope = float(ds.RescaleSlope)\n                        intercept = float(ds.RescaleIntercept)\n                        pixel_array = pixel_array * slope + intercept\n                    \n                    volume_slices.append(pixel_array)\n                except Exception as e:\n                    print(f\"Error processing slice: {e}\")\n                    continue\n            \n            if not volume_slices:\n                raise ValueError(\"No valid slices extracted\")\n            \n            # Stack into 3D volume\n            volume = np.stack(volume_slices, axis=0)  # Shape: (depth, height, width)\n            \n            # Normalize and resize\n            volume = self.preprocess_volume(volume)\n            \n            return volume\n            \n        except Exception as e:\n            print(f\"Error processing series {series_path}: {e}\")\n            # Return zeros if processing fails\n            return np.zeros(self.target_size, dtype=np.float32)\n    \n    def preprocess_volume(self, volume: np.ndarray) -> np.ndarray:\n        \"\"\"Preprocess 3D volume: normalize, clip, resize\"\"\"\n        # Handle potential issues\n        if volume.size == 0:\n            return np.zeros(self.target_size, dtype=np.float32)\n        \n        # Clip extreme values (robust to outliers)\n        p1, p99 = np.percentile(volume, [1, 99])\n        volume = np.clip(volume, p1, p99)\n        \n        # Normalize to [0, 1]\n        volume_min, volume_max = volume.min(), volume.max()\n        if volume_max > volume_min:\n            volume = (volume - volume_min) / (volume_max - volume_min)\n        \n        # Resize to target size\n        if volume.shape != self.target_size:\n            zoom_factors = [\n                self.target_size[i] / volume.shape[i] for i in range(3)\n            ]\n            volume = ndimage.zoom(volume, zoom_factors, order=1)\n        \n        return volume.astype(np.float32)\n\nclass Simple3DCNN(nn.Module):\n    \"\"\"Lightweight 3D CNN for aneurysm detection\"\"\"\n    \n    def __init__(self, num_classes: int = len(LABEL_COLS)):\n        super(Simple3DCNN, self).__init__()\n        \n        # 3D Convolutional layers\n        self.conv1 = nn.Conv3d(1, 16, kernel_size=3, padding=1)\n        self.pool1 = nn.MaxPool3d(2)\n        self.conv2 = nn.Conv3d(16, 32, kernel_size=3, padding=1)\n        self.pool2 = nn.MaxPool3d(2)\n        self.conv3 = nn.Conv3d(32, 64, kernel_size=3, padding=1)\n        self.pool3 = nn.MaxPool3d(2)\n        self.conv4 = nn.Conv3d(64, 128, kernel_size=3, padding=1)\n        self.pool4 = nn.MaxPool3d(2)\n        \n        # Adaptive pooling to handle variable sizes\n        self.adaptive_pool = nn.AdaptiveAvgPool3d((2, 2, 2))\n        \n        # Fully connected layers\n        self.fc1 = nn.Linear(128 * 2 * 2 * 2, 256)\n        self.dropout1 = nn.Dropout(0.5)\n        self.fc2 = nn.Linear(256, 128)\n        self.dropout2 = nn.Dropout(0.3)\n        self.fc3 = nn.Linear(128, num_classes)\n        \n        # Batch normalization\n        self.bn1 = nn.BatchNorm3d(16)\n        self.bn2 = nn.BatchNorm3d(32)\n        self.bn3 = nn.BatchNorm3d(64)\n        self.bn4 = nn.BatchNorm3d(128)\n        \n    def forward(self, x):\n        # Input shape: (batch_size, 1, depth, height, width)\n        x = self.pool1(F.relu(self.bn1(self.conv1(x))))\n        x = self.pool2(F.relu(self.bn2(self.conv2(x))))\n        x = self.pool3(F.relu(self.bn3(self.conv3(x))))\n        x = self.pool4(F.relu(self.bn4(self.conv4(x))))\n        \n        # Adaptive pooling\n        x = self.adaptive_pool(x)\n        \n        # Flatten\n        x = x.view(x.size(0), -1)\n        \n        # Fully connected layers\n        x = F.relu(self.fc1(x))\n        x = self.dropout1(x)\n        x = F.relu(self.fc2(x))\n        x = self.dropout2(x)\n        x = self.fc3(x)\n        \n        return torch.sigmoid(x)\n\nclass AneurysmDataset(Dataset):\n    \"\"\"Dataset for loading training data\"\"\"\n    \n    def __init__(self, data_df: pd.DataFrame, series_dir: str, processor: DICOMProcessor):\n        self.data_df = data_df\n        self.series_dir = series_dir\n        self.processor = processor\n        \n    def __len__(self):\n        return len(self.data_df)\n    \n    def __getitem__(self, idx):\n        row = self.data_df.iloc[idx]\n        series_id = row[ID_COL]\n        \n        # Load volume\n        series_path = os.path.join(self.series_dir, series_id)\n        volume = self.processor.load_dicom_series(series_path)\n        \n        # Get labels\n        labels = row[LABEL_COLS].values.astype(np.float32)\n        \n        # Convert to tensor and add channel dimension\n        volume_tensor = torch.from_numpy(volume).unsqueeze(0)  # Add channel dim\n        labels_tensor = torch.from_numpy(labels)\n        \n        return volume_tensor, labels_tensor\n\n# Global model and processor\nmodel = None\nprocessor = None\n\ndef initialize_model():\n    \"\"\"Initialize model and processor (called once)\"\"\"\n    global model, processor\n    \n    if model is not None:\n        return\n    \n    print(\"Initializing model...\")\n    processor = DICOMProcessor(TARGET_SIZE)\n    model = Simple3DCNN(num_classes=len(LABEL_COLS))\n    \n    # Load pre-trained weights if available\n    try:\n        if os.path.exists('/kaggle/input/model_weights.pth'):\n            model.load_state_dict(torch.load('/kaggle/input/model_weights.pth', map_location='cpu'))\n            print(\"Loaded pre-trained weights\")\n        else:\n            print(\"No pre-trained weights found, using random initialization\")\n    except Exception as e:\n        print(f\"Error loading weights: {e}\")\n    \n    model.to(DEVICE)\n    model.eval()\n    print(f\"Model initialized on {DEVICE}\")\n\ndef predict(series_path: str) -> pl.DataFrame:\n    \"\"\"Make prediction for a single series\"\"\"\n    \n    # Initialize model on first call\n    initialize_model()\n    \n    series_id = os.path.basename(series_path)\n    \n    try:\n        # Process the DICOM series\n        volume = processor.load_dicom_series(series_path)\n        \n        # Convert to tensor and add batch dimension\n        volume_tensor = torch.from_numpy(volume).unsqueeze(0).unsqueeze(0)  # (1, 1, D, H, W)\n        volume_tensor = volume_tensor.to(DEVICE)\n        \n        # Make prediction\n        with torch.no_grad():\n            predictions = model(volume_tensor)\n            predictions = predictions.cpu().numpy().flatten()\n        \n        # Create result DataFrame\n        result_data = [[series_id] + predictions.tolist()]\n        result_df = pl.DataFrame(\n            data=result_data,\n            schema=[ID_COL] + LABEL_COLS,\n            orient='row'\n        )\n        \n        # Clean up memory\n        del volume_tensor\n        torch.cuda.empty_cache()\n        gc.collect()\n        \n    except Exception as e:\n        print(f\"Error predicting for series {series_id}: {e}\")\n        # Return baseline predictions (0.5 for all classes)\n        result_data = [[series_id] + [0.5] * len(LABEL_COLS)]\n        result_df = pl.DataFrame(\n            data=result_data,\n            schema=[ID_COL] + LABEL_COLS,\n            orient='row'\n        )\n    \n    # Mandatory cleanup\n    shutil.rmtree('/kaggle/shared', ignore_errors=True)\n    \n    return result_df.drop(ID_COL)\n\ndef train_model(train_df_path: str, series_dir: \"/kaggle/input/rsna-intracranial-aneurysm-detection/series\", num_epochs: int = 50, batch_size: int = 32):\n    \"\"\"Training function (for reference - would be run separately)\"\"\"\n    \n    # Load training data\n    train_df = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train.csv\")\n    \n    # Initialize components\n    processor = DICOMProcessor(TARGET_SIZE)\n    model = Simple3DCNN(num_classes=len(LABEL_COLS))\n    model.to(DEVICE)\n    \n    # Create dataset and dataloader\n    dataset = AneurysmDataset(train_df, series_dir, processor)\n    dataloader = DataLoader(dataset, batch_size=batch_size, shuffle=True, num_workers=0)\n    \n    # Loss and optimizer\n    criterion = nn.BCELoss()\n    optimizer = optim.Adam(model.parameters(), lr=0.001, weight_decay=1e-5)\n    scheduler = optim.lr_scheduler.ReduceLROnPlateau(optimizer, mode='min', patience=2)\n    \n    # Training loop\n    model.train()\n    for epoch in range(num_epochs):\n        total_loss = 0\n        num_batches = 0\n        \n        for batch_idx, (volumes, labels) in enumerate(dataloader):\n            volumes, labels = volumes.to(DEVICE), labels.to(DEVICE)\n            \n            optimizer.zero_grad()\n            outputs = model(volumes)\n            loss = criterion(outputs, labels)\n            loss.backward()\n            optimizer.step()\n            \n            total_loss += loss.item()\n            num_batches += 1\n            \n            if batch_idx % 10 == 0:\n                print(f'Epoch {epoch+1}/{num_epochs}, Batch {batch_idx+1}, Loss: {loss.item():.4f}')\n        \n        avg_loss = total_loss / num_batches\n        scheduler.step(avg_loss)\n        print(f'Epoch {epoch+1}/{num_epochs} completed, Average Loss: {avg_loss:.4f}')\n    \n    # Save model\n    torch.save(model.state_dict(), 'model_weights.pth')\n    print(\"Model saved to model_weights.pth\")\n\n# Competition server setup\ninference_server = kaggle_evaluation.rsna_inference_server.RSNAInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway()\n    display(pl.read_parquet('/kaggle/working/submission.parquet'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-29T13:14:55.852796Z","iopub.execute_input":"2025-07-29T13:14:55.853124Z","iopub.status.idle":"2025-07-29T13:15:20.120886Z","shell.execute_reply.started":"2025-07-29T13:14:55.853099Z","shell.execute_reply":"2025-07-29T13:15:20.120136Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}