{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":99552,"databundleVersionId":13851420,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 1. Import & Settings","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport matplotlib.pyplot as plt\nfrom ipywidgets import interact, IntSlider\nimport seaborn as sns\nimport plotly.express as px\nimport plotly.io as pio\nimport pydicom\nimport nibabel as nib\nfrom glob import glob\nimport os\n\npio.renderers.default = \"kaggle\" \n\nimport warnings\nwarnings.filterwarnings(\"ignore\", category=FutureWarning)\n\n# Visualisation settings\nplt.style.use('seaborn-v0_8-whitegrid')\nsns.set_palette('viridis')\npd.options.display.float_format = '{:,.4f}'.format\nplt.rcParams['figure.figsize'] = (18, 8)\nplt.rcParams['axes.titlesize'] = 16\nplt.rcParams['axes.labelsize'] = 14\n\n\ndata_dir = '/kaggle/input/rsna-intracranial-aneurysm-detection'\ntrain_df = pd.read_csv(os.path.join(data_dir, 'train.csv'))\nlocalizers_df = pd.read_csv(os.path.join(data_dir, 'train_localizers.csv'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-29T15:28:56.291323Z","iopub.execute_input":"2025-09-29T15:28:56.292181Z","iopub.status.idle":"2025-09-29T15:28:56.328624Z","shell.execute_reply.started":"2025-09-29T15:28:56.292147Z","shell.execute_reply":"2025-09-29T15:28:56.327491Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Initial inspection:","metadata":{"execution":{"iopub.status.busy":"2025-09-29T15:13:54.184238Z","iopub.execute_input":"2025-09-29T15:13:54.185153Z","iopub.status.idle":"2025-09-29T15:13:54.190171Z","shell.execute_reply.started":"2025-09-29T15:13:54.185119Z","shell.execute_reply":"2025-09-29T15:13:54.188958Z"}}},{"cell_type":"code","source":"print(\"Train DataFrame Info:\")\ntrain_df.info()\nprint(\"\\n\" + \"=\"*50 + \"\\n\")\nprint(\"Localizers DataFrame Info:\")\nlocalizers_df.info()\n\nprint(\"\\nTrain DataFrame Head:\")\ndisplay(train_df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-29T15:14:07.044384Z","iopub.execute_input":"2025-09-29T15:14:07.044716Z","iopub.status.idle":"2025-09-29T15:14:07.122556Z","shell.execute_reply.started":"2025-09-29T15:14:07.044693Z","shell.execute_reply":"2025-09-29T15:14:07.121439Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 2. Train.csv analysis","metadata":{}},{"cell_type":"markdown","source":"## Aneurysm Present\nIt is target variable according to the description","metadata":{"execution":{"iopub.status.busy":"2025-09-29T15:16:19.106638Z","iopub.execute_input":"2025-09-29T15:16:19.107087Z","iopub.status.idle":"2025-09-29T15:16:19.11213Z","shell.execute_reply.started":"2025-09-29T15:16:19.106993Z","shell.execute_reply":"2025-09-29T15:16:19.111028Z"}}},{"cell_type":"code","source":"aneurysm_counts = train_df['Aneurysm Present'].value_counts()\n\nfig = px.pie(\n    names=['No Aneurysm', 'Aneurysm Present'],\n    values=aneurysm_counts.values,\n    title='Class Balance for \"Aneurysm Present\"',\n    hole=0.3,\n    color_discrete_sequence=px.colors.sequential.RdBu\n)\nfig.show()\n\nprint(f\"Proportion of patients with aneurysm: {aneurysm_counts[1] / len(train_df):.2%}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-29T15:19:49.333613Z","iopub.execute_input":"2025-09-29T15:19:49.333949Z","iopub.status.idle":"2025-09-29T15:19:49.382841Z","shell.execute_reply.started":"2025-09-29T15:19:49.333926Z","shell.execute_reply":"2025-09-29T15:19:49.381438Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Demographic data (PatientAge, PatientSex)","metadata":{"execution":{"iopub.status.busy":"2025-09-29T15:17:10.602752Z","iopub.execute_input":"2025-09-29T15:17:10.60311Z","iopub.status.idle":"2025-09-29T15:17:10.607835Z","shell.execute_reply.started":"2025-09-29T15:17:10.603083Z","shell.execute_reply":"2025-09-29T15:17:10.606903Z"}}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(16, 6))\n\nsns.histplot(data=train_df, x='PatientAge', hue='Aneurysm Present', kde=True, ax=axes[0])\naxes[0].set_title('Age Distribution by Aneurysm Presence')\n\nsns.countplot(data=train_df, x='PatientSex', hue='Aneurysm Present', ax=axes[1])\naxes[1].set_title('Sex Distribution by Aneurysm Presence')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-29T15:20:02.298164Z","iopub.execute_input":"2025-09-29T15:20:02.298514Z","iopub.status.idle":"2025-09-29T15:20:03.068989Z","shell.execute_reply.started":"2025-09-29T15:20:02.298492Z","shell.execute_reply":"2025-09-29T15:20:03.067745Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Modality","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(\n    train_df,\n    x='Modality',\n    color='Aneurysm Present',\n    barmode='group',\n    title='Modality Distribution by Aneurysm Presence',\n    text_auto=True\n)\nfig.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-29T15:20:08.025216Z","iopub.execute_input":"2025-09-29T15:20:08.025997Z","iopub.status.idle":"2025-09-29T15:20:08.091907Z","shell.execute_reply.started":"2025-09-29T15:20:08.025971Z","shell.execute_reply":"2025-09-29T15:20:08.090828Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Multiclass labels (13 locations)","metadata":{}},{"cell_type":"code","source":"location_cols = [col for col in train_df.columns if 'Artery' in col or 'Tip' in col or 'Circulation' in col]\nlocation_counts = train_df[location_cols].sum().sort_values(ascending=False)\n\nfig = px.bar(\n    x=location_counts.values,\n    y=location_counts.index,\n    orientation='h',\n    title='Number of Aneurysms by Location',\n    labels={'y': 'Location', 'x': 'Count'}\n)\nfig.show()\n\n# Let's check if one patient has aneurysms in several places.\ntrain_df['num_aneurysms'] = train_df[location_cols].sum(axis=1)\nprint(\"Distribution of the number of aneurysms per series:\")\nprint(train_df[train_df['Aneurysm Present'] == 1]['num_aneurysms'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-29T15:20:40.713007Z","iopub.execute_input":"2025-09-29T15:20:40.713381Z","iopub.status.idle":"2025-09-29T15:20:40.785038Z","shell.execute_reply.started":"2025-09-29T15:20:40.713334Z","shell.execute_reply":"2025-09-29T15:20:40.783988Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 3. Visual Image Analysis","metadata":{"execution":{"iopub.status.busy":"2025-09-29T15:21:29.402338Z","iopub.execute_input":"2025-09-29T15:21:29.402711Z","iopub.status.idle":"2025-09-29T15:21:29.407696Z","shell.execute_reply.started":"2025-09-29T15:21:29.40269Z","shell.execute_reply":"2025-09-29T15:21:29.40643Z"}}},{"cell_type":"markdown","source":"## Comparison of Modalities: CTA vs. MRA vs. MRI","metadata":{}},{"cell_type":"code","source":"def show_modality_comparison(df, data_dir):\n    modalities = df['Modality'].unique()\n    \n    fig, axes = plt.subplots(1, len(modalities), figsize=(20, 7))\n    if len(modalities) == 1:\n        axes = [axes]\n        \n    fig.suptitle('Visual Comparison of Imaging Modalities (Corrected)', fontsize=20)\n\n    for i, mod in enumerate(modalities):\n        sample_series_uid = df[df['Modality'] == mod]['SeriesInstanceUID'].iloc[0]\n        series_path = os.path.join(data_dir, 'series', sample_series_uid)\n        \n        dicom_files = glob(os.path.join(series_path, '*.dcm'))\n        \n        if not dicom_files:\n            print(f\"No DICOM files found for series {sample_series_uid} (Modality: {mod})\")\n            continue\n            \n        dcm_path = dicom_files[0]\n        dcm = pydicom.dcmread(dcm_path)\n        \n        pixel_data = dcm.pixel_array\n        \n        if pixel_data.ndim == 3:\n            mid_frame_idx = pixel_data.shape[0] // 2\n            image_to_show = pixel_data[mid_frame_idx, :, :]\n            print(f\"Modality {mod}: Detected multi-frame DICOM with shape {pixel_data.shape}. Displaying frame {mid_frame_idx}.\")\n        else:\n            image_to_show = pixel_data\n            print(f\"Modality {mod}: Detected single-frame DICOM with shape {pixel_data.shape}.\")\n\n        ax = axes[i]\n        ax.imshow(image_to_show, cmap='gray')\n        ax.set_title(f'Modality: {mod}')\n        ax.axis('off')\n        \n    plt.tight_layout(rect=[0, 0.03, 1, 0.95])\n    plt.show()\n\nshow_modality_comparison(train_df, data_dir)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-29T15:24:09.603067Z","iopub.execute_input":"2025-09-29T15:24:09.603438Z","iopub.status.idle":"2025-09-29T15:24:11.02959Z","shell.execute_reply.started":"2025-09-29T15:24:09.60341Z","shell.execute_reply":"2025-09-29T15:24:11.028403Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### **CTA (Computed Tomography Angiography):**\nHow it looks: Very contrasting. The vessels are bright white due to the contrast agent injected, and the bones of the skull are also very bright. The soft tissues of the brain are gray.\nInsight for the model: The model can easily confuse a bright bone with a vessel. Good preprocessing is needed (for example, removing the skull or using RescaleIntercept/Slope to move to Hounsfield units where bones have values > 400-500 HU).\n#### **MRA (Magnetic Resonance Angiography):**\nWhat it looks like: There are many MRA techniques (for example, Time-of-Flight). Vessels usually look bright against the dark background of suppressed static tissues. The contrast is not as sharp as on CT, the image is more \"soft\". The bones are not visible.\nInsight for the model: There are fewer distractions (no bones), but the signal intensity in the vessels may be more heterogeneous.\n#### **MRI (T1/T2):**\nWhat it looks like: A standard MRI scan. The focus here is on the differentiation of brain tissues (gray/white matter), rather than on blood vessels. Vessels can be seen as \"voids\" (flow voids) or can be enhanced by contrast.\nInsight for the model: Finding an aneurysm is the most difficult thing here. The task is almost impossible without understanding the anatomy and context. Perhaps these data serve to complicate the task and check the robustness of the model.","metadata":{}},{"cell_type":"markdown","source":"## Visualization of an aneurysm","metadata":{}},{"cell_type":"code","source":"def load_scan(series_path):\n    \"\"\"\n    Downloads DICOM series from a folder.\n    It can process both folders with multiple files and a single multiframe file.\n    Sorts slices by their position in space.\n    \"\"\"\n    dicom_files = glob(os.path.join(series_path, '*.dcm'))\n    \n    if not dicom_files:\n        print(f\"No DICOM files found in {series_path}\")\n        return None, None\n        \n    if len(dicom_files) == 1:\n        dcm = pydicom.dcmread(dicom_files[0])\n        if hasattr(dcm, 'NumberOfFrames') and dcm.NumberOfFrames > 1:\n            print(f\"Detected multi-frame DICOM with {dcm.NumberOfFrames} frames.\")\n    \n            z_positions = []\n            if 'PerFrameFunctionalGroupsSequence' in dcm:\n                for frame_info in dcm.PerFrameFunctionalGroupsSequence:\n                    z_positions.append(frame_info.PlanePositionSequence[0].ImagePositionPatient[2])\n            \n            pixel_data = dcm.pixel_array\n            \n            if z_positions and len(z_positions) == pixel_data.shape[0]:\n                sorted_indices = np.argsort(z_positions)\n                pixel_data = pixel_data[sorted_indices]\n            \n            return pixel_data, dcm\n\n    slices = [pydicom.dcmread(f) for f in dicom_files]\n    slices.sort(key=lambda x: float(x.ImagePositionPatient[2]))\n    \n    shape = slices[0].pixel_array.shape\n    if not all(s.pixel_array.shape == shape for s in slices):\n        print(\"Warning: Slices have different shapes. This may cause issues.\")\n        return None, None\n\n    volume = np.stack([s.pixel_array for s in slices])\n    \n    return volume, slices[0]\n\nseries_to_show_uid = train_df['SeriesInstanceUID'].iloc[5]\nseries_path = os.path.join(data_dir, 'series', series_to_show_uid)\n\nvolume, meta = load_scan(series_path)\n\nif volume is not None:\n    print(f\"Successfully loaded volume with shape: {volume.shape}\")\n    \n    def plot_slice(slice_index):\n        plt.figure(figsize=(8, 8))\n        plt.imshow(volume[slice_index, :, :], cmap='gray')\n        plt.title(f'Series: {series_to_show_uid[:15]}...\\nSlice: {slice_index}/{volume.shape[0]-1}')\n        plt.axis('off')\n        plt.show()\n\n    interact(\n        plot_slice,\n        slice_index=IntSlider(\n            min=0,\n            max=volume.shape[0] - 1,\n            step=1,\n            value=volume.shape[0] // 2,\n            description='Slice:',\n            continuous_update=False\n        )\n    );\nelse:\n    print(\"Could not load the volume for this series.\")\n\nif volume is not None:\n    mip_axial = np.max(volume, axis=0)\n    mip_coronal = np.max(volume, axis=1)\n    mip_sagittal = np.max(volume, axis=2)\n    \n    fig, axes = plt.subplots(1, 3, figsize=(21, 7))\n    fig.suptitle('Maximum Intensity Projections (MIP)', fontsize=20)\n    \n    axes[0].imshow(mip_axial, cmap='gray')\n    axes[0].set_title('Axial MIP (Top-Down View)')\n    axes[1].imshow(mip_coronal, cmap='gray')\n    axes[1].set_title('Coronal MIP (Front-Back View)')\n    axes[2].imshow(mip_sagittal, cmap='gray')\n    axes[2].set_title('Sagittal MIP (Side View)')\n    \n    for ax in axes:\n        ax.axis('off')\n        \n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-29T15:38:56.950045Z","iopub.execute_input":"2025-09-29T15:38:56.950424Z","iopub.status.idle":"2025-09-29T15:39:11.161465Z","shell.execute_reply.started":"2025-09-29T15:38:56.950394Z","shell.execute_reply":"2025-09-29T15:39:11.160217Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## DICOM Metadata Analysis\nThe scanner settings can greatly affect the image quality. Let's analyze their spread.","metadata":{}},{"cell_type":"code","source":"meta_data = []\nfor series_uid in train_df['SeriesInstanceUID'].head(200):\n    dcm_path = glob(os.path.join(data_dir, 'series', series_uid, '*.dcm'))[0]\n    dcm = pydicom.dcmread(dcm_path, stop_before_pixels=True)\n    \n    meta_data.append({\n        'SeriesInstanceUID': series_uid,\n        'PixelSpacing_row': dcm.PixelSpacing[0] if 'PixelSpacing' in dcm else None,\n        'PixelSpacing_col': dcm.PixelSpacing[1] if 'PixelSpacing' in dcm else None,\n        'SliceThickness': dcm.SliceThickness if 'SliceThickness' in dcm else None,\n        'RescaleIntercept': dcm.RescaleIntercept if 'RescaleIntercept' in dcm else None,\n        'RescaleSlope': dcm.RescaleSlope if 'RescaleSlope' in dcm else None,\n    })\nmeta_df = pd.DataFrame(meta_data).dropna()\n\nfig, axes = plt.subplots(2, 2, figsize=(15, 12))\nsns.histplot(meta_df['PixelSpacing_row'], kde=True, ax=axes[0, 0]).set_title('Pixel Spacing (X) Distribution')\nsns.histplot(meta_df['SliceThickness'], kde=True, ax=axes[0, 1]).set_title('Slice Thickness (Z) Distribution')\nsns.histplot(meta_df['RescaleIntercept'], kde=True, ax=axes[1, 0]).set_title('Rescale Intercept Distribution')\nsns.histplot(meta_df['RescaleSlope'], kde=True, ax=axes[1, 1]).set_title('Rescale Slope Distribution')\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-29T15:33:37.829331Z","iopub.execute_input":"2025-09-29T15:33:37.829706Z","iopub.status.idle":"2025-09-29T15:33:48.161014Z","shell.execute_reply.started":"2025-09-29T15:33:37.829683Z","shell.execute_reply":"2025-09-29T15:33:48.159931Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"PixelSpacing (resolution in the XY plane) and SliceThickness (resolution along the Z axis) have a wide range of values. This means that the scans were done on different hardware with different protocols. This confirms our conclusion No. 1: without resampling to a single physical resolution, the model will compare \"meters with kilograms\".","metadata":{}},{"cell_type":"markdown","source":"# 4. Data Quality Analysis","metadata":{}},{"cell_type":"markdown","source":"## Checking The Consistency Of Labels\nAneurysm Present must be 1 if and only if there is at least one label in the 13 location columns. Let's check it out.","metadata":{}},{"cell_type":"code","source":"location_cols = [col for col in train_df.columns if 'Artery' in col or 'Tip' in col or 'Circulation' in col]\nis_aneurysm_by_loc = (train_df[location_cols].sum(axis=1) > 0).astype(int)\nmismatch = (train_df['Aneurysm Present'] != is_aneurysm_by_loc).sum()\n\nprint(f\"The number of discrepancies between the 'Aneurysm Present' and the total by location: {mismatch}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-29T15:35:35.358086Z","iopub.execute_input":"2025-09-29T15:35:35.358441Z","iopub.status.idle":"2025-09-29T15:35:35.371127Z","shell.execute_reply.started":"2025-09-29T15:35:35.358418Z","shell.execute_reply":"2025-09-29T15:35:35.369202Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Where Are The Coordinates Lost?\nLet's see which aneurysm locations are most often missing coordinates in train_localizers.csv.","metadata":{"execution":{"iopub.status.busy":"2025-09-29T15:35:52.890998Z","iopub.execute_input":"2025-09-29T15:35:52.891345Z","iopub.status.idle":"2025-09-29T15:35:52.899345Z","shell.execute_reply.started":"2025-09-29T15:35:52.891322Z","shell.execute_reply":"2025-09-29T15:35:52.897677Z"}}},{"cell_type":"code","source":"aneurysm_df = train_df.query(\"`Aneurysm Present` == 1\").copy()\naneurysm_df['num_aneurysms'] = aneurysm_df[location_cols].sum(axis=1)\n\ntotal_loc_counts = aneurysm_df[location_cols].sum()\n\nloc_with_coords_counts = localizers_df['location'].value_counts()\n\nmissing_coords_stats = pd.DataFrame({'Total': total_loc_counts, 'WithCoords': loc_with_coords_counts}).fillna(0)\nmissing_coords_stats['Missing_Percent'] = 1 - (missing_coords_stats['WithCoords'] / missing_coords_stats['Total'])\nprint(\"Percentage of missing coordinates by location:\")\nprint(missing_coords_stats.sort_values('Missing_Percent', ascending=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-29T15:36:23.118031Z","iopub.execute_input":"2025-09-29T15:36:23.118346Z","iopub.status.idle":"2025-09-29T15:36:23.146189Z","shell.execute_reply.started":"2025-09-29T15:36:23.118325Z","shell.execute_reply":"2025-09-29T15:36:23.144774Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 5. Summary","metadata":{}},{"cell_type":"markdown","source":"1. **Main Data Characteristics:**\n- Target Variable (Aneurysm Present): Severe imbalance, ~35% of the series contain an aneurysm. This requires a weighted learning loss function.\n    - The metric (weighted AUC ROC) takes this into account, giving the Aneurysm Present target a weight of 13.\n     Multimodality: The data is represented by three main types of scans: CTA (CT angiography), MRA (MR angiography) and MRI (MRI). They are radically different visually (contrast, bone visibility). CTA dominates the dataset.\n    - Multi-label Task: In addition to the binary flag, there are 13 targets for aneurysm localization. One patient may have several aneurysms. Correlations between locations are weak.\n     Demographics: Aneurysms are more common in older age. Age can be a useful sign.\n2. **Image Features (DICOM):**\n- Aneurysms are small objects: Visually, they are tiny \"protrusions\" on blood vessels that are easy to miss. The head-on approach of classifying the entire 3D scan is ineffective.\n   - Anisotropy: DICOM file parameters (PixelSpacing, SliceThickness) vary greatly. This means that the physical dimensions of the voxels are not constant. Resampling to a single isotropic resolution (for example, 1x1x1 mm3) is a mandatory preprocessing step.\n   - Storage format: Series are stored in two ways: as a folder with multiple 2D files or as a single multi-frame DICOM file containing the entire 3D volume. The data loader must handle both scenarios.\n4. **Additional Data and Its Value:**\n- Localizers (train_localizers.csv): They contain 2D coordinates (x, y) for aneurysms.\n    - Problem 1 (Incompleteness): Coordinates are not available for all aneurysms (Weakly Supervised Learning problem).\n    - Problem 2 (Ambiguity): The Z-coordinate (slice number) is not specified for multiframe DICOM files, which complicates the use of these labels.\n    - Segmentation (segments/): Accurate 3D masks of vessels are provided for a subset of the data. This is an extremely valuable resource for teaching an auxiliary segmentation model that can dramatically narrow the search area for aneurysms.","metadata":{}}]}