{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":99552,"databundleVersionId":13441085,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":">  *First time? Me too!*\n> \n> I know this isn’t super polished.\n> This is a **baby-steps walkthrough** to show how to actually get started with a Kaggle competition, without going insane.\n\n> My motivation: I need to learn how to manipulate MRI scans for my thesis. 🧠\n","metadata":{}},{"cell_type":"markdown","source":"TASK : 1. Classification, check if aneurysm is present (train.csv)\n2. Pinpoint the location of the aneurysm (train_localizers.csv)\n","metadata":{}},{"cell_type":"markdown","source":"First read the data, conduct a EDA. I am one of those people who have heard EDA thousand times but have no clue what I am actually supposed to do. Or what should I plot, what should I actually visualise. So here it goes: See the columns of the data, Find out how they reach your end goal. what kind of columns are there, their count, Like in this one maximum problem is found in Anterior Communicating Artery (375). It could mean that is what happens in real life but could also mean the data collected in this dataset has more of these kinda patient. Make graphs, doesnt matter pie, bar, any. ","metadata":{}},{"cell_type":"code","source":"import pandas as pd \nimport numpy as np","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:56:27.141165Z","iopub.execute_input":"2025-08-07T18:56:27.141623Z","iopub.status.idle":"2025-08-07T18:56:29.722725Z","shell.execute_reply.started":"2025-08-07T18:56:27.141589Z","shell.execute_reply":"2025-08-07T18:56:29.721845Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data = pd.read_csv('../input/rsna-intracranial-aneurysm-detection/train.csv')\ntrain_data.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:56:29.723663Z","iopub.execute_input":"2025-08-07T18:56:29.724363Z","iopub.status.idle":"2025-08-07T18:56:29.793261Z","shell.execute_reply.started":"2025-08-07T18:56:29.72434Z","shell.execute_reply":"2025-08-07T18:56:29.792309Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:56:29.795126Z","iopub.execute_input":"2025-08-07T18:56:29.795484Z","iopub.status.idle":"2025-08-07T18:56:29.801845Z","shell.execute_reply.started":"2025-08-07T18:56:29.795453Z","shell.execute_reply":"2025-08-07T18:56:29.800969Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"localizers = pd.read_csv('../input/rsna-intracranial-aneurysm-detection/train_localizers.csv')\nlocalizers.head(2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:56:48.955483Z","iopub.execute_input":"2025-08-07T18:56:48.955784Z","iopub.status.idle":"2025-08-07T18:56:48.990539Z","shell.execute_reply.started":"2025-08-07T18:56:48.955755Z","shell.execute_reply":"2025-08-07T18:56:48.989535Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"So there are two types of data and SeriesInstanceUID is the connection between those, now you have to solve two task of classification and pin pointing the locations. Combine both to create a solid model. Yayy, you got this. Easy peasy.\n\n**Step 1**: “Yes/No” for aneurysm  \n**Step 2**: “If yes, show me where”\n","metadata":{}},{"cell_type":"code","source":"localizers_count = localizers['location'].value_counts()\nlocalizers_count","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:56:55.991653Z","iopub.execute_input":"2025-08-07T18:56:55.992066Z","iopub.status.idle":"2025-08-07T18:56:56.005984Z","shell.execute_reply.started":"2025-08-07T18:56:55.992038Z","shell.execute_reply":"2025-08-07T18:56:56.005005Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Basic EDA to understand data","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt \n\nplt.figure(figsize = (8,8))\nplt.pie(localizers_count, labels=localizers_count.index, autopct='%1.1f%%')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:56:58.515692Z","iopub.execute_input":"2025-08-07T18:56:58.516096Z","iopub.status.idle":"2025-08-07T18:56:58.930575Z","shell.execute_reply.started":"2025-08-07T18:56:58.516067Z","shell.execute_reply":"2025-08-07T18:56:58.929515Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scan_types = train_data['Modality'].unique()\nprint(\"\\nDifferent Scan types: \", scan_types)\nmodality_count = train_data['Modality'].value_counts()\nprint(\"\\n\", modality_count)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:57:02.034438Z","iopub.execute_input":"2025-08-07T18:57:02.034738Z","iopub.status.idle":"2025-08-07T18:57:02.05179Z","shell.execute_reply.started":"2025-08-07T18:57:02.034716Z","shell.execute_reply":"2025-08-07T18:57:02.050488Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize =(6,4))\nmodality_count.plot(kind='bar', color='pink')\nplt.title(\"Modality Distribution\")\nplt.xlabel(\"Modality\")\nplt.ylabel(\"Number of Scans\")\nplt.xticks(rotation=0)\nplt.grid(axis='y', linestyle=' ', alpha=0.7)\nplt.tight_layout()\n\nfor i, v in enumerate(modality_count):\n    plt.text(i, v , str(v), ha='center', va='bottom', fontsize=9, rotation=0)\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:57:04.318874Z","iopub.execute_input":"2025-08-07T18:57:04.319243Z","iopub.status.idle":"2025-08-07T18:57:04.57975Z","shell.execute_reply.started":"2025-08-07T18:57:04.319218Z","shell.execute_reply":"2025-08-07T18:57:04.578645Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"These are just the metadata of the images, now lets try loading the actual image. There are two kinds of files. DICOMs -> raw 2d image slice need to be stacked to form 3D volume and NIFTI(.nii) -> processed and combined into 3D volumes. ","metadata":{}},{"cell_type":"markdown","source":"NOW\n1. Understang the problem\n2. Preprocessing the data\n3. Build the model\n4. Train + validate\n5. Predict and submit","metadata":{}},{"cell_type":"code","source":"import os\nimport pydicom\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom glob import glob\nimport random\nfrom scipy.ndimage import zoom\n\n\ndef resize_volume(img, target_shape=(64, 128, 128)):\n    current_shape = img.shape\n    zoom_factors = [t / c for t, c in zip(target_shape, current_shape)]\n    return zoom(img, zoom_factors, order=1)  # order=1 = linear interpolation\n    \n    \ndef load_dicom_series(series_path, target_shape=(128, 128)):\n    files = glob(os.path.join(series_path, \"*.dcm\"))\n    print(\"Found DICOM files:\", len(files))  \n\n    files = [pydicom.dcmread(f) for f in files]\n    files = [f for f in files if hasattr(f, \"ImagePositionPatient\")]\n    print(\"Files with ImagePositionPatient:\", len(files)) \n\n    files = sorted(files, key=lambda s: float(s.ImagePositionPatient[2]))\n\n    resized_slices = []\n    for s in files:\n        img = s.pixel_array\n        resized_img = resize_volume(img, target_shape)\n        resized_slices.append(resized_img)\n\n    if not resized_slices:\n        raise ValueError(f\"No valid slices in {series_path}\")  \n\n    volume = np.stack(resized_slices)\n    return volume\n\n\ndef normalize(volume):\n    volume = volume.astype(np.float32)\n    volume = (volume - np.min(volume)) / (np.max(volume) - np.min(volume))\n    return volume\n\ndef show_middle_slices(volume, n=9):\n    depth = volume.shape[0]\n    idxs = np.linspace(0.25*depth, 0.75*depth, n).astype(int)\n    plt.figure(figsize=(15, 5))\n    for i, idx in enumerate(idxs):\n        plt.subplot(1, n, i+1)\n        plt.imshow(volume[idx], cmap='gray')\n        plt.title(f\"Slice {idx}\")\n        plt.axis('off')\n    plt.tight_layout()\n    plt.show()\n\n# Picking a random DICOM series to visualise the brain \nseries_root = '../input/rsna-intracranial-aneurysm-detection/series'\nall_series = [d for d in os.listdir(series_root) if os.path.isdir(os.path.join(series_root, d))]\nrandom_series = random.choice(all_series)\nrandom_series_path = os.path.join(series_root, random_series)\n\nprint(\"Visualizing:\", random_series)\n\nvolume = load_dicom_series(random_series_path)\nvolume = normalize(volume)\nshow_middle_slices(volume)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:57:11.207456Z","iopub.execute_input":"2025-08-07T18:57:11.207822Z","iopub.status.idle":"2025-08-07T18:57:24.392693Z","shell.execute_reply.started":"2025-08-07T18:57:11.207798Z","shell.execute_reply":"2025-08-07T18:57:24.391501Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import ipywidgets as widgets\nfrom IPython.display import display\n\ndef interactive_view(volume):\n    def view_slice(slice_idx):\n        plt.imshow(volume[slice_idx], cmap='gray')\n        plt.title(f\"Slice {slice_idx}\")\n        plt.axis('off')\n        plt.show()\n\n    slider = widgets.IntSlider(min=0, max=volume.shape[0]-1, step=1, value=volume.shape[0]//2)\n    widgets.interact(view_slice, slice_idx=slider)\n\ninteractive_view(volume)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:57:24.394478Z","iopub.execute_input":"2025-08-07T18:57:24.395245Z","iopub.status.idle":"2025-08-07T18:57:24.695044Z","shell.execute_reply.started":"2025-08-07T18:57:24.395218Z","shell.execute_reply":"2025-08-07T18:57:24.694023Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Full Mesh 3D Brain using Marching Cubes -> I dont think that turned out good. I was trying to make a 3D brain just to see. ","metadata":{}},{"cell_type":"code","source":"from skimage import measure\nimport plotly.graph_objects as go\n\nv = volume  \n\n# Pick a good isosurface level after normalization (0–1)\niso_level = 0.1  \n\n# Extract mesh surface\nverts, faces, _, _ = measure.marching_cubes(v, level=iso_level)\n\nx, y, z = verts.T\ni, j, k = faces.T\n\nfig = go.Figure(data=[go.Mesh3d(\n    x=x, y=y, z=z,\n    i=i, j=j, k=k,\n    color='lightgray',\n    opacity=1.0,\n    lighting=dict(ambient=0.3, diffuse=1),\n    lightposition=dict(x=100, y=200, z=0)\n)])\n\nfig.update_layout(\n    scene=dict(aspectmode='data'),\n    title=\"Full-Resolution Brain Surface\"\n)\nfig.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:57:58.175028Z","iopub.execute_input":"2025-08-07T18:57:58.175351Z","iopub.status.idle":"2025-08-07T18:57:58.441418Z","shell.execute_reply.started":"2025-08-07T18:57:58.175325Z","shell.execute_reply":"2025-08-07T18:57:58.440437Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(f\"Total slices in this scan: {volume.shape[0]}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:57:25.614548Z","iopub.execute_input":"2025-08-07T18:57:25.614905Z","iopub.status.idle":"2025-08-07T18:57:25.620772Z","shell.execute_reply.started":"2025-08-07T18:57:25.614878Z","shell.execute_reply":"2025-08-07T18:57:25.619671Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"> `volume.shape[0] = 45`\n\nThis means a single brain has 45 image slices, and we need to combine all of them into one NumPy array to train the model.  \n\nHowever, we have a problem: all images are not the same size.  \n\n---\n\n> **Problem:** Different image sizes. For example, one slice might be `128 × 128` pixels, another could be `68 × 68`.\n> **Solution:** Resize all slices to a fixed size  \n> **How to choose a good fixed size?**   → Take a random sample and average the dimensions or trying finding a middle ground.\n","metadata":{}},{"cell_type":"code","source":"shapes = []\n\nfor folder in random.sample(os.listdir(series_root), 5):  # More samples = better view\n    path = os.path.join(series_root, folder)\n    try:\n        slices = [pydicom.dcmread(os.path.join(path, f)).pixel_array for f in sorted(os.listdir(path))]\n        vol = np.stack(slices)\n        shapes.append(vol.shape)  # (depth, height, width)\n    except:\n        continue\n\n# Separate dims\ndepths = [s[0] for s in shapes]\nheights = [s[1] for s in shapes]\nwidths = [s[2] for s in shapes]\n\n# Plot distributions\nplt.figure(figsize=(12, 4))\n\nplt.subplot(1, 3, 1)\nplt.hist(depths, bins=10)\nplt.title(\"Depth Distribution\")\n\nplt.subplot(1, 3, 2)\nplt.hist(heights, bins=10)\nplt.title(\"Height Distribution\")\n\nplt.subplot(1, 3, 3)\nplt.hist(widths, bins=10)\nplt.title(\"Width Distribution\")\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:58:17.101279Z","iopub.execute_input":"2025-08-07T18:58:17.101591Z","iopub.status.idle":"2025-08-07T18:59:11.879972Z","shell.execute_reply.started":"2025-08-07T18:58:17.101568Z","shell.execute_reply":"2025-08-07T18:59:11.878994Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Extract the columns needed for labels; i.e all the positions that might consist of aneurysm. ","metadata":{}},{"cell_type":"code","source":"binary_cols = [col for col in train_data.columns if set(train_data[col].unique()) <= {0, 1, np.nan}]\nbinary_cols.remove('Aneurysm Present')\nprint(\"Coutn of binary col\", len(binary_cols))\nbinary_cols","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:59:18.107394Z","iopub.execute_input":"2025-08-07T18:59:18.107749Z","iopub.status.idle":"2025-08-07T18:59:18.123777Z","shell.execute_reply.started":"2025-08-07T18:59:18.107707Z","shell.execute_reply":"2025-08-07T18:59:18.122756Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Okay, so now we have to combine the labels with the data to get it ready for training.  \nI was confused at first — why aren’t the labels already in the same file with tagged annotations?  \n\nTurns out, most scans in `train.csv` don’t have aneurysms.  \nFor confirmed aneurysms, they have rows in `train_localizations.csv`.  \n\nKeeping them separate makes `train.csv` smaller and faster to load for binary or multilabel classification.\n\nOur task is to combine the two datasets based on the ID.  \nThink for a second — how would you do that logically?\n","metadata":{}},{"cell_type":"code","source":"artery_cols = [col for col in train_data.columns if col != 'SeriesInstanceUID']\nlabel_dict = train_data.set_index('SeriesInstanceUID')[artery_cols].to_dict(orient='index')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:59:21.577583Z","iopub.execute_input":"2025-08-07T18:59:21.577891Z","iopub.status.idle":"2025-08-07T18:59:21.626043Z","shell.execute_reply.started":"2025-08-07T18:59:21.577868Z","shell.execute_reply":"2025-08-07T18:59:21.625235Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now that we have both X and Y ready, our data is ready to be a input for a ML model. ","metadata":{}},{"cell_type":"markdown","source":"### ⚡ Speeding Up Development with Saved Samples\n\nPreprocessing the full dataset takes a long time and can be frustrating during experimentation.  \nTo stay sane, I took a random sample of 100 volumes, preprocessed them, and saved them as `.npy` files.\n\nThis way, I can quickly load the processed data next time without repeating the entire pipeline.\n","metadata":{}},{"cell_type":"code","source":"from tqdm import tqdm\n\nX_data = []\ny_data = []\n\nvalid_folders = [f for f in os.listdir(series_root) if f in label_dict]\n\n# Taking a small random sample for now to move ahead quickly.\nvalid_folders = random.sample(valid_folders, 50)\n\n\nprocessed = 0  # 👈 manual counter\n\nfor folder in tqdm(valid_folders):\n    try:\n        vol = load_dicom_series(os.path.join(series_root, folder), target_shape=(128, 128))\n        vol = normalize(vol)\n        vol = resize_volume(vol, target_shape=(64, 128, 128))\n        vol = vol.astype(np.float16)\n\n        label = label_dict[folder]\n        X_data.append(vol)\n        y_data.append(label)\n\n        processed += 1\n        if processed % 10 == 0:\n            tqdm.write(f\"✅ Processed: {processed} folders\")\n    except Exception as e:\n        tqdm.write(f\"⚠️ Skipped {folder} due to error: {e}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T18:59:35.661579Z","iopub.execute_input":"2025-08-07T18:59:35.661883Z","iopub.status.idle":"2025-08-07T19:04:02.803096Z","shell.execute_reply.started":"2025-08-07T18:59:35.661861Z","shell.execute_reply":"2025-08-07T19:04:02.802125Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Save it and load later because it pre-processing again and again takes a lot of time. \n\n# np.save('/kaggle/working/X_data_100.npy', X_data)\n# np.save('/kaggle/working/y_data_100.npy', y_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-05T17:01:59.626987Z","iopub.execute_input":"2025-08-05T17:01:59.627461Z","iopub.status.idle":"2025-08-05T17:01:59.930826Z","shell.execute_reply.started":"2025-08-05T17:01:59.627429Z","shell.execute_reply":"2025-08-05T17:01:59.928907Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now changing them into Np array suitable for model. Yes, using np.stack over np.array itched my brain for a while Turns out, np.stack gives you a proper tensor that can be passed to the model, while np.array might return a garbage object array if the shapes aren't exactly the same.","metadata":{}},{"cell_type":"code","source":"X_data = np.stack(X_data)[..., np.newaxis]  # shape: (N, 64, 128, 128, 1)\ny_data = [label['Aneurysm Present'] for label in y_data] ## TODO. do this while loading the DCOM \ny_data = np.array(y_data) ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T19:04:02.804928Z","iopub.execute_input":"2025-08-07T19:04:02.80524Z","iopub.status.idle":"2025-08-07T19:04:02.832702Z","shell.execute_reply.started":"2025-08-07T19:04:02.805219Z","shell.execute_reply":"2025-08-07T19:04:02.831908Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX_train, X_val, y_train, y_val = train_test_split(\n    X_data, y_data, \n    test_size=0.2, \n    stratify=y_data,  # Keeps class balance\n    random_state=42\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T19:30:15.855554Z","iopub.execute_input":"2025-08-07T19:30:15.85968Z","iopub.status.idle":"2025-08-07T19:30:17.230134Z","shell.execute_reply.started":"2025-08-07T19:30:15.8596Z","shell.execute_reply":"2025-08-07T19:30:17.228778Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(X_data.shape)\nprint(X_train.shape)\nprint(X_train[0].shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T19:30:19.445186Z","iopub.execute_input":"2025-08-07T19:30:19.447042Z","iopub.status.idle":"2025-08-07T19:30:19.452911Z","shell.execute_reply.started":"2025-08-07T19:30:19.447007Z","shell.execute_reply":"2025-08-07T19:30:19.451869Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from tensorflow.keras import layers, models\n\nmodel = models.Sequential([\n    layers.Input(shape=(64, 128, 128, 1)),\n    layers.Conv3D(16, 3, activation='relu', padding='same'),\n    layers.MaxPool3D(2),\n    layers.Conv3D(32, 3, activation='relu', padding='same'),\n    layers.MaxPool3D(2),\n    layers.Conv3D(64, 3, activation='relu', padding='same'),\n    layers.GlobalAveragePooling3D(),\n    layers.Dense(64, activation='relu'),\n    layers.Dense(1, activation='sigmoid')\n])\n\nmodel.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T19:30:20.636702Z","iopub.execute_input":"2025-08-07T19:30:20.637767Z","iopub.status.idle":"2025-08-07T19:30:39.671261Z","shell.execute_reply.started":"2025-08-07T19:30:20.637731Z","shell.execute_reply":"2025-08-07T19:30:39.670148Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"history = model.fit(X_train, y_train, validation_data=(X_val, y_val), epochs=10)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-07T19:30:39.672608Z","iopub.execute_input":"2025-08-07T19:30:39.673622Z","execution_failed":"2025-08-07T21:08:09.354Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.evaluate(X_val, y_val)","metadata":{"trusted":true,"execution":{"execution_failed":"2025-08-07T21:08:09.356Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\nplt.plot(history.history['accuracy'], label='train acc')\nplt.plot(history.history['val_accuracy'], label='val acc')\nplt.legend()\nplt.show()","metadata":{"trusted":true,"execution":{"execution_failed":"2025-08-07T21:08:09.356Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.save('aneurysm_model.h5')","metadata":{"trusted":true,"execution":{"execution_failed":"2025-08-07T21:08:09.356Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import accuracy_score, classification_report\n\n# Convert predicted probabilities to binary labels (0 or 1)\npred_labels = (preds > 0.5).astype(int).flatten()\nprint(\"Accuracy:\", accuracy_score(y_val, pred_labels))\n\n# Full classification report\nprint(classification_report(y_val, pred_labels))","metadata":{"trusted":true,"execution":{"execution_failed":"2025-08-07T21:08:09.356Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Report was something like: \n\nAccuracy: 0.5789473684210527\n              precision    recall  f1-score   support\n\n           0       0.58      1.00      0.73        11\n           1       0.00      0.00      0.00         8\n\n    accuracy                           0.58        19\n   macro avg       0.29      0.50      0.37        19\nweighted avg       0.34      0.58      0.42        19\n","metadata":{}},{"cell_type":"markdown","source":"📉 The model achieves around 58% accuracy, but it's heavily biased toward predicting only class 0 (no aneurysm).  \n","metadata":{}},{"cell_type":"markdown","source":"## TAKE AWAY\n\nYou’ve built your first model, it was confusing but you did it. \nNow it’s time to iterate: try different architectures, tune hyperparameters, and start submitting to Kaggle.  \nExperiment boldly! you’ve got this. **Good luck!**\n","metadata":{}},{"cell_type":"code","source":"# --- Submission maker (drop this after training) ---\nimport os, numpy as np, pandas as pd\nfrom tqdm import tqdm\n\n# Paths\nDATA_ROOT = \"../input/rsna-intracranial-aneurysm-detection\"\nTEST_ROOT = os.path.join(DATA_ROOT, \"test\")\nSUB_PATH  = os.path.join(DATA_ROOT, \"sample_submission.csv\")\n\n# 1) Read sample to get the exact required columns\nsub = pd.read_csv(SUB_PATH)\nlabel_cols = [c for c in sub.columns if c != \"ID\"]  # 13 location columns (or 1 \"Label\" if that's the format)\n\n# 2) Collect test IDs (SeriesInstanceUID folders)\ntest_ids = [d for d in os.listdir(TEST_ROOT) if os.path.isdir(os.path.join(TEST_ROOT, d))]\ntest_ids = sorted(test_ids)\n\ndef predict_series(series_id: str) -> float:\n    \"\"\"Load → preprocess → predict single series → return one probability.\"\"\"\n    series_path = os.path.join(TEST_ROOT, series_id)\n    vol = load_dicom_series(series_path, target_shape=(128, 128))  # uses your existing helpers\n    vol = normalize(vol)\n    vol = resize_volume(vol, target_shape=(64, 128, 128)).astype(np.float32)\n    vol = np.expand_dims(vol, axis=-1)  # (D,H,W,1)\n    vol = np.expand_dims(vol, axis=0)   # (1,D,H,W,1) batch\n    prob = float(model.predict(vol, verbose=0)[0, 0])  # sigmoid output\n    return prob\n\n# 3) Run inference and fill submission\nrows = []\nfor sid in tqdm(test_ids, desc=\"Predicting\"):\n    p = predict_series(sid)\n    row = {\"ID\": sid}\n    # If competition needs 13 columns, we broadcast the same scalar for now.\n    # If it's a single 'Label', this still works because label_cols == ['Label'].\n    for c in label_cols:\n        row[c] = p\n    rows.append(row)\n\nsubmission = pd.DataFrame(rows, columns=[\"ID\"] + label_cols)\nsubmission.to_csv(\"submission.csv\", index=False)\nprint( Saved submission.csv with\", len(submission), \"rows and\", len(label_cols), \"label column(s).\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-21T17:20:12.511957Z","iopub.execute_input":"2025-08-21T17:20:12.512211Z","iopub.status.idle":"2025-08-21T17:20:15.534061Z","shell.execute_reply.started":"2025-08-21T17:20:12.512189Z","shell.execute_reply":"2025-08-21T17:20:15.532652Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}