{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":99552,"databundleVersionId":13851420,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"- Title “RSNA Project Notes”\n- A bullet list with “goal”, “data”, “next steps”\n- Make the word \"goal\" bold\n- Add the competition link, name it \"RSNA Intracranial Aneurysm Detection\" https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection","metadata":{},"attachments":{}},{"cell_type":"markdown","source":"[RSNA Intracranial Aneurysm Detection](https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection)","metadata":{}},{"cell_type":"markdown","source":"# Heading\n## Subheading\n### subsection\n\n**bold** v.s. plain text  \n*italic* v.s. plain text\n\n**LIST:**\n* first item\n* second item\n* third item\n\nfor more imformation about this competition, click [here](https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/data)","metadata":{}},{"cell_type":"markdown","source":"# What you can edit (simple knobs)\n\nThis block defines a few settings you can change to explore different scans or show more images.","metadata":{}},{"cell_type":"code","source":"# --- SETTINGS STUDENTS CAN TUNE ---\n\n# If True, try to pick a series with an aneurysm. If False, pick any series.\nPICK_POSITIVE = True\n\n# Optionally force a specific SeriesInstanceUID (otherwise we'll sample one).\nSERIES_TO_VIEW = None  # e.g., \"1.2.840.113619....\"  <<-- put a UID string here to lock a case\n\n# How many slices to show in the grid, and how many columns in that grid.\nNUM_SLICES = 18\nGRID_COLS  = 6\n\n# Random seed so results are repeatable when sampling.\nRANDOM_SEED = 42","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2026-01-24T05:40:54.553628Z","iopub.execute_input":"2026-01-24T05:40:54.554278Z","iopub.status.idle":"2026-01-24T05:40:54.559161Z","shell.execute_reply.started":"2026-01-24T05:40:54.55424Z","shell.execute_reply":"2026-01-24T05:40:54.558109Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Chunk 1 — Imports & basic setup\n\nWe load the libraries we’ll use, including tools for reading DICOM (medical images) and NIfTI (segmentations).","metadata":{}},{"cell_type":"code","source":"import os, glob, ast\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport pydicom\n\n# NIfTI is optional; only needed if you want to overlay vessel segmentations.\ntry:\n    import nibabel as nib\n    HAVE_NIB = True\nexcept Exception:\n    HAVE_NIB = False\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-24T05:40:54.560791Z","iopub.execute_input":"2026-01-24T05:40:54.561375Z","iopub.status.idle":"2026-01-24T05:40:54.577791Z","shell.execute_reply.started":"2026-01-24T05:40:54.561347Z","shell.execute_reply":"2026-01-24T05:40:54.576395Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Chunk 2 — Read the label CSV files\n\nWe open the main labels file and the “localizer” file that gives exact XY points for some aneurysms.","metadata":{}},{"cell_type":"code","source":"# Change this to your competition input path on Kaggle\n# /kaggle/input/rsna-intracranial-aneurysm-detection\n\nBASE_DIR = \"/kaggle/input/rsna-intracranial-aneurysm-detection\"  # <<-- EDIT if your path name differs\n\n# train.csv = one row per series (scan), with 13 location labels + the main \"Aneurysm Present\".\ndf  = pd.read_csv(os.path.join(BASE_DIR, \"train.csv\"))\n\n# train_localizers.csv = points marking aneurysm centers (x, y) on specific slices.\nloc = pd.read_csv(os.path.join(BASE_DIR, \"train_localizers.csv\"))\n\nSERIES_DIR = os.path.join(BASE_DIR, \"series\")\nSEG_DIR    = os.path.join(BASE_DIR, \"segmentations\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-24T05:40:54.57911Z","iopub.execute_input":"2026-01-24T05:40:54.579398Z","iopub.status.idle":"2026-01-24T05:40:54.626275Z","shell.execute_reply.started":"2026-01-24T05:40:54.579369Z","shell.execute_reply":"2026-01-24T05:40:54.625202Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.head(8)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-24T05:40:54.628394Z","iopub.execute_input":"2026-01-24T05:40:54.628822Z","iopub.status.idle":"2026-01-24T05:40:54.644843Z","shell.execute_reply.started":"2026-01-24T05:40:54.628796Z","shell.execute_reply":"2026-01-24T05:40:54.643863Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Chunk 3 — Pick one series to visualize\n\nWe choose a scan. You can hard-set SERIES_TO_VIEW to a known ID, or let the code pick one (preferably positive cases if available).","metadata":{}},{"cell_type":"code","source":"if SERIES_TO_VIEW:\n    row = df[df[\"SeriesInstanceUID\"] == SERIES_TO_VIEW].iloc[0]\nelse:\n    if PICK_POSITIVE and \"Aneurysm Present\" in df.columns and df[\"Aneurysm Present\"].sum() > 0:\n        cand = df[df[\"Aneurysm Present\"] == 1]\n    else:\n        cand = df\n    row = cand.sample(1, random_state=0).iloc[0]\n\nsid = row[\"SeriesInstanceUID\"]\nmodality = row.get(\"Modality\", \"Unknown\")\nseries_path = os.path.join(SERIES_DIR, sid)\nassert os.path.isdir(series_path), f\"Series folder not found: {series_path}\"\n# print(f\"Series Path: {series_path}\")\nprint(f\"Chosen SeriesInstanceUID: {sid}\")\n# print(f\"Modality: {modality}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-24T05:40:54.645855Z","iopub.execute_input":"2026-01-24T05:40:54.646144Z","iopub.status.idle":"2026-01-24T05:40:54.659044Z","shell.execute_reply.started":"2026-01-24T05:40:54.646121Z","shell.execute_reply":"2026-01-24T05:40:54.657935Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from IPython.display import display\ndisplay(row.to_frame().T)\n\n# print(row.to_frame().T.to_string(index=False))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-24T05:40:54.660062Z","iopub.execute_input":"2026-01-24T05:40:54.660425Z","iopub.status.idle":"2026-01-24T05:40:54.686266Z","shell.execute_reply.started":"2026-01-24T05:40:54.660393Z","shell.execute_reply":"2026-01-24T05:40:54.68534Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# files = sorted(glob.glob(os.path.join(series_path, \"*.dcm\")))  \n# print(f\"Total slices: {len(files)}\")\n# names = [os.path.basename(p) for p in files]\n# print(names[:8] + [\"...\"] + names[-4:]) # show the first 8 images and the last 4 images","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-24T05:40:54.687267Z","iopub.execute_input":"2026-01-24T05:40:54.687557Z","iopub.status.idle":"2026-01-24T05:40:54.700967Z","shell.execute_reply.started":"2026-01-24T05:40:54.68753Z","shell.execute_reply":"2026-01-24T05:40:54.699987Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Orientation** = which side is up/left. It keeps all slices facing the same way so nothing is flipped or rotated weirdly.\n\n**Position** = which “layer” this slice sits on along the stack (like how high or low it is).\n\n**Sorting idea**: first use Orientation to find the direction the stack grows, then use Position to compare slices—who’s higher, who’s lower—and sort by that.\n\nIf those geometry tags are **missing**, fall back to **InstanceNumber** (the scanner’s “page number”). It often works, but it isn’t guaranteed to be perfect.\n\nLuckily, in this case we have both orientation and position information.","metadata":{}},{"cell_type":"code","source":"# dcm_path = files[len(files)//2] \n# ds = pydicom.dcmread(dcm_path, stop_before_pixels=True)\n# print(\"Modality:\", getattr(ds,\"Modality\",None))\n# print(\"PatientAge:\", getattr(ds,\"PatientAge\",None), \"PatientSex:\", getattr(ds,\"PatientSex\",None))\n# print(\"Rows x Cols:\", getattr(ds,\"Rows\",None),\"x\",getattr(ds,\"Columns\",None))\n# print(\"PixelSpacing (mm):\", getattr(ds,\"PixelSpacing\",None))        \n# print(\"SliceThickness (mm):\", getattr(ds,\"SliceThickness\",None))\n# print(\"ImagePositionPatient (mm):\", getattr(ds,\"ImagePositionPatient\",None))  \n# print(\"ImageOrientationPatient:\", getattr(ds,\"ImageOrientationPatient\",None)) \n# print(\"InstanceNumber:\", getattr(ds,\"InstanceNumber\",None))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-24T05:40:54.702047Z","iopub.execute_input":"2026-01-24T05:40:54.702337Z","iopub.status.idle":"2026-01-24T05:40:54.718093Z","shell.execute_reply.started":"2026-01-24T05:40:54.702314Z","shell.execute_reply":"2026-01-24T05:40:54.717034Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Chunk 4 — Load the DICOM slices and stack them into a 3D volume\n\nA scan is a “stack of slices”. We load all slices, sort them in the right order, and stack into a 3D array [Z, Y, X].","metadata":{}},{"cell_type":"markdown","source":"What does `pydicom.dcmread(fp)` do?\nIt opens a DICOM file (the standard file type for medical images) and parses it into a Python object you can work with.","metadata":{}},{"cell_type":"code","source":"def sort_key(fp):\n    ds = pydicom.dcmread(fp, stop_before_pixels=True)\n    # Use physical position (z) if available; else fall back to instance number.\n    if \"ImagePositionPatient\" in ds:\n        return float(ds.ImagePositionPatient[2])\n    if \"InstanceNumber\" in ds:\n        return int(ds.InstanceNumber)\n    return 0\n\ndcm_files = sorted(glob.glob(os.path.join(series_path, \"*.dcm\")), key=sort_key)\nassert len(dcm_files) > 0, \"No DICOM files found in this series.\"\n\nslices = []\nfor fp in dcm_files:\n    ds = pydicom.dcmread(fp)\n    arr = ds.pixel_array.astype(np.float32)\n\n    # Most CT/CTA store a linear transform (slope/intercept) to convert to Hounsfield Units (HU).\n    slope     = float(getattr(ds, \"RescaleSlope\", 1.0))\n    intercept = float(getattr(ds, \"RescaleIntercept\", 0.0))\n    arr = arr * slope + intercept\n\n    slices.append(arr)\n\nvol = np.stack(slices, axis=0)  # shape: [Z, Y, X]\nprint(\"Volume shape [Z, Y, X]:\", vol.shape)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-24T05:40:54.71917Z","iopub.execute_input":"2026-01-24T05:40:54.719559Z","iopub.status.idle":"2026-01-24T05:40:55.498222Z","shell.execute_reply.started":"2026-01-24T05:40:54.719521Z","shell.execute_reply":"2026-01-24T05:40:55.497156Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- **(CT/CTA) Convert to HU**  \n   Many CT images store values that must be unpacked:  \n   `image = image * RescaleSlope + RescaleIntercept` → Hounsfield Units (HU).  \n   (For MRI/MRA this step usually leaves the image unchanged.)\n\n- **Collect slices in order**  \n   Add each 2D slice to a list. The **order matters** — sort files by geometry (best: `ImageOrientationPatient` + `ImagePositionPatient`) or fall back to `InstanceNumber`.\n\n- **Stack into 3D**  \n   `vol = np.stack(slices, axis=0)` puts slices along the first axis.  \n   `vol.shape` is `[Z, Y, X]` → **Z = number of slices**, **Y = rows**, **X = columns**.\n\n- **Quick sanity checks**  \n   - Show first / middle / last slice to see if the stack looks continuous.  \n   - Print the shape and (for CT) min/max HU to confirm values look reasonable.\n","metadata":{}},{"cell_type":"markdown","source":"# Chunk 5 — Make the volume viewable (simple brightness/contrast)\n\nRaw medical data is too dark/bright. We map values to 0–255 for display. For CT we use a typical “brain window”; for MRI we use percentile stretch.","metadata":{}},{"cell_type":"code","source":"def to_uint8(img3d, center=None, width=None):\n    img = img3d.copy()\n    if (center is not None) and (width is not None):\n        vmin = center - width/2.0\n        vmax = center + width/2.0\n    else:\n        vmin, vmax = np.percentile(img, (1, 99))  # generic for MRI/MRA\n    img = np.clip(img, vmin, vmax)\n    img = (img - vmin) / (vmax - vmin + 1e-6)\n    return (img * 255).astype(np.uint8)\n\nuse_ct_window = (str(modality).upper() in [\"CT\", \"CTA\"])\nvol_u8 = to_uint8(vol, center=40, width=300) if use_ct_window else to_uint8(vol)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-24T05:40:55.500781Z","iopub.execute_input":"2026-01-24T05:40:55.501143Z","iopub.status.idle":"2026-01-24T05:40:55.607437Z","shell.execute_reply.started":"2026-01-24T05:40:55.501121Z","shell.execute_reply":"2026-01-24T05:40:55.606429Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Shape of 'vol_u8' after windowing:\", vol_u8.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-24T05:40:55.608604Z","iopub.execute_input":"2026-01-24T05:40:55.608892Z","iopub.status.idle":"2026-01-24T05:40:55.614658Z","shell.execute_reply.started":"2026-01-24T05:40:55.608869Z","shell.execute_reply":"2026-01-24T05:40:55.613521Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Chunk 6 — A simple “slice gallery” (grid of images)\n\nWe show a grid of evenly spaced slices to build intuition that “3D = many 2D slices”.","metadata":{}},{"cell_type":"markdown","source":"NUM_SLICES = 12\n\n\nGRID_COLS  = 6\n\nRANDOM_SEED = 42","metadata":{}},{"cell_type":"code","source":"def show_montage(vol_u8, nslices=NUM_SLICES, ncols=GRID_COLS, title=\"\"):\n    idx = np.linspace(0, vol_u8.shape[0]-1, nslices).astype(int)\n    nrows = int(np.ceil(nslices / ncols))\n    fig, axes = plt.subplots(nrows, ncols, figsize=(2.5*ncols, 2.5*nrows))\n    axes = axes.ravel()\n    for i, ax in enumerate(axes):\n        ax.axis(\"off\")\n        if i < len(idx):\n            ax.imshow(vol_u8[idx[i]], cmap=\"gray\")\n            ax.set_title(f\"z={idx[i]}\")\n    fig.suptitle(title)\n    plt.show()\n\n# show_montage(vol_u8, nslices=NUM_SLICES, ncols=GRID_COLS,\n#              title=f\"Series {sid} | {modality} | shape {vol_u8.shape}\")\n\nshow_montage(vol_u8, title=f\"Series {sid} | {modality} | shape {vol_u8.shape}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-24T05:40:55.615954Z","iopub.execute_input":"2026-01-24T05:40:55.616216Z","iopub.status.idle":"2026-01-24T05:40:56.85209Z","shell.execute_reply.started":"2026-01-24T05:40:55.616194Z","shell.execute_reply":"2026-01-24T05:40:56.850875Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Homework\n\nTry changing the parameters in the first code cell:\n\n- NUM_SLICES = 15\n- GRID_COLS = 6\n- RANDOM_SEED = 0\n\nThink about this:\n- What do NUM_SLICES and GRID_COLS control? When you adjust these values, which code cell do you need to rerun to see the effect?\n\n- What is RANDOM_SEED? When you change it, what differences do you notice? (If you are not sure about this, ask AI for help)\n\nWrite down your observations in a new markdown cell below.","metadata":{}}]}