{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":99552,"databundleVersionId":13762876,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This is a project for the course \"Deep Learning for Visual Recognition\" in Aarhus university\n\nWe will start by looking at the data we have","metadata":{}},{"cell_type":"markdown","source":"# Data exploration","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport nibabel as nib\nimport numpy as np\nimport os\nimport pandas as pd \nimport pydicom\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-09-26T12:01:38.790925Z","iopub.execute_input":"2025-09-26T12:01:38.791235Z","iopub.status.idle":"2025-09-26T12:01:41.966896Z","shell.execute_reply.started":"2025-09-26T12:01:38.791203Z","shell.execute_reply":"2025-09-26T12:01:41.965597Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"folder_path = '/kaggle/input/rsna-intracranial-aneurysm-detection'\nos.listdir(folder_path)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-26T12:02:17.494376Z","iopub.execute_input":"2025-09-26T12:02:17.494786Z","iopub.status.idle":"2025-09-26T12:02:17.50262Z","shell.execute_reply.started":"2025-09-26T12:02:17.494762Z","shell.execute_reply":"2025-09-26T12:02:17.501397Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We have 2 csv files with 3 more folders, first, we will explore the csv files:","metadata":{}},{"cell_type":"markdown","source":"## train.csv","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv(folder_path + '/train.csv')\ntrain","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-25T08:26:11.076123Z","iopub.execute_input":"2025-09-25T08:26:11.076427Z","iopub.status.idle":"2025-09-25T08:26:11.143231Z","shell.execute_reply.started":"2025-09-25T08:26:11.076405Z","shell.execute_reply":"2025-09-25T08:26:11.14235Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We have 4348 observations in our training set, that will be divided into training and validation. Each sample contains 18 columns which are:\n\n- SeriesInstanceUID: An unique identifier for each instance, which is associated with a folder into the series folder, that allows to extract the medical images for each patient\n- PatientAge\n- PatientSex\n- Modality: The modality of the medical image, this field can be CTA,MRA, MIR T1post or MRI 2\n\nFollowed by 13 binary columns of the possible locations of the aneurysm, which are:\n- Left Infraclinoid Internal Carotid Artery\n- Right Infraclinoid Internal Carotid Artery\n- Left Supraclinoid Internal Carotid Artery\n- Right Supraclinoid Internal Carotid Artery\n- Left Middle Cerebral Artery\n- Right Middle Cerebral Artery\n- Anterior Communicating Artery\n- Left Anterior Cerebral Artery\n- Right Anterior Cerebral Artery\n- Left Posterior Communicating Artery\n- Right Posterior Communicating Artery\n- Basilar Tip\n- Other Posterior Circulation\n\n![brain arteries](https://upload.wikimedia.org/wikipedia/commons/9/99/Arteries_beneath_brain.png)\n\nAnd finally, a binary column indicating the presence of an aneurysm named\n- Aneurysm present\n\n","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize = (10,30))\n\ngrid = plt.GridSpec(9, 2, wspace = .25, hspace = .25)\n\nplt.subplot(grid[0,0])\nplt.hist(train['PatientAge'], bins = 20)\nplt.title(\"Patient Age\")\n\nplt.subplot(grid[0,1])\nplt.hist(train[\"PatientSex\"])\nplt.title(\"Patient Sex\")\n\nplt.subplot(grid[1,0])\nplt.hist(train[\"Modality\"])\nplt.title(\"Modality\")\n\nplt.subplot(grid[1,1])\nplt.hist(train['Left Infraclinoid Internal Carotid Artery'])\nplt.title(\"Location 1\")\n\nplt.subplot(grid[2,0])\nplt.hist(train['Right Infraclinoid Internal Carotid Artery'])\nplt.title(\"Location 2\")\n\nplt.subplot(grid[2,1])\nplt.hist(train['Left Supraclinoid Internal Carotid Artery'])\nplt.title('Location 3')\n\nplt.subplot(grid[3,0])\nplt.hist(train['Right Infraclinoid Internal Carotid Artery'])\nplt.title(\"Location 4\")\n\nplt.subplot(grid[3,1])\nplt.hist(train['Left Middle Cerebral Artery'])\nplt.title(\"Location 5\")\n\nplt.subplot(grid[4,0])\nplt.hist(train['Right Middle Cerebral Artery'])\nplt.title('Location 6')\n\nplt.subplot(grid[4,1])\nplt.hist(train['Anterior Communicating Artery'])\nplt.title(\"Location 7\")\n\nplt.subplot(grid[5,0])\nplt.hist(train['Left Anterior Cerebral Artery'])\nplt.title('Location 8')\n\nplt.subplot(grid[5,1])\nplt.hist(train['Right Anterior Cerebral Artery'])\nplt.title('Location 9')\n\nplt.subplot(grid[6,0])\nplt.hist(train['Left Posterior Communicating Artery'])\nplt.title('Location 10')\n\nplt.subplot(grid[6,1])\nplt.hist(train['Right Posterior Communicating Artery'])\nplt.title('Location 11')\n\nplt.subplot(grid[7,0])\nplt.hist(train['Basilar Tip'])\nplt.title('Location 12')\n\nplt.subplot(grid[7,1])\nplt.hist(train['Other Posterior Circulation'])\nplt.title('Location 13')\n\nplt.subplot(grid[8,0])\nplt.hist(train['Aneurysm Present'])\nplt.title('Aneurysm Present')\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:42:35.936471Z","iopub.execute_input":"2025-09-23T12:42:35.936785Z","iopub.status.idle":"2025-09-23T12:42:38.240239Z","shell.execute_reply.started":"2025-09-23T12:42:35.936761Z","shell.execute_reply":"2025-09-23T12:42:38.238863Z"},"collapsed":true,"jupyter":{"source_hidden":true,"outputs_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The age of the patients go from 18 up to 89 years, with an average age around 60 years.\n\nThere are more females than males, however, at this point, we are not sure if this is a determining factor to take into acount for the model.\n\nThe modality most used for taking the images is CTA, its probably better to start focusing on these type of images.\n\nThe initial target class(Aneurysm Present) is well distributed, however, there aren't many images for each class, so it will probably be necessary to use data augmentation when trying to classify the regions of the aneurysms.","metadata":{}},{"cell_type":"code","source":"len(train[train['Aneurysm Present'] == 1])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-21T18:22:17.835894Z","iopub.execute_input":"2025-09-21T18:22:17.836178Z","iopub.status.idle":"2025-09-21T18:22:17.843513Z","shell.execute_reply.started":"2025-09-21T18:22:17.836156Z","shell.execute_reply":"2025-09-21T18:22:17.842704Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Train localizers.csv","metadata":{}},{"cell_type":"code","source":"localizers = pd.read_csv(folder_path + '/train_localizers.csv')\nlocalizers","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-21T18:21:25.927151Z","iopub.execute_input":"2025-09-21T18:21:25.927797Z","iopub.status.idle":"2025-09-21T18:21:25.956414Z","shell.execute_reply.started":"2025-09-21T18:21:25.92777Z","shell.execute_reply":"2025-09-21T18:21:25.955702Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This csv contains 2251 observations with 4 columns corresponding to:\n- SeriesInstanceUID: The unique identifier of the patient\n- SOPInstanceUID: The path to the image where the aneurysm was detected\n- coordinates: The coordinates of the aneurysm in the respective image\n- location: The location of the aneurysm in the 13 possible regions defined in the train.csv file","metadata":{}},{"cell_type":"code","source":"len(localizers['SeriesInstanceUID'].unique())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-21T18:24:28.649363Z","iopub.execute_input":"2025-09-21T18:24:28.649719Z","iopub.status.idle":"2025-09-21T18:24:28.658909Z","shell.execute_reply.started":"2025-09-21T18:24:28.649694Z","shell.execute_reply":"2025-09-21T18:24:28.658123Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There are only 1862 patients with aneurysms detected but 2251 localizers, which implies that several patients have more than one aneurysm","metadata":{}},{"cell_type":"code","source":"localizers['SeriesInstanceUID'].value_counts().value_counts()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-21T18:48:35.226096Z","iopub.execute_input":"2025-09-21T18:48:35.226821Z","iopub.status.idle":"2025-09-21T18:48:35.23663Z","shell.execute_reply.started":"2025-09-21T18:48:35.226789Z","shell.execute_reply":"2025-09-21T18:48:35.235714Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"A patient can have up to 5 aneurysms in this dataset","metadata":{}},{"cell_type":"markdown","source":"## Series","metadata":{}},{"cell_type":"code","source":"len(os.listdir(f'{folder_path}/series'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:43:28.425749Z","iopub.execute_input":"2025-09-23T12:43:28.426142Z","iopub.status.idle":"2025-09-23T12:43:28.43611Z","shell.execute_reply.started":"2025-09-23T12:43:28.426114Z","shell.execute_reply":"2025-09-23T12:43:28.435078Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Series is a folder containing the images taken for all the patients, let's look at the images of a patient taken with each modality","metadata":{}},{"cell_type":"code","source":"modalities = train['Modality'].unique()\nmodalities","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:43:28.745658Z","iopub.execute_input":"2025-09-23T12:43:28.746024Z","iopub.status.idle":"2025-09-23T12:43:28.756557Z","shell.execute_reply.started":"2025-09-23T12:43:28.745997Z","shell.execute_reply":"2025-09-23T12:43:28.755317Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"mra = train[train['Modality'] == modalities[0]].iloc[0]\ncta = train[train['Modality'] == modalities[1]].iloc[0]\nmri2 = train[train['Modality'] == modalities[2]].iloc[0]\nmri1 = train[train['Modality'] == modalities[3]].iloc[0]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:43:28.931198Z","iopub.execute_input":"2025-09-23T12:43:28.931509Z","iopub.status.idle":"2025-09-23T12:43:28.945766Z","shell.execute_reply.started":"2025-09-23T12:43:28.931485Z","shell.execute_reply":"2025-09-23T12:43:28.944362Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"mra_images = os.listdir(f'{folder_path}/series/{mra[\"SeriesInstanceUID\"]}')\nlen(mra_images)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:43:29.161946Z","iopub.execute_input":"2025-09-23T12:43:29.162317Z","iopub.status.idle":"2025-09-23T12:43:29.230136Z","shell.execute_reply.started":"2025-09-23T12:43:29.16229Z","shell.execute_reply":"2025-09-23T12:43:29.229061Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"cta_images = os.listdir(f'{folder_path}/series/{cta[\"SeriesInstanceUID\"]}')\nlen(cta_images)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:43:30.918068Z","iopub.execute_input":"2025-09-23T12:43:30.918439Z","iopub.status.idle":"2025-09-23T12:43:30.983368Z","shell.execute_reply.started":"2025-09-23T12:43:30.918413Z","shell.execute_reply":"2025-09-23T12:43:30.982344Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"mri1_images = os.listdir(f'{folder_path}/series/{mri1[\"SeriesInstanceUID\"]}')\nlen(mri1_images)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:43:31.272988Z","iopub.execute_input":"2025-09-23T12:43:31.273333Z","iopub.status.idle":"2025-09-23T12:43:31.926551Z","shell.execute_reply.started":"2025-09-23T12:43:31.273309Z","shell.execute_reply":"2025-09-23T12:43:31.925346Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"mri2_images = os.listdir(f'{folder_path}/series/{mri2[\"SeriesInstanceUID\"]}')\nlen(mri2_images)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:43:31.928325Z","iopub.execute_input":"2025-09-23T12:43:31.928621Z","iopub.status.idle":"2025-09-23T12:43:31.939649Z","shell.execute_reply.started":"2025-09-23T12:43:31.928595Z","shell.execute_reply":"2025-09-23T12:43:31.938179Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It stands out the radically difference in the number of files on directories for different patients, the next question that arises is, is this difference due to the method used? Or is it different from patient to patient. This is important because if we use a 3D UNet, we will need to preprocess our data to have the same input size and we will have to choose a proper value for the dimensions\n\nFirst, lets check the files in each folder:","metadata":{}},{"cell_type":"code","source":"{mra[\"SeriesInstanceUID\"]}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:44:07.011034Z","iopub.execute_input":"2025-09-23T12:44:07.011402Z","iopub.status.idle":"2025-09-23T12:44:07.018013Z","shell.execute_reply.started":"2025-09-23T12:44:07.011374Z","shell.execute_reply":"2025-09-23T12:44:07.017181Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"os.listdir(f'{folder_path}/series/{mra[\"SeriesInstanceUID\"]}')[:10]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:44:07.335009Z","iopub.execute_input":"2025-09-23T12:44:07.335403Z","iopub.status.idle":"2025-09-23T12:44:07.343929Z","shell.execute_reply.started":"2025-09-23T12:44:07.335373Z","shell.execute_reply":"2025-09-23T12:44:07.342863Z"},"collapsed":true,"jupyter":{"outputs_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"os.listdir(f'{folder_path}/series/{cta[\"SeriesInstanceUID\"]}')[:10]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:44:07.694587Z","iopub.execute_input":"2025-09-23T12:44:07.694957Z","iopub.status.idle":"2025-09-23T12:44:07.703226Z","shell.execute_reply.started":"2025-09-23T12:44:07.694922Z","shell.execute_reply":"2025-09-23T12:44:07.70205Z"},"collapsed":true,"jupyter":{"outputs_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"os.listdir(f'{folder_path}/series/{mri1[\"SeriesInstanceUID\"]}')[:10]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:44:08.2971Z","iopub.execute_input":"2025-09-23T12:44:08.29743Z","iopub.status.idle":"2025-09-23T12:44:08.30493Z","shell.execute_reply.started":"2025-09-23T12:44:08.297395Z","shell.execute_reply":"2025-09-23T12:44:08.303983Z"},"collapsed":true,"jupyter":{"outputs_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"os.listdir(f'{folder_path}/series/{mri2[\"SeriesInstanceUID\"]}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:44:09.820873Z","iopub.execute_input":"2025-09-23T12:44:09.821829Z","iopub.status.idle":"2025-09-23T12:44:09.82879Z","shell.execute_reply.started":"2025-09-23T12:44:09.821786Z","shell.execute_reply":"2025-09-23T12:44:09.827719Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"All of them are .dcm files, meaning that there is just one image particullary for the MRI T2 observation taken, lets read a couple of the dcm files and see the images","metadata":{}},{"cell_type":"markdown","source":"### MRI 2","metadata":{}},{"cell_type":"code","source":"dcm_img = pydicom.dcmread(f'{folder_path}/series/{mri2[\"SeriesInstanceUID\"]}/' + mri2_images[0], force=True)\ndcm_img","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:44:13.872723Z","iopub.execute_input":"2025-09-23T12:44:13.873655Z","iopub.status.idle":"2025-09-23T12:44:14.021554Z","shell.execute_reply.started":"2025-09-23T12:44:13.873624Z","shell.execute_reply":"2025-09-23T12:44:14.020487Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"img_array = dcm_img.pixel_array\nprint(img_array.shape)\n\nplt.imshow(img_array[3], cmap = 'gray')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:44:14.271016Z","iopub.execute_input":"2025-09-23T12:44:14.271344Z","iopub.status.idle":"2025-09-23T12:44:14.855638Z","shell.execute_reply.started":"2025-09-23T12:44:14.271319Z","shell.execute_reply":"2025-09-23T12:44:14.85436Z"},"jupyter":{"source_hidden":true,"outputs_hidden":true},"collapsed":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In this case, even if there is only one file, this file contains 25 images of the patient, we will have to check other MRI T2 images to see if this is the general rule","metadata":{}},{"cell_type":"markdown","source":"### MRI 1","metadata":{}},{"cell_type":"code","source":"dcm_img = pydicom.dcmread(f'{folder_path}/series/{mri1[\"SeriesInstanceUID\"]}/' + mri1_images[0], force=True)\ndcm_img","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:47:25.675511Z","iopub.execute_input":"2025-09-23T12:47:25.675945Z","iopub.status.idle":"2025-09-23T12:47:25.702616Z","shell.execute_reply.started":"2025-09-23T12:47:25.675878Z","shell.execute_reply":"2025-09-23T12:47:25.701129Z"},"jupyter":{"source_hidden":true,"outputs_hidden":true},"collapsed":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"img_array = dcm_img.pixel_array\nprint(img_array.shape)\n\nplt.imshow(img_array, cmap = 'gray')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:47:25.867554Z","iopub.execute_input":"2025-09-23T12:47:25.867867Z","iopub.status.idle":"2025-09-23T12:47:26.146757Z","shell.execute_reply.started":"2025-09-23T12:47:25.867847Z","shell.execute_reply":"2025-09-23T12:47:26.145781Z"},"jupyter":{"source_hidden":true,"outputs_hidden":true},"collapsed":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### CTA","metadata":{}},{"cell_type":"code","source":"dcm_img = pydicom.dcmread(f'{folder_path}/series/{cta[\"SeriesInstanceUID\"]}/' + cta_images[0], force=True)\ndcm_img","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:47:29.309003Z","iopub.execute_input":"2025-09-23T12:47:29.309348Z","iopub.status.idle":"2025-09-23T12:47:29.336251Z","shell.execute_reply.started":"2025-09-23T12:47:29.309323Z","shell.execute_reply":"2025-09-23T12:47:29.335226Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"img_array = dcm_img.pixel_array\nprint(img_array.shape)\n\nplt.imshow(img_array, cmap = 'gray')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:47:29.471625Z","iopub.execute_input":"2025-09-23T12:47:29.471945Z","iopub.status.idle":"2025-09-23T12:47:29.728086Z","shell.execute_reply.started":"2025-09-23T12:47:29.471922Z","shell.execute_reply":"2025-09-23T12:47:29.726827Z"},"jupyter":{"source_hidden":true,"outputs_hidden":true},"collapsed":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### MRA","metadata":{}},{"cell_type":"code","source":"dcm_img = pydicom.dcmread(f'{folder_path}/series/{mra[\"SeriesInstanceUID\"]}/' + mra_images[0], force=True)\ndcm_img","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:47:41.914249Z","iopub.execute_input":"2025-09-23T12:47:41.91475Z","iopub.status.idle":"2025-09-23T12:47:41.941427Z","shell.execute_reply.started":"2025-09-23T12:47:41.914706Z","shell.execute_reply":"2025-09-23T12:47:41.939605Z"},"jupyter":{"source_hidden":true,"outputs_hidden":true},"collapsed":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"img_array = dcm_img.pixel_array\nprint(img_array.shape)\n\nplt.imshow(img_array, cmap = 'gray')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-23T12:47:42.085486Z","iopub.execute_input":"2025-09-23T12:47:42.085817Z","iopub.status.idle":"2025-09-23T12:47:42.368502Z","shell.execute_reply.started":"2025-09-23T12:47:42.085793Z","shell.execute_reply":"2025-09-23T12:47:42.367208Z"},"jupyter":{"source_hidden":true,"outputs_hidden":true},"collapsed":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It appears that only the MRI images come compressed into one file, so we would have to move it to multiple image files\n\nAll of the .dcm files come with many unused fields in the metadata, which can be erased during the preprocessing to reduce the size of the training files and make it portable to move to colab or other cloud servers.\n\nSome other information as the weight of the patient may be worth storing so we could add it to the train csv if its present for all the observations, otherwise we might erase this data.","metadata":{}},{"cell_type":"markdown","source":"## Segmentations","metadata":{}},{"cell_type":"code","source":"segmentations = os.listdir(f'{folder_path}/segmentations')\nlen(segmentations)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-21T20:23:49.987708Z","iopub.execute_input":"2025-09-21T20:23:49.988438Z","iopub.status.idle":"2025-09-21T20:23:49.995636Z","shell.execute_reply.started":"2025-09-21T20:23:49.988411Z","shell.execute_reply":"2025-09-21T20:23:49.994962Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"segmentations = os.listdir(f'{folder_path}/segmentations')\nsegmentations_uid = [i[:len(segmentations[0]) - 4] for i in segmentations]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-21T19:42:51.575558Z","iopub.execute_input":"2025-09-21T19:42:51.575978Z","iopub.status.idle":"2025-09-21T19:42:51.581709Z","shell.execute_reply.started":"2025-09-21T19:42:51.575936Z","shell.execute_reply":"2025-09-21T19:42:51.580768Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"len(segmentations)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-21T20:28:19.086519Z","iopub.execute_input":"2025-09-21T20:28:19.087137Z","iopub.status.idle":"2025-09-21T20:28:19.09254Z","shell.execute_reply.started":"2025-09-21T20:28:19.087109Z","shell.execute_reply":"2025-09-21T20:28:19.091679Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train[train['SeriesInstanceUID'].isin(segmentations_uid)]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-21T19:43:07.601766Z","iopub.execute_input":"2025-09-21T19:43:07.602619Z","iopub.status.idle":"2025-09-21T19:43:07.61944Z","shell.execute_reply.started":"2025-09-21T19:43:07.602594Z","shell.execute_reply":"2025-09-21T19:43:07.618515Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"path = f'{folder_path}/segmentations/{segmentations[0]}'\nimg = nib.load(path).get_fdata()\nimg.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-21T20:25:54.539103Z","iopub.execute_input":"2025-09-21T20:25:54.5394Z","iopub.status.idle":"2025-09-21T20:25:56.90122Z","shell.execute_reply.started":"2025-09-21T20:25:54.539378Z","shell.execute_reply":"2025-09-21T20:25:56.90033Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"segmentations_uid[0] in os.listdir(f'{folder_path}/series')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-21T20:29:15.78124Z","iopub.execute_input":"2025-09-21T20:29:15.781542Z","iopub.status.idle":"2025-09-21T20:29:15.790977Z","shell.execute_reply.started":"2025-09-21T20:29:15.781519Z","shell.execute_reply":"2025-09-21T20:29:15.790007Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"len(os.listdir(f'{folder_path}/series/{segmentations_uid[0]}'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-21T20:29:55.126758Z","iopub.execute_input":"2025-09-21T20:29:55.127049Z","iopub.status.idle":"2025-09-21T20:29:55.133691Z","shell.execute_reply.started":"2025-09-21T20:29:55.127027Z","shell.execute_reply":"2025-09-21T20:29:55.132851Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The segmentations folder is a redundancy of the series folder, which stores all the images taked in a single nii file instead of multiple dcm files, so it can be deleted","metadata":{}},{"cell_type":"markdown","source":"## kaggle_evaluation","metadata":{}},{"cell_type":"code","source":"os.listdir(f'{folder_path}/kaggle_evaluation')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-21T20:35:13.963957Z","iopub.execute_input":"2025-09-21T20:35:13.964587Z","iopub.status.idle":"2025-09-21T20:35:13.971437Z","shell.execute_reply.started":"2025-09-21T20:35:13.964563Z","shell.execute_reply":"2025-09-21T20:35:13.970557Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"os.listdir(f'{folder_path}/kaggle_evaluation/series')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-21T20:36:23.138018Z","iopub.execute_input":"2025-09-21T20:36:23.138309Z","iopub.status.idle":"2025-09-21T20:36:23.145192Z","shell.execute_reply.started":"2025-09-21T20:36:23.138286Z","shell.execute_reply":"2025-09-21T20:36:23.144472Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pd.read_csv(f'{folder_path}/kaggle_evaluation/test.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-21T20:36:01.951956Z","iopub.execute_input":"2025-09-21T20:36:01.952241Z","iopub.status.idle":"2025-09-21T20:36:01.968713Z","shell.execute_reply.started":"2025-09-21T20:36:01.952218Z","shell.execute_reply":"2025-09-21T20:36:01.967877Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This is a folder meant to do the submissions of the test cases, it has only 3 cases since the rest are hidden for the contest, then, this folder wont be used","metadata":{}},{"cell_type":"markdown","source":"# Data preprocessing","metadata":{}},{"cell_type":"code","source":"import torch\nfrom torchvision.io import read_image\nfrom sklearn.model_selection import train_test_split","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-25T08:26:29.889248Z","iopub.execute_input":"2025-09-25T08:26:29.889596Z","iopub.status.idle":"2025-09-25T08:26:40.15982Z","shell.execute_reply.started":"2025-09-25T08:26:29.88954Z","shell.execute_reply":"2025-09-25T08:26:40.158853Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Make an isotropic transformation for the data and convert to NIFT1 files\n\nWe want the spacing between voxels to be the same across all images, this will ease the training for the network, in some images, this will imply to add more entries, and more images, in other cases, this will imply to downsize the images.\n\nThis will be done using a linear interpolation from the itk library, which will allow us to read the images, determine the spacing between voxels and resample the image to the desired space between voxels.\n\nAfter looking at the size of the images in Mb after the resampling, its best to keep the spacing between the voxels in 1mm, since reducing this space increases dramatically the size of the images, making the dataset too heavy to be workable.\n\nWhen performing the resampling operation, all the slices will be joint into one simgle 3D image, which has to be stored in a different format than the multilpe dicom files in the original dataset. Thus, after performing the resampling, these will be stored in .nii files","metadata":{}},{"cell_type":"code","source":"!pip install itk","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-26T12:23:50.552501Z","iopub.execute_input":"2025-09-26T12:23:50.55291Z","iopub.status.idle":"2025-09-26T12:24:25.073243Z","shell.execute_reply.started":"2025-09-26T12:23:50.552879Z","shell.execute_reply":"2025-09-26T12:24:25.071939Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import itk\n\ndef resample_image_and_convert_to_nii(serie, voxel_size = 1):\n    # Load your 3D image (from DICOM or already built)\n    image = itk.imread(folder_path+\"/series/\" + serie, itk.F)  # itk.imread can read a folder of DICOMs too\n    \n    # Original info\n    original_spacing = image.GetSpacing()\n    original_size = image.GetLargestPossibleRegion().GetSize()\n    original_origin = image.GetOrigin()\n    original_direction = image.GetDirection()\n    \n    # Desired new spacing\n    new_spacing = [voxel_size, voxel_size, voxel_size] #Make distances to be 1mm in each dimension\n    \n    # Compute new size to keep the same physical extent\n    new_size = [\n        int(round(orig_sz * orig_spc / new_spc))\n        for orig_sz, orig_spc, new_spc in zip(original_size, original_spacing, new_spacing)\n    ]\n    \n    # Resampling\n    resample = itk.ResampleImageFilter.New(Input=image)\n    resample.SetInterpolator(itk.LinearInterpolateImageFunction.New(InputImage=image))\n    resample.SetOutputSpacing(new_spacing)\n    resample.SetSize(new_size)\n    resample.SetOutputOrigin(original_origin)\n    resample.SetOutputDirection(original_direction)\n    resample.SetTransform(itk.IdentityTransform[itk.D, 3].New())\n    \n    resampled_image = resample.GetOutput()\n    \n    # Save result\n    itk.imwrite(resampled_image, serie + \"resampled_0.5mm.nii.gz\")\n\nfor serie in tqdm(series[:10]):\n    resample_image_and_convert_to_nii(serie)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-26T13:12:44.462978Z","iopub.execute_input":"2025-09-26T13:12:44.463303Z","iopub.status.idle":"2025-09-26T13:16:13.635159Z","shell.execute_reply.started":"2025-09-26T13:12:44.463281Z","shell.execute_reply":"2025-09-26T13:16:13.63427Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.imshow(image[101], cmap = 'gray')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-26T12:56:47.31682Z","iopub.execute_input":"2025-09-26T12:56:47.317174Z","iopub.status.idle":"2025-09-26T12:56:47.566981Z","shell.execute_reply.started":"2025-09-26T12:56:47.317152Z","shell.execute_reply":"2025-09-26T12:56:47.565549Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"resampled_image.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-26T12:56:38.842606Z","iopub.execute_input":"2025-09-26T12:56:38.843172Z","iopub.status.idle":"2025-09-26T12:56:38.851194Z","shell.execute_reply.started":"2025-09-26T12:56:38.843138Z","shell.execute_reply":"2025-09-26T12:56:38.850043Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.imshow(resampled_image[101], cmap = 'gray')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-26T12:56:56.119443Z","iopub.execute_input":"2025-09-26T12:56:56.119857Z","iopub.status.idle":"2025-09-26T12:56:56.349854Z","shell.execute_reply.started":"2025-09-26T12:56:56.11983Z","shell.execute_reply":"2025-09-26T12:56:56.348628Z"}},"outputs":[],"execution_count":null}]}