{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# From 2D to 3D : A Beginner's Guide to Transforming DICOM images into NIFTI volumes!\n\n## What is NIFTI?\n* NIFTI (Neuroimaging Informatics Technology Initiative) is a file format commonly used for storing medical imaging data, particularly in the field of neuroimaging. It's an extension of the older analyze format but includes several additional features like storing metadata. NIFTI files usually have the extensions .nii or .nii.gz for compressed versions.\n\n* The NIFTI format can store 3D volumes (like a CT or MRI scan) or even 4D data (like a time series of 3D volumes). It's widely used because it can also store additional information like the spatial resolution of the images, making it easier to perform subsequent analyses that require real-world spatial dimensions.","metadata":{"papermill":{"duration":0.005157,"end_time":"2023-10-08T02:40:07.363793","exception":false,"start_time":"2023-10-08T02:40:07.358636","status":"completed"},"tags":[]}},{"cell_type":"code","source":"try:\n    import pydicom\n    \nexcept:\n    !pip install pydicom","metadata":{"execution":{"iopub.status.busy":"2023-10-19T03:38:52.864129Z","iopub.execute_input":"2023-10-19T03:38:52.864414Z","iopub.status.idle":"2023-10-19T03:38:53.025895Z","shell.execute_reply.started":"2023-10-19T03:38:52.864392Z","shell.execute_reply":"2023-10-19T03:38:53.025017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\nimport pydicom\nimport nibabel as nib\n\n","metadata":{"execution":{"iopub.status.busy":"2023-10-19T03:38:53.027354Z","iopub.execute_input":"2023-10-19T03:38:53.027591Z","iopub.status.idle":"2023-10-19T03:38:53.387752Z","shell.execute_reply.started":"2023-10-19T03:38:53.027573Z","shell.execute_reply":"2023-10-19T03:38:53.386652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Let's check what we got!","metadata":{}},{"cell_type":"code","source":"TRAIN_PATH = \"/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images\"\npatient_id = [f for f in os.listdir(\"/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images\")]","metadata":{"execution":{"iopub.status.busy":"2023-10-19T03:38:53.389244Z","iopub.execute_input":"2023-10-19T03:38:53.389678Z","iopub.status.idle":"2023-10-19T03:38:53.719897Z","shell.execute_reply.started":"2023-10-19T03:38:53.389654Z","shell.execute_reply":"2023-10-19T03:38:53.71855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"series_idx = []\n\nfor i in range(len(patient_id)):\n\n    patients_path = os.path.join(TRAIN_PATH, patient_id[i])\n    \n    if not len(os.listdir(patients_path)) ==0:\n    \n        for j in os.listdir(patients_path):\n            \n            series_idx.append(j)\n            \n    else:\n        print(f\"There is an empty patient file! {patients_path}\")","metadata":{"execution":{"iopub.status.busy":"2023-10-19T03:38:53.721307Z","iopub.execute_input":"2023-10-19T03:38:53.721633Z","iopub.status.idle":"2023-10-19T03:39:10.281045Z","shell.execute_reply.started":"2023-10-19T03:38:53.721605Z","shell.execute_reply":"2023-10-19T03:39:10.280047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"You have {len(patient_id)} PATIENTS in image folders.\")\nprint(f\"You have {len(series_idx)} SERIES in image folders.\")","metadata":{"execution":{"iopub.status.busy":"2023-10-19T03:39:10.283549Z","iopub.execute_input":"2023-10-19T03:39:10.284267Z","iopub.status.idle":"2023-10-19T03:39:10.289682Z","shell.execute_reply.started":"2023-10-19T03:39:10.284233Z","shell.execute_reply":"2023-10-19T03:39:10.288509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.read_csv(\"/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv\")\n\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-19T03:39:10.290927Z","iopub.execute_input":"2023-10-19T03:39:10.291242Z","iopub.status.idle":"2023-10-19T03:39:10.34316Z","shell.execute_reply.started":"2023-10-19T03:39:10.291211Z","shell.execute_reply":"2023-10-19T03:39:10.34256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print( f\"You have {df['patient_id'].nunique()} PATIENTS in the Train Dataset\")","metadata":{"execution":{"iopub.status.busy":"2023-10-19T03:39:10.344065Z","iopub.execute_input":"2023-10-19T03:39:10.344394Z","iopub.status.idle":"2023-10-19T03:39:10.35398Z","shell.execute_reply.started":"2023-10-19T03:39:10.344373Z","shell.execute_reply":"2023-10-19T03:39:10.352777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## .nii / .nii.gz\n\n* Both .nii and .nii.gz are valid NIFTI file extensions, but they serve different purposes:\n\n* .nii: This is the uncompressed NIFTI file format. It stores all the header and image data in a single file. While it's straightforward and quickly readable, it takes up more disk space.\n\n* .nii.gz: This is the gzip-compressed version of the NIFTI file. It significantly reduces the file size, which is beneficial when storing or transferring large volumes of data. The cost is a slightly longer time to read or write the file due to the compression and decompression steps. However, modern computers are generally fast enough that this time is often negligible.","metadata":{}},{"cell_type":"code","source":"for i in range(2): #You can change here to deal with large data\n    \n    patient_path = os.path.join(TRAIN_PATH,patient_id[i])\n    \n    if not os.path.isdir(patient_path):\n        continue\n    \n    # Create a new directory for the patient to store NIFTI files\n    nifti_patient_path = os.path.join('/kaggle/working/nifti', patient_id[i])\n    Path(nifti_patient_path).mkdir(parents=True, exist_ok=True)\n    \n    for j in os.listdir(patient_path):\n        series_path = os.path.join(patient_path,j)\n        \n        if not os.path.isdir(series_path):\n            continue\n        \n        dicom_file_names = [f for f in os.listdir(series_path) if os.path.isfile(os.path.join(series_path, f))]\n        dicom_files = [pydicom.dcmread(os.path.join(series_path, f)) for f in dicom_file_names]\n        dicom_files.sort(key=lambda x: int(x.InstanceNumber))\n        \n        volume_shape = (len(dicom_files), dicom_files[0].Rows, dicom_files[0].Columns)\n        volume_3d = np.zeros(volume_shape, dtype=np.int16)\n\n        # Fill the 3D array with the pixel arrays from each DICOM file\n        for c, ds in enumerate(dicom_files):\n            volume_3d[c, :, :] = ds.pixel_array\n\n        # Create NIFTI image (using an identity affine matrix for simplicity)\n        affine = np.eye(4)\n        img = nib.Nifti1Image(volume_3d, affine)\n        \n        # Save NIFTI file\n        nifti_series_path = os.path.join(nifti_patient_path, j + '.nii.gz')\n        nib.save(img, nifti_series_path)\n        print(f\"Saved NIFTI for patient {patient_id[i]} and series {j} to {nifti_series_path}\")","metadata":{"papermill":{"duration":22.863866,"end_time":"2023-10-08T02:40:31.506501","exception":false,"start_time":"2023-10-08T02:40:08.642635","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-10-19T03:39:10.35502Z","iopub.execute_input":"2023-10-19T03:39:10.355946Z","iopub.status.idle":"2023-10-19T03:41:19.578769Z","shell.execute_reply.started":"2023-10-19T03:39:10.35591Z","shell.execute_reply":"2023-10-19T03:41:19.577634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Reasons for using .nii.gz:\n\n1. **Disk Space:** Medical images can be large, especially for 3D volumes or time-series data. Using .nii.gz can save a substantial amount of disk space.\n\n2. **Data Transfer:** Smaller files are quicker to transfer over a network, which can be crucial in settings where large datasets are routinely shared.\n\n3. **Standard Practice:** The compressed format is commonly used in various neuroimaging research and pipelines, making it a widely-accepted standard.\n\n4. **Compatibility:** Most modern neuroimaging software can read compressed NIFTI files directly without requiring manual decompression first.\n\n* *In summary* .nii.gz is often chosen as a trade-off between disk space and read/write speed, and it has become a widely-used standard in the field. If disk space and data transfer are not concerns, and you prioritize read/write speed, you might opt for the .nii format.","metadata":{}},{"cell_type":"code","source":"# Read the NIFTI file\nimg = nib.load(nifti_series_path)\n\n# Convert image to numpy array\nimg_data = img.get_fdata()\n\n# Plot a slice in the middle along the Z-axis\nx_middle = img_data.shape[0] // 2\nplt.imshow(img_data[x_middle, :, :], cmap='gray')\nplt.title(f'Slice at {x_middle}')\nplt.axis('off')\n\n# Show the plot\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-19T03:41:19.580481Z","iopub.execute_input":"2023-10-19T03:41:19.581611Z","iopub.status.idle":"2023-10-19T03:41:22.509168Z","shell.execute_reply.started":"2023-10-19T03:41:19.581572Z","shell.execute_reply":"2023-10-19T03:41:22.50829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## IPywidgets\n\n* You could add interactive widgets using libraries like ipywidgets to let users dynamically select which slice of the 3D volume they'd like to view. This would give beginners a better understanding of how 3D volumes are stacked from 2D slices.","metadata":{}},{"cell_type":"code","source":"import ipywidgets as widgets\nfrom ipywidgets import interact\n\n# Read the NIFTI file\nimg = nib.load(nifti_series_path)\n\n# Convert image to numpy array\nimg_data = img.get_fdata()\n\n@interact(Slices=widgets.IntSlider(min=0, max=img_data.shape[2]-1, step=1, value=0))\ndef plot_slice(Slices):\n    plt.imshow(img_data[Slices, :, :], cmap='gray')\n    plt.axis('off')\n    plt.title(nifti_series_path)\n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-10-19T03:41:22.51031Z","iopub.execute_input":"2023-10-19T03:41:22.510734Z","iopub.status.idle":"2023-10-19T03:41:25.502411Z","shell.execute_reply.started":"2023-10-19T03:41:22.510708Z","shell.execute_reply":"2023-10-19T03:41:25.501803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Thanks for reading\n\n* I hope that notebook would be beneficial for your future studies!!","metadata":{}}]}