{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Preprocessing DCM Images\n\nFollowing the Notebook on TorchIO ([here](https://www.kaggle.com/code/fepegar/torchio-3d-loading-preprocessing-augmentation)), one can learn some useful things for working with and processing the dcm files. Inspired by it, I would like to show my way of preprocessing the images in this competition. If you run this notebook locally, you can produce an **optimized data set for training**. The size of the full data set is too large to publish it in one go here. However, it is still much less heavy than the original one.\n\nThe Preprocessing contains:\n- Intensity normalization to Hounsfield scale\n- Resetting image orientation to canonical one\n- Reducing image step accuracy to 1 mm\n- Cropping and padding of the images\n\nThis preprocessing allowed reducing the size of the images and set them to the same scale and dimention. This accelerates and eases further model training.\n\nFollowing, torchio is installed. The full rsna-2022-cervical-spine-fracture-detection directory is used as an input because on ([TorchIO](https://torchio.readthedocs.io)) there is a special library developed for this competition and this data directory. It creates an instance of <code>RSNACervicalSpineFracture</code>, which inherits from <code>SubjectsDataset</code> and from <code>torch.utils.data.Dataset</code>.\n","metadata":{}},{"cell_type":"code","source":"# Install torchio\n!pip install torchio\n\n# Import environment\nimport torchio as tio\nfrom torchio.datasets import RSNACervicalSpineFracture\nfrom tqdm import tqdm\nimport os\n\n# Read dataset\ntioDataset = RSNACervicalSpineFracture(root_dir=\"/kaggle/input/rsna-2022-cervical-spine-fracture-detection\",\n                                       add_segmentations=False, add_bounding_boxes=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-31T20:16:03.224976Z","iopub.execute_input":"2022-08-31T20:16:03.225537Z","iopub.status.idle":"2022-08-31T20:16:17.098978Z","shell.execute_reply.started":"2022-08-31T20:16:03.225485Z","shell.execute_reply":"2022-08-31T20:16:17.097362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Intensity Rescaling\n\nFollowing the best practices described in [this notebook](https://www.kaggle.com/code/fepegar/torchio-3d-loading-preprocessing-augmentation#Intensity-preprocessing), the intensity has been transformed to [Hounsfield Scale](https://en.wikipedia.org/wiki/Hounsfield_scale) and rescaled to be in the range between 0 and 1. Note, that everything below 5% and above 95% percentile has been cut off to avoid extreme outliers.","metadata":{}},{"cell_type":"code","source":"HOUNSFIELD_AIR, HOUNSFIELD_BONE = -1000, 1900\n# Transform to Honsfield Scale\nclamp = tio.Clamp(out_min=HOUNSFIELD_AIR, out_max=HOUNSFIELD_BONE)\n# Rescale to 0..1 scale and remove outliers\nrescale = tio.RescaleIntensity(percentiles=(0.5, 99.5))\npreprocess_intensity = tio.Compose([clamp, rescale,])","metadata":{"execution":{"iopub.status.busy":"2022-08-31T20:16:17.102311Z","iopub.execute_input":"2022-08-31T20:16:17.103623Z","iopub.status.idle":"2022-08-31T20:16:17.110214Z","shell.execute_reply.started":"2022-08-31T20:16:17.103536Z","shell.execute_reply":"2022-08-31T20:16:17.1089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Space Orintation\n\nThe 3D MRI images may have multiple spacial orientations, which can be very roughly imagined as different angles of view on the same object. This may result into different location of important elements on the images. Thus, it is important to transform all images to the same Spacial Orientation. This is done in the following code part:","metadata":{}},{"cell_type":"code","source":"# Transform all images to canonical spacial orientation\nnormalize_orientation = tio.ToCanonical()\ntype(normalize_orientation)","metadata":{"execution":{"iopub.status.busy":"2022-08-31T20:16:17.111619Z","iopub.execute_input":"2022-08-31T20:16:17.111977Z","iopub.status.idle":"2022-08-31T20:16:17.125387Z","shell.execute_reply.started":"2022-08-31T20:16:17.111945Z","shell.execute_reply":"2022-08-31T20:16:17.124044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Normalizing Image Size\n\nThe given step of the MRI scan is 0.4 mm. This is a quite generous resolution which can be downsampled to 1 mm. However, at the same time one should not forget that all the images have different size and the downsampling may not really fit the image size (one can't devide the size of all images to 1mm slices without remainder). That is why, one can resize all the images to the same dimention. Here, 20x20x20 cm is chosen. It makes sense also to crop and pad all the images after downsampling and resizing. This will ensure same image sizes by cutting off extra parts or filling up with zeroes the missing parts. It is better to train on a set of images of the same size.","metadata":{}},{"cell_type":"code","source":"# Downsample to the resolution of 1 mm\ndownsample = tio.Resample(1) # mm\ncrop_or_pad = tio.CropOrPad((200,200,200)) # mm\n\npreprocess_spatial = tio.Compose([\n       normalize_orientation,\n       downsample,\n       crop_or_pad\n           ])","metadata":{"execution":{"iopub.status.busy":"2022-08-31T20:16:17.127083Z","iopub.execute_input":"2022-08-31T20:16:17.127899Z","iopub.status.idle":"2022-08-31T20:16:17.136973Z","shell.execute_reply.started":"2022-08-31T20:16:17.127858Z","shell.execute_reply":"2022-08-31T20:16:17.135658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Put it all together\n\nAfter main steps are done, one can put the whole preprocessing together. \n\nFirst, combine all the steps above:","metadata":{}},{"cell_type":"code","source":"preprocess = tio.Compose([\n        preprocess_intensity,\n        preprocess_spatial,\n            ])","metadata":{"execution":{"iopub.status.busy":"2022-08-31T20:16:17.141385Z","iopub.execute_input":"2022-08-31T20:16:17.141818Z","iopub.status.idle":"2022-08-31T20:16:17.151053Z","shell.execute_reply.started":"2022-08-31T20:16:17.141782Z","shell.execute_reply":"2022-08-31T20:16:17.149572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next step is to iterate through all the dcm file-sets in the input data, preprocess them and write them the new preprocessed <code>.nii</code> (NifTi) files. Unlike in the input, the new data set contains one <code>.nii</code> file per patient.\n\nNote, the last code part is in \"demonstration mode\" and will not work on the fly. Replace the <code>path_output</code> with your local path where you will save the files on your local machine. Also remove the break statement in the loop to let it run completely. Here, it is inserted to not overload the resource given by kaggle.","metadata":{}},{"cell_type":"code","source":"# Create directory for output data\npath_output = \"/kaggle/working/rsna2022/preprocessed_1mm_200x200x200mm/\"\nif not os.path.exists(path_output):\n        os.makedirs(path_output)\n\n#Iterate through all the input and preprocess it\ni = 0\nfor subj in tqdm(tioDataset.dry_iter()):    \n    i += 1\n    if(i == 10): #REMOVE !!\n        break # REMOVE !!\n    file_to_save = path_output+subj.StudyInstanceUID+\".nii\"\n    print(subj.StudyInstanceUID)\n    if (not os.path.exists(file_to_save)):\n        # Preprocessing\n        subj_preprocessed = preprocess(subj)\n        # Saving\n        subj_preprocessed.ct.save(path_output+subj.StudyInstanceUID+\".nii\" )","metadata":{"execution":{"iopub.status.busy":"2022-08-31T20:16:17.152852Z","iopub.execute_input":"2022-08-31T20:16:17.153906Z","iopub.status.idle":"2022-08-31T20:18:21.247387Z","shell.execute_reply.started":"2022-08-31T20:16:17.153861Z","shell.execute_reply":"2022-08-31T20:18:21.245349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conclusions\n\nWith the data produced by this notebook I was able to reduce the time of reading the data **from 7 seconds per iteration (pation) to 7 iteration per second**. This means the time reduction of almost by 50 times.","metadata":{}}]}