{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:bold; letter-spacing: 2px; color:#006600; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #003300\">DCIM to NPY</p>","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pydicom\nimport os\nimport tqdm\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\n\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2023-08-01T15:13:32.259384Z","iopub.execute_input":"2023-08-01T15:13:32.260138Z","iopub.status.idle":"2023-08-01T15:13:32.4782Z","shell.execute_reply.started":"2023-08-01T15:13:32.260065Z","shell.execute_reply":"2023-08-01T15:13:32.477114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px; border:#DEB887 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n    \nWe are given a huge dataset of almost $460$ $GigaBytes(GB)$, but we cannot directly send this data to model pipelines, first we need to convert this to something `numerical`. There are a lot of numerical procedures out there, but we will be focusing on `NPY` files","metadata":{}},{"cell_type":"markdown","source":"<div style=\"border-radius:10px; border:#DEB887 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\nBefore diving into how we can convert the DCIM to NPY format, first lets understand what information does this file holds. If we open a file we get this information \n\n|||||\n|---|---|---|---|\n|File Meta Information Version|`OB: b'\\x00\\x01'`|`OB` - Data is stored in `Ocetet` sequence of `Bytes`|\n|||`b'\\x00\\x01'` - DICOM standars file version|\n|Media Storage SOP Class UID |`UI: CT Image Storage`|`UI` - Unique Identifier\n|||`CT Image Storage` - Computed Topography Image Storage\n|Media Storage SOP Instance UID|`UI: 1.2.123.12345.1.2.3.10004.1.1000`|`UI` - Unique Identifier\n|||`1.2.123.12345.1.2.3.10004.1.1000` - Unique Identifier for every file\n|Transfer Syntax UID|`UI: RLE Lossless`|`UI` - Unique Identifier\n|||`RLE Lossless` -Run Length Encoding , a lossless compression method, works by replacing repeated sequences of bytes with a single byte, significantly reducing the size of the image file without losing any data.\n|Implementation Class UID|`UI: 1.2.3.123456.4.5.1234.1.12.0`|`UI` - Unique Identifier\n|||`1.2.3.123456.4.5.1234.1.12.0` - Unique Identifier for Machine \n|Implementation Version Name|`SH: 'PYDICOM 2.4.0'`|`Pydicom` version used \n|SOP Instance UID|`UI: 1.2.123.12345.1.2.3.10004.1.1000`| Same as `Media Storage SOP Instance UID`\n|Content Date|`DA : '20230721'`|Date - July 21, 2023\n|Content Time|`TM : '232853.126781'`| Time - 23:28:53.126781\n|Patient ID|\n|Slice Thikness|`DS : 1.0`|`DS` - DataSet\n||`1.0` | 1mm\n|KVP|`DS : '90.0'`|`DS` - DataSet\n|||`90.0` Kilo Voltage Peak\n|Patient Position | `CS : 'HFS'`|`CS` - Code String\n||`HFS`| Head First Supine\n|Image Position(Patient)|`DS: [-249.02441, -392.5244, -1596.8]`|x , y , z coordinations of the Patient\n|Image Orientation(Patient)|`DS: [1.0, 0.0, 0.0, 0.0, 1.0, 0.0]`| Lookup Table for Value Of Interest Lookup\n|Samples Per Pixel|`US : '1'`| Number of samples taken per pixel\n|Photometric Interpretation | `CS : 'MONOCHROME1'`| Type of PhotoMetirc Interpretation\n|Rows|`US : '512'`| Number of Rows\n|Columns|`US : '512'`|Number of Columns|\n|Pixel Spacing| `DS: [0.9511719, 0.9511719]`| Distance between Pixels\n|Bits Allocated| `US : 16`| Number of bits allocated\n|Bits Stored|`US : 12`| Number of bits where the data is stored\n|High Bit| `US : 11`| Where the highest(most valuable) bit is stored\n|Pixel Representation|`US : 0`| How pixel is represented/stored, 0 represents un-asignined \n|Window Center|`DS : -50`|Range of pixel values that are mapped to a specific range of display values\n|Window Width| `DS : 400`|Range of pixel values that are displayed\n|Rescale Intercept|`DS : '-1024'`|Value that is subtracted from the pixel values in the image before they are displayed.\n|Rescale Slope|`DS : 1`|Value that is used to convert the pixel values in the image to display values\n|Rescale Type|`LS : 'HU'`|Units of the rescale slope and intercept `Hounsfield Units`\n|Pixel Data|`OB: Array of 252188 elements`| Number of elements in the pixel\n    \n<img src = \"https://miro.medium.com/v2/resize:fit:828/format:webp/0*k6A9sw3-EQQOq9J9.png\">","metadata":{}},{"cell_type":"markdown","source":"<div style=\"border-radius:10px; border:#DEB887 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\nIf we try to read a sample file like ","metadata":{}},{"cell_type":"code","source":"file = pydicom.read_file('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/10004/21057/1000.dcm')\nfile","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2023-08-01T15:13:33.448635Z","iopub.execute_input":"2023-08-01T15:13:33.449082Z","iopub.status.idle":"2023-08-01T15:13:33.478577Z","shell.execute_reply.started":"2023-08-01T15:13:33.449039Z","shell.execute_reply":"2023-08-01T15:13:33.477202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px; border:#DEB887 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\nIt gives us a lot of information\n\nWe can easily convert this to `NPY` array","metadata":{}},{"cell_type":"code","source":"file.pixel_array , file.pixel_array.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-01T15:13:33.909251Z","iopub.execute_input":"2023-08-01T15:13:33.910506Z","iopub.status.idle":"2023-08-01T15:13:33.938608Z","shell.execute_reply.started":"2023-08-01T15:13:33.910455Z","shell.execute_reply":"2023-08-01T15:13:33.936795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px; border:#DEB887 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\n$VOI$ $LUT$ is a `Look up table` provided which `become important peice of information` when dealing with this type of data. We can easily apply every attribute of this function like this ","metadata":{}},{"cell_type":"code","source":"apply_voi_lut(file.pixel_array , file)","metadata":{"execution":{"iopub.status.busy":"2023-08-01T15:13:34.380386Z","iopub.execute_input":"2023-08-01T15:13:34.381236Z","iopub.status.idle":"2023-08-01T15:13:34.396644Z","shell.execute_reply.started":"2023-08-01T15:13:34.381196Z","shell.execute_reply":"2023-08-01T15:13:34.395412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px; border:#DEB887 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\nFurther we need to check if the file is `MonoChrome` or not. Monochrome files are basically inverted in color, which can lead to issues","metadata":{}},{"cell_type":"code","source":"file.PhotometricInterpretation","metadata":{"execution":{"iopub.status.busy":"2023-08-01T15:13:34.939446Z","iopub.execute_input":"2023-08-01T15:13:34.939918Z","iopub.status.idle":"2023-08-01T15:13:34.94825Z","shell.execute_reply.started":"2023-08-01T15:13:34.939881Z","shell.execute_reply":"2023-08-01T15:13:34.946772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px; border:#DEB887 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\nNow lets try to make a function for processing the files","metadata":{}},{"cell_type":"code","source":"def read_xray(path, voi_lut = True, fix_monochrome = True):\n    \n    file = pydicom.read_file(path)\n    \n    data = file.pixel_array\n    \n    if voi_lut : data = apply_voi_lut(data , file)\n    if fix_monochrome and file.PhotometricInterpretation == \"MONOCHROME2\":data = np.amax(data) - data\n        \n    data = data - np.min(data)\n    data = data / np.max(data)\n    \n    data = (data * 255).astype(np.uint8)\n        \n    return data","metadata":{"execution":{"iopub.status.busy":"2023-08-01T15:13:35.761403Z","iopub.execute_input":"2023-08-01T15:13:35.761789Z","iopub.status.idle":"2023-08-01T15:13:35.770054Z","shell.execute_reply.started":"2023-08-01T15:13:35.76176Z","shell.execute_reply":"2023-08-01T15:13:35.768369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px; border:#DEB887 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\nIf we choose a sample image and see the difference for MonoChrome","metadata":{}},{"cell_type":"code","source":"img = read_xray('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/10004/21057/1000.dcm' , fix_monochrome = False)\nplt.figure(figsize = (6,6))\nplt.imshow(img, 'gray')\nplt.title(\"Image before fixing MonoChrome\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-08-01T15:13:36.324503Z","iopub.execute_input":"2023-08-01T15:13:36.324956Z","iopub.status.idle":"2023-08-01T15:13:36.788592Z","shell.execute_reply.started":"2023-08-01T15:13:36.324922Z","shell.execute_reply":"2023-08-01T15:13:36.787209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img = read_xray('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/10004/21057/1000.dcm')\nplt.figure(figsize = (6,6))\nplt.imshow(img, 'gray')\nplt.title(\"Image after fixing MonoChrome\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-08-01T15:13:36.79082Z","iopub.execute_input":"2023-08-01T15:13:36.79121Z","iopub.status.idle":"2023-08-01T15:13:37.192584Z","shell.execute_reply.started":"2023-08-01T15:13:36.791176Z","shell.execute_reply":"2023-08-01T15:13:37.191396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px; border:#DEB887 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\nSo how much would be a `NPY` file cost in memory compared to `DCM` file ","metadata":{}},{"cell_type":"code","source":"sample_file = file.pixel_array\n \nprint(\"Memory size of one array element in bytes: \", sample_file.itemsize)\n \nprint(\"Memory size of numpy array in bytes:\", sample_file.size * sample_file.itemsize)\n\nprint(\"Memory size of numpy array in Mega Bytes\" , (sample_file.size * sample_file.itemsize)/1048576)","metadata":{"execution":{"iopub.status.busy":"2023-08-01T15:13:37.440774Z","iopub.execute_input":"2023-08-01T15:13:37.441876Z","iopub.status.idle":"2023-08-01T15:13:37.448739Z","shell.execute_reply.started":"2023-08-01T15:13:37.441839Z","shell.execute_reply":"2023-08-01T15:13:37.447515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px; border:#DEB887 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\nWe can see the NPY file costs almost double of the DICOM file, which can raise the dataset cost upto $1TB$, whereas the maximum amount of dataset that a profile on Kaggle can hold is $109GB$ and one notebook is $19.5GB$, ","metadata":{}},{"cell_type":"code","source":"sample_dicom = '/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/10004/21057/1000.dcm'\n\nos.stat(sample_dicom).st_size","metadata":{"execution":{"iopub.status.busy":"2023-08-01T15:13:38.051889Z","iopub.execute_input":"2023-08-01T15:13:38.052315Z","iopub.status.idle":"2023-08-01T15:13:38.060852Z","shell.execute_reply.started":"2023-08-01T15:13:38.052282Z","shell.execute_reply":"2023-08-01T15:13:38.059516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"border-radius:10px; border:#DEB887 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n\nDespite of resoruces this conversion also takes a lot of time $79$ iterations, taking space upto $19.5GB$ , cost upto $17$ $Minutes$, and almost $7.7GB$ of RAM, I am trying to find a way to shift this to $CUDA$","metadata":{}},{"cell_type":"code","source":"train_list = sorted(os.listdir('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images'))[:79]\n\nfor folder_1 in tqdm.tqdm(train_list , total = len(train_list)):\n    \n    folder_1_list = sorted(os.listdir('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/' + folder_1))\n    \n    os.makedirs('/kaggle/working/Inputs/' + folder_1 + '/')\n    lis = list()\n    \n    for folder_2 in folder_1_list:\n        \n        folder_2_list = sorted(os.listdir('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/' + folder_1 + '/' + folder_2))\n        \n        for files in folder_2_list:\n            \n            file = pydicom.read_file('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/' + folder_1 + '/' + folder_2 + '/' + files)\n            \n            arr = file.pixel_array\n            arr = np.resize(arr , new_shape = (512 , 512))\n            lis.append(arr)\n            \n        np.save('/kaggle/working/Inputs/' + folder_1 + '/' + 'file', np.stack(lis , -1))","metadata":{"execution":{"iopub.status.busy":"2023-08-01T15:13:38.712415Z","iopub.execute_input":"2023-08-01T15:13:38.712814Z","iopub.status.idle":"2023-08-01T15:15:35.871121Z","shell.execute_reply.started":"2023-08-01T15:13:38.712783Z","shell.execute_reply":"2023-08-01T15:15:35.869318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:bold; letter-spacing: 2px; color:#E77200; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #E77200\">1 | TO DO LIST 📑</p>\n\n<div style=\"border-radius:10px; border:#E77200 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n    \n* $TO$ $DO$ $1$ $:$ $TRY$ $TO$ $CONVERT$ $ALL$ $VALUE$ $TO$ $NUMPY$\n* $TO$ $DO$ $1$ $:$ $DANCE$","metadata":{}},{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:bold; letter-spacing: 2px; color:#FF9980; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #FF9980\">2 | Ending 🎭</p>\n\n<div style=\"border-radius:10px; border:#FF9980 solid; padding: 15px; background-color: #F3f9ed; font-size:100%; text-align:left\">\n    \n**THAT IT FOR TODAY GUYS**\n\n**WE WILL GO DEEPER INTO THE DATA IN THE UPCOMING VERSIONS**\n\n**PLEASE COMMENT YOUR THOUGHTS, HIHGLY APPRICIATED**\n\n**DONT FORGET TO MAKE AN UPVOTE, IF YOU LIKED MY WORK $:)$**\n    \n<img src = \"https://i.imgflip.com/19aadg.jpg\">\n    \n**PEACE OUT $!!!$**","metadata":{}}]}