{"cells":[{"metadata":{},"cell_type":"markdown","source":"The purpose of this kernel is to get a better idea of what kind of data we are working with."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport pydicom\nimport os\nimport matplotlib.pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## What are DICOMs??\nDICOMs are a format used for storing medical scanning data. They store a collection of metadata and image data. Let's have a look."},{"metadata":{"trusted":true},"cell_type":"code","source":"ds = pydicom.dcmread('../input/rsna-intracranial-hemorrhage-detection/stage_1_train_images/ID_000039fa0.dcm')\nds","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I think what we can see there that this is a CT scan (modality CT), the image stored is a 512 by 512 pixels. Not sure what the other fields mean yet.\n\nThe image is stored in the Pixel Data, let's take a look."},{"metadata":{"trusted":true},"cell_type":"code","source":"print('pixel_array:', ds.pixel_array)\nprint('center:',ds.pixel_array[206:306,206:306])\nprint('dimensions:', ds.pixel_array.shape)\nplt.imshow(ds.pixel_array, cmap=plt.cm.bone)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Yep, that's a cranium. The image is indeed a 512 by 512 image but those are some odd pixel values.\n\nTurns out they are probably [Hounsfield units](https://en.wikipedia.org/wiki/Hounsfield_scale) and it basically measures how much radiation is passing right through. Air has an HU value of -1000, bone has values between 300 and 400, or 1800 and 1900, depending on the type. Water is 0. The values we see here do not quite correspond to these because they need to be [rescaled](https://blog.kitware.com/dicom-rescale-intercept-rescale-slope-and-itk/)."},{"metadata":{"trusted":true},"cell_type":"code","source":"rescale_intercept = ds[('0028','1052')].value\nrescale_slope = ds[('0028','1053')].value\n\nrescaled_hu = rescale_slope * ds.pixel_array + rescale_intercept","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The black areas in the image have an unscaled value of -2000. these actually provide us with no information. The outer grey areas have an unscaled value of 0, and a scaled value of -1024, which is basically the value of air. We can see this transition if we zoom in close enough.\n\nWe will do this with the unscaled data as the tansition is visually clearer."},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Transition zone:')\nprint(ds.pixel_array[75:85,65:75])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can also take a look at the transition between the air and the cranium. The transition is a bit blurry, possibly due to hair."},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Air-to-cranium transition zone:')\nprint(rescaled_hu[105:115,160:170])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Honestly, the machine learning probably does not care that much about this rescaling, but it helps a bit in comprehension. The rescaled image looks the same as the unscaled one."},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.imshow(rescaled_hu, cmap=plt.cm.bone)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## What is windowing?\nAs best I can tell, windowing are some guidance values which help the viewer clarify which values to focus on. Essentially, if a value is beyond the \"window\", assign a max value to it.\n\nIn this image, we have a windowing center of 30 and a windowing width of 80, giving us an effective range of -10 to 70."},{"metadata":{"trusted":true},"cell_type":"code","source":"y_min = 0\ny_max = 255\nwindow_center = ds[('0028','1050')].value\nwindow_width = ds[('0028','1051')].value\n\nwindowed_hu = rescaled_hu.copy()\nmin_val = window_center - window_width / 2\nmax_val = window_center + window_width / 2\n\n# we want pixels with min_val to have score zero\nwindowed_hu = windowed_hu - min_val;\n\n# we want pixels with the max value to have a score of 255\nwindowed_hu = windowed_hu * y_max / max_val\n\n# we want to contrain all other values\nwindowed_hu = np.clip(windowed_hu, y_min, y_max)\n\n## have a look\nplt.imshow(windowed_hu, cmap=plt.cm.bone)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Well then, those are eyes. \n\nI'm not sure this particular windowing is very helpful. \n\nRyan Epp's [Gradient & Sigmoid Windowing](https://www.kaggle.com/reppic/gradient-sigmoid-windowing) kernel has a lot more useful forms of windowing, including using multiple windowing ranges and keeping them all as \"channels\" in an image."},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}