{"cells":[{"metadata":{},"cell_type":"markdown","source":""},{"metadata":{},"cell_type":"markdown","source":"# RSNA: the worst cases\nMost images in this dataset are normal.  \nWhen they are not, one hemorrhage type is detected most of the time.  \nBut sometimes there are multiple hemorrhages or even all of them.  \nI have counted those cases, and looked at them.  "},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"from glob import glob\nimport os\nimport pandas as pd\nimport numpy as np\nimport re\nfrom PIL import Image\nimport seaborn as sns\nimport pydicom\nimport matplotlib.pyplot as plt\nfrom sklearn.model_selection import train_test_split\nfrom tqdm import tqdm_notebook as tqdm\nimport cv2\n\n#checnking the input files\n# print(os.listdir(\"../input/rsna-intracranial-hemorrhage-detection/\"))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Load data"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# Load datasets\ntrain_images_dir = '../input/rsna-intracranial-hemorrhage-detection/stage_1_train_images/'\ntest_images_dir = '../input/rsna-intracranial-hemorrhage-detection/stage_1_test_images/'\n\ntrain = pd.read_csv('../input/rsna-intracranial-hemorrhage-detection/stage_1_train.csv')\ntest = pd.read_csv('../input/rsna-intracranial-hemorrhage-detection/stage_1_sample_submission.csv')\n\n# Transform training set. Code from https://www.kaggle.com/taindow/pytorch-efficientnet-b0-benchmark\ntrain[['ID', 'Image', 'Diagnosis']] = train['ID'].str.split('_', expand=True)\ntrain = train[['Image', 'Diagnosis', 'Label']]\ntrain.drop_duplicates(inplace=True)\ntrain = train.pivot(index='Image', columns='Diagnosis', values='Label').reset_index()\ntrain['Image'] = 'ID_' + train['Image']\ntrain.head(10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Look at some normal cases\nHere are multiple images where `any == 0`."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"def _get_first_of_dicom_field_as_int(x):\n    #get x[0] as in int is x is a 'pydicom.multival.MultiValue', otherwise get int(x)\n    if type(x) == pydicom.multival.MultiValue:\n        return int(x[0])\n    else:\n        return int(x)\n\ndef _get_windowing(data):\n    dicom_fields = [data[('0028','1050')].value, #window center\n                    data[('0028','1051')].value, #window width\n                    data[('0028','1052')].value, #intercept\n                    data[('0028','1053')].value] #slope\n    return [_get_first_of_dicom_field_as_int(x) for x in dicom_fields]\n\ndef get_image(data, windowing=None):\n    window_center, window_width, intercept, slope = windowing or _get_windowing(data)\n    img = data.pixel_array\n    img = (img*slope +intercept)\n    img_min = window_center - window_width//2\n    img_max = window_center + window_width//2\n    img[img<img_min] = img_min\n    img[img>img_max] = img_max\n    return img","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(nrows=4, ncols=4, figsize=(30,30))\n\nfor i in range(16):\n    idx = train[train['any'] == 0]['Image'].iloc[i]\n    data = pydicom.dcmread(train_images_dir + idx + '.dcm')\n    img = get_image(data)\n    ax[i//4, i%4].set_title(idx)\n    ax[i//4, i%4].imshow(img, cmap=plt.cm.bone)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Look at the bad cases\nLet's count images in which there are multiple hemorrhages detected."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"for n in range(6):\n    many = train[train[['epidural', 'intraparenchymal', 'intraventricular', 'subarachnoid', 'subdural']].sum(1) == n].copy()\n    print('Number of hemorrhages: {}, amount of such images: {}, fraction: {:.3f}%'.format(n, len(many), 100 * len(many) / len(train)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Most images (86%) show normal brains.  \n10% of images have only one type of hemorrhage.  \nThere are **20 samples** having all the hemorrhage types at the same time:"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"many","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's look closer at these samples: extract `Patient ID` and see if they belong to same patient"},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"many['Patient ID'] = many['Image'].map(lambda image: pydicom.dcmread(train_images_dir + image + '.dcm')[('0010', '0020')].repval.replace(\"'\", ''))\nmany","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"many['Patient ID'].unique()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are 8 unique patients to whom those 20 images belong.  \n\n## Print bad scans\n\nLet's look at each of them."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"log = []\n\nfor i, patient in enumerate(many['Patient ID'].unique()):\n    print('Patient №{}: {}'.format(i + 1, patient))\n    ids = many[many['Patient ID'] == patient]['Image']\n    fig, ax = plt.subplots(nrows=1, ncols=len(ids), figsize=(30,10))\n    for j, idx in enumerate(ids):\n        data = pydicom.dcmread(train_images_dir + idx + '.dcm')\n        log.append(('Patient №{}: {}'.format(i + 1, patient), idx, data))\n        img = get_image(data)\n        if len(ids) == 1:\n            ax.set_title(idx, fontsize=40)\n            ax.imshow(img, cmap=plt.cm.bone)\n        else:\n            ax[j].set_title(idx, fontsize=30)\n            ax[j].imshow(img, cmap=plt.cm.bone)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I am not a doctor, so I am not exactly sure what do I see.  \nFor patients №2,4,5,7,8 I suppose there are some serious skull injuries; no surprise that all types of hemorrhage were detected...  \n\nI am not sure about other ones, especially №3.  \n№3 looks asymetric a bit, but maybe it's just angle of view.  \nI cannot see any obvious defects there.   \nMaybe the windowing in not correct?  \n\n# Windowing\nLet's try drawing this image with different windowings.  \nI used windows from kernel https://www.kaggle.com/dcstang/see-like-a-radiologist-with-systematic-windowing"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# Get id of image of patient №3\nidx = many[many['Patient ID'] == many['Patient ID'].unique()[2]]['Image'].iloc[0]\ndata = pydicom.dcmread(train_images_dir + idx + '.dcm')\nc, w, intercept, slope = _get_windowing(data)\n\nknown_windows = [('Default window', c, w),\n           ('Brain Matter window', 40, 80),\n           ('Blood/subdural window', 80, 200),\n           ('Soft tissue window', 40, 375),\n           ('Bone window', 600, 2800),\n           ('Grey-white differentiation window', 32, 8)]\nfig, ax = plt.subplots(nrows=2, ncols=3, figsize=(30,20))\n\nfor i, (window_name, window_center, window_width) in enumerate(known_windows):\n    img = get_image(data, [window_center, window_width, intercept, slope])\n    ax[i//3, i%3].set_title(window_name, fontsize=40)\n    ax[i//3, i%3].imshow(img, cmap=plt.cm.bone)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now I can see, using the bone window, that there is probably a crack at the top left corner.  \n\nLet's look at all those 20 images with all windows that I know:"},{"metadata":{"trusted":true},"cell_type":"code","source":"for patient, idx, data in log:\n    print(patient, 'Image', idx)\n    c, w, intercept, slope = _get_windowing(data)\n\n    fig, ax = plt.subplots(nrows=2, ncols=3, figsize=(30,20))\n\n    for i, (window_name, window_center, window_width) in enumerate(known_windows):\n        img = get_image(data, [window_center, window_width, intercept, slope])\n        ax[i//3, i%3].set_title(window_name, fontsize=40)\n        ax[i//3, i%3].imshow(img, cmap=plt.cm.bone)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It was not clear what happens at image №1. Now, through the bone window I can clearly see what is wrong.  \nAll the patients in this set, except №6, have clearly visible cracks.  \n\nI am not sure what exactly is wrong with patient №6.  \nBut, thanks to https://www.kaggle.com/dcstang/see-like-a-radiologist-with-systematic-windowing, I can see with my own eyes that there is not only something big outside the scull on the left, but also a big \"bad\" spot in `Soft tissue window` inside the brain. I couldn't see this using `Default window`.  \n\n# Conclusion\nThose pictures look interesting but also scary...\n\nNow I know that:\n* There are actually a lot of samples when it is 2 or more types of hemorrhage\n* I probably want to use multiple windows to feed multiple inputs to a neural network. If I can't see anything on the default windowed image, why should model? :)\n\n### I wish health to everyone, and good luck with this challenge!\nIf you have something to add, especially from medical point, please leave comments :)"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}