{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# **Pulmonary embolism**\n\nPulmonary embolism is a blockage in one of the pulmonary arteries in your lungs. In most cases, pulmonary embolism is caused by blood clots that travel to the lungs from deep veins in the legs or, rarely, from veins in other parts of the body (deep vein thrombosis).\n\nBecause the clots block blood flow to the lungs, pulmonary embolism can be life-threatening. However, prompt treatment greatly reduces the risk of death. Taking measures to prevent blood clots in your legs will help protect you against pulmonary embolism"},{"metadata":{},"cell_type":"markdown","source":"# **Symptoms**\n\nPulmonary embolism symptoms can vary greatly, depending on how much of your lung is involved, the size of the clots, and whether you have underlying lung or heart disease.\n\n\nCommon signs and symptoms include:\n\nShortness of breath. This symptom typically appears suddenly and always gets worse with exertion.\nChest pain. You may feel like you're having a heart attack. The pain is often sharp and felt when you breathe in deeply, often stopping you from being able to take a deep breath. It can also be felt when you cough, bend or stoop.\nCough. The cough may produce bloody or blood-streaked sputum.\nOther signs and symptoms that can occur with pulmonary embolism include:\n\nRapid or irregular heartbeat\nLightheadedness or dizziness\nExcessive sweating\nFever\nLeg pain or swelling, or both, usually in the calf caused by a deep vein thrombosis\nClammy or discolored skin (cyanosis)"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pydicom\nimport os\nimport matplotlib.pyplot as plt\nfrom glob import glob\nfrom mpl_toolkits.mplot3d.art3d import Poly3DCollection\nimport scipy.ndimage\nfrom skimage import morphology\nfrom skimage import measure\nfrom skimage.transform import resize\nfrom sklearn.cluster import KMeans\nfrom plotly import __version__\nfrom plotly.offline import download_plotlyjs, init_notebook_mode, plot, iplot\nimport plotly.figure_factory as ff\nfrom plotly.graph_objs import *\ninit_notebook_mode(connected=True)\nimport pandas as pd\nfrom tqdm import tqdm\nimport seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train = pd.read_csv('/kaggle/input/rsna-str-pulmonary-embolism-detection/train.csv')\ndf_test = pd.read_csv('/kaggle/input/rsna-str-pulmonary-embolism-detection/test.csv')\n\nPATH = \"../input/rsna-str-pulmonary-embolism-detection/\"\nTRAIN_PATH = PATH + \"train/\"\nTEST_PATH = PATH + \"test/\"\nsub = pd.read_csv(PATH + \"sample_submission.csv\")\ntrain_image_file_paths = glob(TRAIN_PATH + '/*/*/*.dcm')\ntest_image_file_paths = glob(TEST_PATH + '/*/*/*.dcm')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train.head(5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We see that there are 16 fields in total among them the first three are the identifiers (IDs). Here the first one is the StudyInstanceUID, SeriesInstanceUID and SOPInstanceUID are strings and the rest 13 are int64 which are essentially boolian data. Now let's have a look at the actual data itself."},{"metadata":{"trusted":true},"cell_type":"code","source":"df_test.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Data fields\n\n\n•\tStudyInstanceUID - unique ID for each study (exam) in the data.\n\n•\tSeriesInstanceUID - unique ID for each series within the study.\n\n•\tSOPInstanceUID - unique ID for each image within the study (and data).\n\n•\tpe_present_on_image - image-level, notes whether any form of PE is present on the image.\n\n•\tnegative_exam_for_pe - exam-level, whether there are any images in the study that have PE present.\n\n•\tqa_motion - informational, indicates whether radiologists noted an issue with motion in the study.\n\n•\tqa_contrast - informational, indicates whether radiologists noted an issue with contrast in the study.\n\n•\tflow_artifact - informational\n\n•\trv_lv_ratio_gte_1 - exam-level, indicates whether the RV/LV ratio present in the study is >= 1\n\n•\trv_lv_ratio_lt_1 - exam-level, indicates whether the RV/LV ratio present in the study is < 1\n\n•\tleftsided_pe - exam-level, indicates that there is PE present on the left side of the images in the study\n\n•\tchronic_pe - exam-level, indicates that the PE in the study is chronic\n\n•\ttrue_filling_defect_not_pe - informational, indicates a defect that is NOT PE\n\n•\trightsided_pe - exam-level, indicates that there is PE present on the right side of the images in the study\n\n•\tacute_and_chronic_pe - exam-level, indicates that the PE present in the study is both acute AND chronic\n\n•\tcentral_pe - exam-level, indicates that there is PE present in the center of the images in the study\n\n•\tindeterminate -exam-level, indicates that while the study is not negative for PE, an ultimate set of exam-\n\n•\tlevel labels could not be created, due to QA issues\n"},{"metadata":{},"cell_type":"markdown","source":"The DICOM files contains a lot of infomation in addition to the raw pixel values. If you want to have a in depth look inside reading dicom files, feel free to search on google."},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train.shape\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_test.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# File descriptions in the given dataset is as follows:\n\ntest - all test images directory\n\ntrain - all train images directory (note that your submission kernels will NOT have access to this set of images, so you must build your models elsewhere and incorporate them into your submissions)\n\nsample_submission.csv - contains rows for each UID+label combination that requires a prediction. Therefore it has a row for each image (for which you will be predicting the existence of a pulmonary embolism within the image) and row for each study+label that requires a study-level prediction.\n\ntrain.csv - contains UIDs and all labels.\n\ntest.csv - contains UIDs"},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_submission = pd.read_csv(\"../input/rsna-str-pulmonary-embolism-detection/sample_submission.csv\")\nsample_submission.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# **Submission size calculation for correctness:**\n\nHere among the labels given in the training data, pe_present_on_image is the image level feature that needs to be predicted for all the images.\n\n\nAnd the rest of the features will be predicted for only the exam. In that case each exam will have multiple images. But for that whole group we will submit only one set of prediction for those following labels:\n\n•\tnegative_exam_for_pe\n\n•\trv_lv_ratio_gte_1\n\n•\trv_lv_ratio_lt_1\n\n•\tleftsided_pe\n\n•\tchronic_pe\n\n•\trightsided_pe\n\n•\tacute_and_chronic_pe\n\n•\tcentral_pe\n\n•\tindeterminate\n"},{"metadata":{},"cell_type":"markdown","source":"Here for each of the image we must have to predict the property pe_present_on_image which actually indicates wherther Pulmonary Embolism (PE) is present in the image."},{"metadata":{"trusted":true},"cell_type":"code","source":"x = df_train.pe_present_on_image.value_counts()\n\nx.plot(kind='barh')\n#x.label('pe_present_on_image')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about pe_present_on_image.\nplt.pie(df_train[\"pe_present_on_image\"].value_counts(),labels=[\"0\",\"1\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of pe_present_on_image\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Draw a pie chart about negative_exam_for_pe.\nplt.pie(df_train[\"negative_exam_for_pe\"].value_counts(),labels=[\"0\",\"1\"],autopct=\"%.1f%%\")\nplt.title(\"Ratio of negative_exam_for_pe\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"x = df_train.negative_exam_for_pe.value_counts()\nprint(x)\nx.plot(kind='barh')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dcm_file","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here among these large number of parameter, there are several paramers that a radiologist must have a good understanding of. Some of those parameters are listed below.\n\nField Code\tVariable Name\tValue\n\n(0020, 0013)\tInstance Number\tIS: \"40\"\n\n(0028, 0030)\tPixel Spacing\tDS: [0.871094,0.871094]\n\n(0028, 1050)\tWindow Center\tDS: \"40.0\"\n\n(0028, 1051)\tWindow Width\tDS: \"400.0\"\n\n(0028, 1052)\tRescale Intercept\tDS: \"-1024.0\"\n\n(0028, 1053)\tRescale Slope\tDS: \"1.0\""},{"metadata":{},"cell_type":"markdown","source":"These parameters are must needed for preprocessing the CT-Scans. Without these parameters it will be very difficult to fully utilize the potential of the CT-Scans."},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(2,1,figsize=(20,10))\nfor file in train_image_file_paths[0:10]:\n    dataset = pydicom.read_file(file)\n    image = dataset.pixel_array.flatten()\n    rescaled_image = image * dataset.RescaleSlope + dataset.RescaleIntercept\n    sns.distplot(image.flatten(), ax=ax[0]);\n    sns.distplot(rescaled_image.flatten(), ax=ax[1])\nax[0].set_title(\"Raw pixel array distributions for 10 examples\");","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# View the correlation heat map\ncorr_mat = df_train.corr(method='pearson')\nsns.heatmap(corr_mat,\n            vmin=-1.0,\n            vmax=1.0,\n            center=0,\n            annot=True, # True:Displays values in a grid\n            fmt='.1f',\n            xticklabels=corr_mat.columns.values,\n            yticklabels=corr_mat.columns.values\n           )\nplt.show()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}