{"cells":[{"metadata":{},"cell_type":"markdown","source":"###       Reading the Attributes from Dicom File\n\nIn the VinBigData Chect x-ray Competition, we are dealing with dicom files for the x-ray chest images. The train.cvs does not contain all the patient data that might be important for the localization and detction of the abnormalities.\n\nIn the notebook you will find a code for reading some of the immportant patient attributes "},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\n#for dirname, _, filenames in os.walk('/kaggle/input'):\n#    for filename in filenames:\n#        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"import matplotlib.pyplot as plt\n%matplotlib inline\nimport pydicom\nimport warnings\nwarnings.filterwarnings(\"ignore\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#\n# The path to the dataset\n#\nDataDir = \"../input/vinbigdata-chest-xray-abnormalities-detection/\"\n!ls {DataDir}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#\n# Reading the train.cvs data\n#\ntrain = pd.read_csv(DataDir+'train.csv')\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#\n# Finding the keywords for accessing the data elements in a dicom file.\n#\ndcm_file = pydicom.dcmread(DataDir+ 'train/'+train['image_id'][2]+'.dicom')\ndcm_file.dir()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#\n# Here another way for accessing the image data as a numpy array\n#\ndcm_pixels = dcm_file.pixel_array\ndcm_pixels","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#\n# Displaying the image from the pixel array\n#\nplt.figure(figsize=(12,10))\nplt.imshow(dcm_pixels, cmap=plt.cm.gray)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#\n# Here is the function for reading the patients' attributes \n# from the dicom images.\n# \ndef get_dcm_attributes(path):\n\n    df = pd.DataFrame(columns=['image_id', 'Age', 'Gender','Image_Hieght',\n                    'ImageWidth','x_spacing','y_spacing'])\n    #Read some files for testing\n    files = list(os.listdir(path))[0:10]\n    #Read All files\n    #files = list(os.listdir(path))\n   \n    try:\n        i = 0\n        for file in files:\n\n            file_path = os.path.join(path,file)\n            dcmData = pydicom.dcmread(file_path,stop_before_pixels=True)\n\n            file_name = file.split(\".\")[0]\n\n            attributes = dcmData.dir()\n            if 'PatientAge' in attributes:\n                age_str = dcmData.PatientAge\n                if age_str != '' and age_str != 'Y':\n                    age = int(age_str[:-1])\n                else:\n                    age = np.NaN\n            else:\n                age = np.NaN\n            if 'PatientSex' in attributes:\n                gender = dcmData.PatientSex\n                if gender =='' : gender = np.NaN\n            else:\n                gender = np.NaN\n            if 'Rows' in attributes:\n                rows = dcmData.Rows\n            else:\n                rows = np.NaN\n            if 'Columns' in attributes:\n                clmns = dcmData.Columns\n            else:\n                clmns = np.NaN\n            if 'PixelSpacing' in attributes:\n                ps = dcmData.PixelSpacing\n            else:\n                ps = [np.NaN,np.NaN]\n\n            df = df.append(pd.DataFrame({'image_id': file_name, \n                    'Age': age, 'Gender': gender,'Image_Hieght': rows,\n                    'ImageWidth': clmns,\n                    'x_spacing': ps[0],'y_spacing': ps[1]}, index=[i]))\n            i+=1\n    except ValueError:\n            print('age_str',\"   \", age_str)\n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#\n# Reading some image attributes. (it takes several minutes for the whole dataset)\n#\nTrainDir = DataDir+'train/'\ndcm_attr = get_dcm_attributes(TrainDir)\ndcm_attr.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"np.sum(dcm_attr.isna())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#\n# Now Join this info with the data in the train.cvs\n#\ntrain_mrg = pd.merge(train, dcm_attr, on = 'image_id')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_mrg.head(20)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Don't forget to upvote ^_^"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}