{"cells":[{"metadata":{},"cell_type":"markdown","source":"It's really handy to have all the DICOM info available in a single DataFrame, so let's create that! In this notebook, we'll just create the DICOM DataFrames. To see how to use them to analyze the competition data, see [this followup notebook](https://www.kaggle.com/jhoward/some-dicom-gotchas-to-be-aware-of-fastai).\n\nFirst, we'll install the latest versions of pytorch and fastai v2 (not officially released yet) so we can use the fastai medical imaging module."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"!pip install torch torchvision feather-format kornia pyarrow --upgrade   > /dev/null\n!pip install git+https://github.com/fastai/fastai_dev             > /dev/null","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"from fastai2.basics import *\nfrom fastai2.medical.imaging import *","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's take a look at what files we have in the dataset."},{"metadata":{"trusted":true},"cell_type":"code","source":"path = Path('../input/rsna-intracranial-hemorrhage-detection/')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Most lists in fastai v2, including that returned by `Path.ls`, are returned as a [fastai.core.L](http://dev.fast.ai/core.html#L), which has lots of handy methods, such as `attrgot` used here to grab file names."},{"metadata":{"trusted":true},"cell_type":"code","source":"path_trn = path/'stage_1_train_images'\nfns_trn = path_trn.ls()\nfns_trn[:5].attrgot('name')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"path_tst = path/'stage_1_test_images'\nfns_tst = path_tst.ls()\nlen(fns_trn),len(fns_tst)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can grab a file and take a look inside using the `dcmread` method that fastai v2 adds."},{"metadata":{"trusted":true},"cell_type":"code","source":"fn = fns_trn[0]\ndcm = fn.dcmread()\ndcm","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Labels"},{"metadata":{},"cell_type":"markdown","source":"Before we pull the metadata out of the DIMCOM files, let's process the labels into a convenient format and save it for later. We'll use *feather* format because it's lightning fast!"},{"metadata":{"trusted":true},"cell_type":"code","source":"def save_lbls():\n    path_lbls = path/'stage_1_train.csv'\n    lbls = pd.read_csv(path_lbls)\n    lbls[[\"ID\",\"htype\"]] = lbls.ID.str.rsplit(\"_\", n=1, expand=True)\n    lbls.drop_duplicates(['ID','htype'], inplace=True)\n    pvt = lbls.pivot('ID', 'htype', 'Label')\n    pvt.reset_index(inplace=True)    \n    pvt.to_feather('labels.fth')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"save_lbls()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_lbls = pd.read_feather('labels.fth').set_index('ID')\ndf_lbls.head(8)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_lbls.mean()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There's not much RAM on these kaggle kernel instances, so we'll clean up as we go."},{"metadata":{"trusted":true},"cell_type":"code","source":"del(df_lbls)\nimport gc; gc.collect();","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# DICOM Meta"},{"metadata":{},"cell_type":"markdown","source":"To turn the DICOM file metadata into a DataFrame we can use the `from_dicoms` function that fastai v2 adds. By passing `px_summ=True` summary statistics of the image pixels (mean/min/max/std) will be added to the DataFrame as well (although it takes much longer if you include this, since the image data has to be uncompressed)."},{"metadata":{"trusted":true},"cell_type":"code","source":"%time df_tst = pd.DataFrame.from_dicoms(fns_tst, px_summ=True)\ndf_tst.to_feather('df_tst.fth')\ndf_tst.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"del(df_tst)\ngc.collect();","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%time df_trn = pd.DataFrame.from_dicoms(fns_trn, px_summ=True)\ndf_trn.to_feather('df_trn.fth')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"There is one corrupted DICOM in the competition data, so the command above prints out the information about this file. Despite the error message show above, the command completes successfully, and the data from the corrupted file is not included in the output DataFrame."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}