{"cells":[{"metadata":{},"cell_type":"markdown","source":"### what \"any\" label means?\n\nstage_1_train.csv contains data labeled with epidural, intraparenchymal,intraventricular, subarachnoid, subdural and any.\n\n\"any\" means the image includes any types of IH (intracranial hemorrhage). Therefore \"any\" should be sum or maximum of probability of 5 types of IH (less than 1) and it isn't needed for training our model.\n\nRather than that, when we train our model to predict whether an image includes IH(and what type is it), we need \"normal\" label which means the image doesn't includes any types of IH. "},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"INPUT_PATH = \"../input/rsna-intracranial-hemorrhage-detection/\"\ntrain_df = pd.read_csv(INPUT_PATH + \"stage_1_train.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"label = train_df.Label.values\ntrain_df = train_df.ID.str.rsplit(\"_\", n=1, expand=True)\ntrain_df.loc[:, \"label\"] = label","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = train_df.rename({0: \"id\", 1: \"subtype\"}, axis=1)\ntrain_df.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_pivot_df = pd.pivot_table(train_df, index=\"id\", columns=\"subtype\", values=\"label\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_pivot_df.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_pivot_df[\"normal\"] = 1 - train_pivot_df[\"any\"]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_pivot_df.head()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}