{"cells":[{"metadata":{},"cell_type":"markdown","source":"### InceptionV3 and EfficientNetB0 \n\nThis notebook takes you through some important steps in building a deep convnet in Keras for multilabel classification of brain CT scans. \n\n\n*Update (1):*\n* *training for 4 epochs instead of 3.*\n* *batch size lowered to 16 from 32.*\n* *training without learning rate decay.*\n* *Weighted BCE instead of \"plain\" BCE*\n* *training data lowered to 80% from 90%.*\n\n\n*Update (2):*\n* *adding competition metric for training*\n* *using custom Callback for validation and test sets instead of the `run()` function and 'global epochs'*\n* *training with \"plain\" BCE again*\n* *merging TestDataGenerator and TrainDataGenerator into one*\n* *adding undersampling (see inside `on_epoch_end`), will now run 6 epochs*\n\n*Update (3):*\n* *skipping/removing windowing (value clipping), but the transformation to Hounsfield Units is kept*\n* *removing initial layer (doing np.stack((img,)&ast;3, axis=-1)) instead*\n* *reducing learning rate to 5e-4 and add decay*\n* *increasing batch size to 32 from 16*\n* *Increasing training set to 90% of the data (10% for validation)*\n* *slight increase in undersampling*\n* *fixed some hardcoding for input dims/sizes*\n* *training with weighted BCE again*\n\n*Update (4):*\n* *Trying out InceptionV3, instead of ResNet50*\n* *undersampling without weights*\n* *adding dense layers with dropout before output*\n* *clipping HUs between -50 and 450 (probably the most relevant value-space?)*\n* *normalization is now mapping input to 0 to 1 range, instead of -1 to 1.*\n* *doing 5 epochs instead of 6*\n\n*Update (5):*\n* *Got some inspiration from [this great](https://www.kaggle.com/reppic/gradient-sigmoid-windowing) kernel by Ryan Epp*\n* *Thus I'm trying out the sigmoid (brain + subdural + bone) to see if it improves the log loss*\n* *Number of epochs reduced to 4, increased undersampling, and validation predictions removed due to limited time*\n\n*Update (6) (did not improve from (5)):*\n* *Going back to raw HUs (with a bit of clipping)*\n* *Together with a first initial conv layer with sigmoid activation*\n* *epochs increased to 6 from 4, and input size increased to (256, 256) from (224, 224)*\n* *simple average of epochs (>1) for the test predictions*\n\n*Update (7):*\n* *Trying windowing based on [appian42's repo](https://github.com/appian42/kaggle-rsna-intracranial-hemorrhage/) instead*\n* *I also include some cleaning based on [Jeremy's kernel](https://www.kaggle.com/jhoward/cleaning-the-data-for-rapid-prototyping-fastai) (hopefully it's correct, atleast the visualization looked good =)*\n* *weighted average of the epochs (>1) for the test predictions*\n* *reducing number of epochs to 5*\n\n*Update (8):*\n* *Removing the extra dense layer before output layer (keeping everything else the same)*"},{"metadata":{},"cell_type":"markdown","source":"\n    EfficientNetB0 - (224, 224, 3)\n    EfficientNetB1 - (240, 240, 3)\n    EfficientNetB2 - (260, 260, 3)\n    EfficientNetB3 - (300, 300, 3)\n    EfficientNetB4 - (380, 380, 3)\n    EfficientNetB5 - (456, 456, 3)\n    EfficientNetB6 - (528, 528, 3)\n    EfficientNetB7 - (600, 600, 3)\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Install Modules from internet\n!pip install efficientnet\n!pip install iterative-stratification","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport pydicom\nimport os\nimport matplotlib.pyplot as plt\nimport collections\nfrom tqdm import tqdm_notebook as tqdm\nfrom datetime import datetime\n\nfrom math import ceil, floor, log\nimport cv2\n\nimport tensorflow as tf\nimport keras\n\nimport sys\n\nimport efficientnet.keras as efn\n#sys.path.append(os.path.abspath('../input/efficientnet/efficientnet-master/efficientnet-master/'))\n#from efficientnet import EfficientNetB0\n\n# from keras_applications.resnet import ResNet50\nfrom keras_applications.inception_v3 import InceptionV3\n\nfrom sklearn.model_selection import ShuffleSplit\n\ntest_images_dir = '../input/rsna-intracranial-hemorrhage-detection/rsna-intracranial-hemorrhage-detection/stage_2_test/'\ntrain_images_dir = '../input/rsna-intracranial-hemorrhage-detection/rsna-intracranial-hemorrhage-detection/stage_2_train/'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"os.listdir('../input/rsna-models/stage2')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 0. Preprocessing (brain + subudral + soft)\n(REMOVED) Many thanks to [Ryan Epp](https://www.kaggle.com/reppic/gradient-sigmoid-windowing). Code is taken from his kernel (see his kernel for more information and other peoples work --- for example [David Tang](https://www.kaggle.com/dcstang/see-like-a-radiologist-with-systematic-windowing), [Marco](https://www.kaggle.com/marcovasquez/basic-eda-data-visualization), [Nanashi](https://www.kaggle.com/jesucristo/rsna-introduction-eda-models), and [Richard McKinley](https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing)). At first I thought I couldn't use sigmoid windowing for this kernel because of how expensive it is to do, but I could resize the image prior to the transformation to save a lot of computation. Not sure how much this will affect the performance of the training, but it really speeded it up.<br>\n(NEW) Based on two great kernels: [appian42's repo](https://github.com/appian42/kaggle-rsna-intracranial-hemorrhage/) (windowing), [Jeremy's kernel](https://www.kaggle.com/jhoward/cleaning-the-data-for-rapid-prototyping-fastai) (cleaning)"},{"metadata":{"trusted":true},"cell_type":"code","source":"os.listdir('../input/rsna-intracranial-hemorrhage-detection/rsna-intracranial-hemorrhage-detection/')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def correct_dcm(dcm):\n    x = dcm.pixel_array + 1000\n    px_mode = 4096\n    x[x>=px_mode] = x[x>=px_mode] - px_mode\n    dcm.PixelData = x.tobytes()\n    dcm.RescaleIntercept = -1000\n\ndef window_image(dcm, window_center, window_width):\n    \n    if (dcm.BitsStored == 12) and (dcm.PixelRepresentation == 0) and (int(dcm.RescaleIntercept) > -100):\n        correct_dcm(dcm)\n    \n    img = dcm.pixel_array * dcm.RescaleSlope + dcm.RescaleIntercept\n    img_min = window_center - window_width // 2\n    img_max = window_center + window_width // 2\n    img = np.clip(img, img_min, img_max)\n\n    return img\n\ndef bsb_window(dcm):\n    brain_img = window_image(dcm, 40, 80)\n    subdural_img = window_image(dcm, 80, 200)\n    soft_img = window_image(dcm, 40, 380)\n    \n    brain_img = (brain_img - 0) / 80\n    subdural_img = (subdural_img - (-20)) / 200\n    soft_img = (soft_img - (-150)) / 380\n    bsb_img = np.array([brain_img, subdural_img, soft_img]).transpose(1,2,0)\n\n    return bsb_img\n\n# Sanity Check\n# Example dicoms: ID_2669954a7, ID_5c8b5d701, ID_52c9913b1\n\ndicom = pydicom.dcmread(train_images_dir + 'ID_e1b0c40f8' + '.dcm')\n#                                     ID  Label\n# 4045566          ID_5c8b5d701_epidural      0\n# 4045567  ID_5c8b5d701_intraparenchymal      1\n# 4045568  ID_5c8b5d701_intraventricular      0\n# 4045569      ID_5c8b5d701_subarachnoid      1\n# 4045570          ID_5c8b5d701_subdural      1\n# 4045571               ID_5c8b5d701_any      1\nplt.imshow(bsb_window(dicom), cmap=plt.cm.bone);\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Check (with an example) if the correction works (visually)"},{"metadata":{"trusted":true},"cell_type":"code","source":"def window_with_correction(dcm, window_center, window_width):\n    if (dcm.BitsStored == 12) and (dcm.PixelRepresentation == 0) and (int(dcm.RescaleIntercept) > -100):\n        correct_dcm(dcm)\n    img = dcm.pixel_array * dcm.RescaleSlope + dcm.RescaleIntercept\n    img_min = window_center - window_width // 2\n    img_max = window_center + window_width // 2\n    img = np.clip(img, img_min, img_max)\n    return img\n\ndef window_without_correction(dcm, window_center, window_width):\n    img = dcm.pixel_array * dcm.RescaleSlope + dcm.RescaleIntercept\n    img_min = window_center - window_width // 2\n    img_max = window_center + window_width // 2\n    img = np.clip(img, img_min, img_max)\n    return img\n\ndef window_testing(img, window):\n    brain_img = window(img, 40, 80)\n    subdural_img = window(img, 80, 200)\n    soft_img = window(img, 40, 380)\n    \n    brain_img = (brain_img - 0) / 80\n    subdural_img = (subdural_img - (-20)) / 200\n    soft_img = (soft_img - (-150)) / 380\n    bsb_img = np.array([brain_img, subdural_img, soft_img]).transpose(1,2,0)\n\n    return bsb_img\n\n# example of a \"bad data point\" (i.e. (dcm.BitsStored == 12) and (dcm.PixelRepresentation == 0) and (int(dcm.RescaleIntercept) > -100) == True)\ndicom = pydicom.dcmread(train_images_dir + \"ID_036db39b7\" + \".dcm\")\n\nfig, ax = plt.subplots(1, 2)\n\nax[0].imshow(window_testing(dicom, window_without_correction), cmap=plt.cm.bone);\nax[0].set_title(\"original\")\nax[1].imshow(window_testing(dicom, window_with_correction), cmap=plt.cm.bone);\nax[1].set_title(\"corrected\");","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 1. Helper functions\n\n* read and transform dcms to 3-channel inputs for e.g. InceptionV3. \n* uses `bsb_window` from previous cell\n\n\\* (REMOVED) Source for windowing (although now partly removed from this kernel): https://www.kaggle.com/omission/eda-view-dicom-images-with-correct-windowing"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"def _read(path, desired_size):\n    \"\"\"Will be used in DataGenerator\"\"\"\n    \n    dcm = pydicom.dcmread(path)\n    \n    try:\n        img = bsb_window(dcm)\n    except:\n        img = np.zeros(desired_size)\n    \n    \n    img = cv2.resize(img, desired_size[:2], interpolation=cv2.INTER_LINEAR)\n    \n    return img\n\n# Another sanity check \nplt.imshow(\n    _read(train_images_dir+'ID_5c8b5d701'+'.dcm', (256, 256)), cmap=plt.cm.bone\n);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 2. Data generators\n\nInherits from keras.utils.Sequence object and thus should be safe for multiprocessing.\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"class DataGenerator(keras.utils.Sequence):\n\n    def __init__(self, list_IDs, labels=None, batch_size=1, img_size=(512, 512, 1), \n                 img_dir=train_images_dir, *args, **kwargs):\n\n        self.list_IDs = list_IDs\n        self.labels = labels\n        self.batch_size = batch_size\n        self.img_size = img_size\n        self.img_dir = img_dir\n        self.on_epoch_end()\n\n    def __len__(self):\n        return int(ceil(len(self.indices) / self.batch_size))\n\n    def __getitem__(self, index):\n        indices = self.indices[index*self.batch_size:(index+1)*self.batch_size]\n        list_IDs_temp = [self.list_IDs[k] for k in indices]\n        \n        if self.labels is not None:\n            X, Y = self.__data_generation(list_IDs_temp)\n            return X, Y\n        else:\n            X = self.__data_generation(list_IDs_temp)\n            return X\n        \n    def on_epoch_end(self):\n        if self.labels is not None: # for training phase we undersample and shuffle\n            # keep probability of any=0 and any=1\n            keep_prob = self.labels.iloc[:, 0].map({0: 0.35, 1: 0.5})\n            keep = (keep_prob > np.random.rand(len(keep_prob)))\n            self.indices = np.arange(len(self.list_IDs))[keep]\n            np.random.shuffle(self.indices)\n        else:\n            self.indices = np.arange(len(self.list_IDs))\n\n    def __data_generation(self, list_IDs_temp):\n        X = np.empty((self.batch_size, *self.img_size))\n        \n        if self.labels is not None: # training phase\n            Y = np.empty((self.batch_size, 6), dtype=np.float32)\n        \n            for i, ID in enumerate(list_IDs_temp):\n                X[i,] = _read(self.img_dir+ID+\".dcm\", self.img_size)\n                Y[i,] = self.labels.loc[ID].values\n        \n            return X, Y\n        \n        else: # test phase\n            for i, ID in enumerate(list_IDs_temp):\n                X[i,] = _read(self.img_dir+ID+\".dcm\", self.img_size)\n            \n            return X","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 3a. loss function, metric & Optimizer"},{"metadata":{"trusted":true},"cell_type":"code","source":"from keras import backend as K\n\ndef weighted_log_loss(y_true, y_pred):\n    \"\"\"\n    Can be used as the loss function in model.compile()\n    ---------------------------------------------------\n    \"\"\"\n    \n    class_weights = np.array([2., 1., 1., 1., 1., 1.])\n    \n    eps = K.epsilon()\n    \n    y_pred = K.clip(y_pred, eps, 1.0-eps)\n\n    out = -(         y_true  * K.log(      y_pred) * class_weights\n            + (1.0 - y_true) * K.log(1.0 - y_pred) * class_weights)\n    \n    return K.mean(out, axis=-1)\n\n\ndef _normalized_weighted_average(arr, weights=None):\n    \"\"\"\n    A simple Keras implementation that mimics that of \n    numpy.average(), specifically for this competition\n    \"\"\"\n    \n    if weights is not None:\n        scl = K.sum(weights)\n        weights = K.expand_dims(weights, axis=1)\n        return K.sum(K.dot(arr, weights), axis=1) / scl\n    return K.mean(arr, axis=1)\n\n\ndef weighted_loss(y_true, y_pred):\n    \"\"\"\n    Will be used as the metric in model.compile()\n    ---------------------------------------------\n    \n    Similar to the custom loss function 'weighted_log_loss()' above\n    but with normalized weights, which should be very similar \n    to the official competition metric:\n        https://www.kaggle.com/kambarakun/lb-probe-weights-n-of-positives-scoring\n    and hence:\n        sklearn.metrics.log_loss with sample weights\n    \"\"\"\n    \n    class_weights = K.variable([2., 1., 1., 1., 1., 1.])\n    \n    eps = K.epsilon()\n    \n    y_pred = K.clip(y_pred, eps, 1.0-eps)\n\n    loss = -(        y_true  * K.log(      y_pred)\n            + (1.0 - y_true) * K.log(1.0 - y_pred))\n    \n    loss_samples = _normalized_weighted_average(loss, class_weights)\n    \n    return K.mean(loss_samples)\n\n\ndef weighted_log_loss_metric(trues, preds):\n    \"\"\"\n    Will be used to calculate the log loss \n    of the validation set in PredictionCheckpoint()\n    ------------------------------------------\n    \"\"\"\n    class_weights = [2., 1., 1., 1., 1., 1.]\n    \n    epsilon = 1e-7\n    \n    preds = np.clip(preds, epsilon, 1-epsilon)\n    loss = trues * np.log(preds) + (1 - trues) * np.log(1 - preds)\n    loss_samples = np.average(loss, axis=1, weights=class_weights)\n\n    return - loss_samples.mean()\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Radam Optimizer\nfrom keras.optimizers import Optimizer\n\n# Ported from https://github.com/LiyuanLucasLiu/RAdam/blob/master/radam.py\nclass RectifiedAdam(Optimizer):\n    \"\"\"RectifiedAdam optimizer.\n    Default parameters follow those provided in the original paper.\n    # Arguments\n        lr: float >= 0. Learning rate.\n        final_lr: float >= 0. Final learning rate.\n        beta_1: float, 0 < beta < 1. Generally close to 1.\n        beta_2: float, 0 < beta < 1. Generally close to 1.\n        gamma: float >= 0. Convergence speed of the bound function.\n        epsilon: float >= 0. Fuzz factor. If `None`, defaults to `K.epsilon()`.\n        decay: float >= 0. Learning rate decay over each update.\n        weight_decay: Weight decay weight.\n        amsbound: boolean. Whether to apply the AMSBound variant of this\n            algorithm.\n    # References\n        - [On the Variance of the Adaptive Learning Rate and Beyond]\n          (https://arxiv.org/abs/1908.03265)\n        - [Adam - A Method for Stochastic Optimization]\n          (https://arxiv.org/abs/1412.6980v8)\n        - [On the Convergence of Adam and Beyond]\n          (https://openreview.net/forum?id=ryQu7f-RZ)\n    \"\"\"\n\n    def __init__(self, lr=0.001, beta_1=0.9, beta_2=0.999,\n                 epsilon=None, decay=0., weight_decay=0.0, **kwargs):\n        super(RectifiedAdam, self).__init__(**kwargs)\n\n        with K.name_scope(self.__class__.__name__):\n            self.iterations = K.variable(0, dtype='int64', name='iterations')\n            self.learning_rate = K.variable(lr, name='lr')\n            self.beta_1 = K.variable(beta_1, name='beta_1')\n            self.beta_2 = K.variable(beta_2, name='beta_2')\n            self.decay = K.variable(decay, name='decay')\n\n        if epsilon is None:\n            epsilon = K.epsilon()\n        self.epsilon = epsilon\n        self.initial_decay = decay\n\n        self.weight_decay = float(weight_decay)\n\n    def get_updates(self, loss, params):\n        grads = self.get_gradients(loss, params)\n        self.updates = [K.update_add(self.iterations, 1)]\n\n        lr = self.learning_rate\n        if self.initial_decay > 0:\n            lr = lr * (1. / (1. + self.decay * K.cast(self.iterations,\n                                                      K.dtype(self.decay))))\n\n        t = K.cast(self.iterations, K.floatx()) + 1\n\n        ms = [K.zeros(K.int_shape(p), dtype=K.dtype(p)) for p in params]\n        vs = [K.zeros(K.int_shape(p), dtype=K.dtype(p)) for p in params]\n        self.weights = [self.iterations] + ms + vs\n\n        for p, g, m, v in zip(params, grads, ms, vs):\n            m_t = (self.beta_1 * m) + (1. - self.beta_1) * g\n            v_t = (self.beta_2 * v) + (1. - self.beta_2) * K.square(g)\n\n            beta2_t = self.beta_2 ** t\n            N_sma_max = 2 / (1 - self.beta_2) - 1\n            N_sma = N_sma_max - 2 * t * beta2_t / (1 - beta2_t)\n\n            # apply weight decay\n            if self.weight_decay != 0.:\n                p_wd = p - self.weight_decay * lr * p\n            else:\n                p_wd = None\n\n            if p_wd is None:\n                p_ = p\n            else:\n                p_ = p_wd\n\n            def gt_path():\n                step_size = lr * K.sqrt(\n                    (1 - beta2_t) * (N_sma - 4) / (N_sma_max - 4) * (N_sma - 2) / N_sma * N_sma_max /\n                    (N_sma_max - 2)) / (1 - self.beta_1 ** t)\n\n                denom = K.sqrt(v_t) + self.epsilon\n                p_t = p_ - step_size * (m_t / denom)\n\n                return p_t\n\n            def lt_path():\n                step_size = lr / (1 - self.beta_1 ** t)\n                p_t = p_ - step_size * m_t\n\n                return p_t\n\n            p_t = K.switch(N_sma > 5, gt_path, lt_path)\n\n            self.updates.append(K.update(m, m_t))\n            self.updates.append(K.update(v, v_t))\n            new_p = p_t\n\n            # Apply constraints.\n            if getattr(p, 'constraint', None) is not None:\n                new_p = p.constraint(new_p)\n\n            self.updates.append(K.update(p, new_p))\n        return self.updates\n\n    def get_config(self):\n        config = {'lr': float(K.get_value(self.learning_rate)),\n                  'beta_1': float(K.get_value(self.beta_1)),\n                  'beta_2': float(K.get_value(self.beta_2)),\n                  'decay': float(K.get_value(self.decay)),\n                  'epsilon': self.epsilon,\n                  'weight_decay': self.weight_decay}\n        base_config = super(RectifiedAdam, self).get_config()\n        return dict(list(base_config.items()) + list(config.items()))\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 3b. Model\n\nModel is divided into three parts: <br> \n\n* (REMOVED) The initial layer, which will transform/map input image of shape (\\_, \\_, 1) to another \"image\" of shape (\\_, \\_, 3).\n\n* The new input image is then passed through InceptionV3 (which I named \"engine\"). InceptionV3 could be replaced by any of the available architectures in keras_application.\n\n* Finally, the output from InceptionV3 goes through average pooling followed by two dense layers (including output layer)."},{"metadata":{"trusted":true},"cell_type":"code","source":"\nclass PredictionCheckpoint(keras.callbacks.Callback):\n    \n    def __init__(self, test_df, valid_df, \n                 test_images_dir=test_images_dir, \n                 valid_images_dir=train_images_dir, \n                 batch_size=32, input_size=(224, 224, 3)):\n        \n        self.test_df = test_df\n        self.valid_df = valid_df\n        self.test_images_dir = test_images_dir\n        self.valid_images_dir = valid_images_dir\n        self.batch_size = batch_size\n        self.input_size = input_size\n        \n    def on_train_begin(self, logs={}):\n        self.test_predictions = []\n        self.valid_predictions = []\n        \n    def on_epoch_end(self,batch, logs={}):\n        self.test_predictions.append(\n            self.model.predict_generator(\n                DataGenerator(self.test_df.index, \n                              None, \n                              self.batch_size, \n                              self.input_size, \n                              self.test_images_dir), verbose=2)[:len(self.test_df)])\n\n\n\nclass MyDeepModel:\n    \n    def __init__(self, engine, input_dims, batch_size=5, num_epochs=4, learning_rate=1e-3, \n                 decay_rate=1.0, decay_steps=1, weights=\"imagenet\", verbose=1, predefined = False):\n        \n        self.engine = engine\n        self.input_dims = input_dims\n        self.batch_size = batch_size\n        self.num_epochs = num_epochs\n        self.learning_rate = learning_rate\n        self.decay_rate = decay_rate\n        self.decay_steps = decay_steps\n        self.weights = weights\n        self.verbose = verbose\n        self.predefined = predefined\n        self._build()\n\n    def _build(self):\n        \n        if self.predefined:\n            engine = self.engine\n            \n        else:\n            engine = self.engine(include_top=False, weights=self.weights, input_shape=self.input_dims)\n        \n        x = keras.layers.GlobalAveragePooling2D(name='avg_pool')(engine.output)\n        out = keras.layers.Dense(6, activation=\"sigmoid\", name='dense_output')(x)\n\n        self.model = keras.models.Model(inputs=engine.input, outputs=out)\n\n        self.model.compile(loss=\"binary_crossentropy\", optimizer=RectifiedAdam(), metrics=[weighted_loss])\n    \n\n    def fit_and_predict(self, train_df, valid_df, test_df):\n        \n        # callbacks\n        pred_history = PredictionCheckpoint(test_df, valid_df, input_size=self.input_dims)\n        #checkpointer = keras.callbacks.ModelCheckpoint(filepath='%s-{epoch:02d}.hdf5' % self.engine.__name__, verbose=1, save_weights_only=True, save_best_only=False)\n        scheduler = keras.callbacks.LearningRateScheduler(lambda epoch: self.learning_rate * pow(self.decay_rate, floor(epoch / self.decay_steps)))\n        \n        self.model.fit_generator(\n            DataGenerator(\n                train_df.index, \n                train_df, \n                self.batch_size, \n                self.input_dims, \n                train_images_dir\n            ),\n            epochs=self.num_epochs,\n            verbose=self.verbose,\n            use_multiprocessing=True,\n            workers=4,\n            callbacks=[pred_history, scheduler]\n        )\n        \n        return pred_history\n    \n    def predict(self, test_df, test_image_dir):\n        y_pred = self.model.predict_generator(\n                                  DataGenerator(test_df.index, \n                                      None, \n                                      self.batch_size, \n                                      self.input_dims, \n                                      test_images_dir), verbose=2)[:len(test_df)]\n        \n        return y_pred\n    \n    def save(self, path):\n        self.model.save_weights(path)\n    \n    def load(self, path):\n        self.model.load_weights(path)\n        \n    def summary(self):\n        self.model.summary()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 4. Read csv files\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"def read_testset(filename=\"../input/rsna-intracranial-hemorrhage-detection/rsna-intracranial-hemorrhage-detection/stage_2_sample_submission.csv\"):\n    df = pd.read_csv(filename)\n    df[\"Image\"] = df[\"ID\"].str.slice(stop=12)\n    df[\"Diagnosis\"] = df[\"ID\"].str.slice(start=13)\n    \n    df = df.loc[:, [\"Label\", \"Diagnosis\", \"Image\"]]\n    df = df.set_index(['Image', 'Diagnosis']).unstack(level=-1)\n    \n    return df\n    \ntest_df = read_testset()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df.head(3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 5. Train model and predict\n\n*Using train, validation and test set* <br>\n\nTraining for 5 epochs with Adam optimizer, with a learning rate of 0.0005 and decay rate of 0.8. The validation predictions are \\[exponentially weighted\\] averaged over all 5 epochs (not in this commit). `fit_and_predict` returns validation and test predictions for all epochs.\n"},{"metadata":{},"cell_type":"markdown","source":"## Model 1: EfficientNetB0"},{"metadata":{"trusted":true},"cell_type":"code","source":"# obtain model\nmodel = MyDeepModel(engine=efn.EfficientNetB0, input_dims=(224, 224, 3), batch_size=32, learning_rate=5e-4,\n                    num_epochs=5, decay_rate=0.8, decay_steps=1, weights=None, verbose=1)\n\n# Load model\nmodel.load('../input/rsna-models/stage2/model1_stage2.h5')\n\n# obtain test + validation predictions (history.test_predictions, history.valid_predictions)\n#history = model.fit_and_predict(df.iloc[train_idx], df.iloc[valid_idx], test_df)\ny_pred1 = model.predict(test_df, test_images_dir)\n\ndel model\nK.clear_session()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"y_pred1.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Model 2: InceptionV3"},{"metadata":{"trusted":true},"cell_type":"code","source":"\n# obtain model\ninceptionv3 = InceptionV3(include_top=False, weights=None, input_shape= (256, 256, 3),\n                             backend = keras.backend, layers = keras.layers,\n                             models = keras.models, utils = keras.utils)\n\nmodel = MyDeepModel(engine=inceptionv3, input_dims=(256, 256, 3), batch_size=32, learning_rate=5e-4,\n                    num_epochs=5, decay_rate=0.8, decay_steps=1, weights=\"imagenet\", verbose=1, predefined=True)\n\n# Load model\nmodel.load('../input/rsna-models/stage2/model2_stage2.h5')\n\n# obtain test + validation predictions (history.test_predictions, history.valid_predictions)\n#history = model.fit_and_predict(df.iloc[train_idx], df.iloc[valid_idx], test_df)\ny_pred2 = model.predict(test_df, test_images_dir)\n\ndel model\nK.clear_session()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"y_pred2.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Model 3: EfficientNetB3"},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport pydicom\nimport os\nimport collections\nimport sys\nimport glob\nimport random\nimport cv2\nimport tensorflow as tf\nimport multiprocessing\n\nfrom math import ceil, floor\nfrom copy import deepcopy\nfrom tqdm import tqdm\nfrom imgaug import augmenters as iaa\n\nimport keras\nimport keras.backend as K\nfrom keras.callbacks import Callback, ModelCheckpoint\nfrom keras.layers import Dense, Flatten, Dropout\nfrom keras.models import Model, load_model\nfrom keras.utils import Sequence\nfrom keras.losses import binary_crossentropy\nfrom keras.optimizers import Adam\n\n# Seed\nSEED = 12345\nnp.random.seed(SEED)\n\n# Constants\nTEST_SIZE = 0.01\nHEIGHT = 256\nWIDTH = 256\nCHANNELS = 3\nTRAIN_BATCH_SIZE = 32\nVALID_BATCH_SIZE = 64\nSHAPE = (HEIGHT, WIDTH, CHANNELS)\n\n# Folders\nDATA_DIR = '../input/rsna-intracranial-hemorrhage-detection/rsna-intracranial-hemorrhage-detection/'\nTEST_IMAGES_DIR = DATA_DIR + 'stage_2_test/'\nTRAIN_IMAGES_DIR = DATA_DIR + 'stage_2_train/'\n\ndef correct_dcm(dcm):\n    x = dcm.pixel_array + 1000\n    px_mode = 4096\n    x[x>=px_mode] = x[x>=px_mode] - px_mode\n    dcm.PixelData = x.tobytes()\n    dcm.RescaleIntercept = -1000\n\ndef window_image(dcm, window_center, window_width):    \n    if (dcm.BitsStored == 12) and (dcm.PixelRepresentation == 0) and (int(dcm.RescaleIntercept) > -100):\n        correct_dcm(dcm)\n    img = dcm.pixel_array * dcm.RescaleSlope + dcm.RescaleIntercept\n    \n    # Resize\n    img = cv2.resize(img, SHAPE[:2], interpolation = cv2.INTER_LINEAR)\n   \n    img_min = window_center - window_width // 2\n    img_max = window_center + window_width // 2\n    img = np.clip(img, img_min, img_max)\n    return img\n\ndef bsb_window(dcm):\n    brain_img = window_image(dcm, 40, 80)\n    subdural_img = window_image(dcm, 80, 200)\n    soft_img = window_image(dcm, 40, 380)\n    \n    brain_img = (brain_img - 0) / 80\n    subdural_img = (subdural_img - (-20)) / 200\n    soft_img = (soft_img - (-150)) / 380\n    bsb_img = np.array([brain_img, subdural_img, soft_img]).transpose(1,2,0)\n    return bsb_img\n\ndef _read(path, SHAPE):\n    dcm = pydicom.dcmread(path)\n    try:\n        img = bsb_window(dcm)\n    except:\n        img = np.zeros(SHAPE)\n    return img\n\n# Image Augmentation\nsometimes = lambda aug: iaa.Sometimes(0.25, aug)\naugmentation = iaa.Sequential([ iaa.Fliplr(0.25),\n                                iaa.Flipud(0.10),\n                                sometimes(iaa.Crop(px=(0, 25), keep_size = True, sample_independently = False))   \n                            ], random_order = True)       \n        \n    \nclass TestDataGenerator(keras.utils.Sequence):\n    def __init__(self, dataset, labels, batch_size = 16, img_size = SHAPE, img_dir = TEST_IMAGES_DIR, *args, **kwargs):\n        self.dataset = dataset\n        self.ids = dataset.index\n        self.labels = labels\n        self.batch_size = batch_size\n        self.img_size = img_size\n        self.img_dir = img_dir\n        self.on_epoch_end()\n\n    def __len__(self):\n        return int(ceil(len(self.ids) / self.batch_size))\n\n    def __getitem__(self, index):\n        indices = self.indices[index*self.batch_size:(index+1)*self.batch_size]\n        X = self.__data_generation(indices)\n        return X\n\n    def on_epoch_end(self):\n        self.indices = np.arange(len(self.ids))\n    \n    def __data_generation(self, indices):\n        X = np.empty((self.batch_size, *self.img_size))\n        \n        for i, index in enumerate(indices):\n            ID = self.ids[index]\n            image = _read(self.img_dir+ID+\".dcm\", self.img_size)\n            X[i,] = image              \n        return X\n    \n\n\ndef read_testset(filename = DATA_DIR + \"stage_2_sample_submission.csv\"):\n    df = pd.read_csv(filename)\n    df[\"Image\"] = df[\"ID\"].str.slice(stop=12)\n    df[\"Diagnosis\"] = df[\"ID\"].str.slice(start=13)\n    df = df.loc[:, [\"Label\", \"Diagnosis\", \"Image\"]]\n    df = df.set_index(['Image', 'Diagnosis']).unstack(level=-1)\n    return df\n\n# Read Train and Test Datasets\ntest_df = read_testset()\n\ndef predictions(test_df, model):    \n    test_preds = model.predict_generator(TestDataGenerator(test_df, None, 5, SHAPE, TEST_IMAGES_DIR), verbose = 1)\n    return test_preds[:test_df.iloc[range(test_df.shape[0])].shape[0]]\n\ndef ModelCheckpointFull(model_name):\n    return ModelCheckpoint(model_name, \n                            monitor = 'val_loss', \n                            verbose = 1, \n                            save_best_only = False, \n                            save_weights_only = True, \n                            mode = 'min', \n                            period = 1)\n\n# Create Model\ndef create_model():\n    K.clear_session()\n    \n    base_model =  efn.EfficientNetB3(weights = None, include_top = False, pooling = 'avg', input_shape = SHAPE)\n    x = base_model.output\n    # Lets try without Dropout\n    # x = Dropout(0.125)(x)\n    y_pred = Dense(6, activation = 'sigmoid')(x)\n\n    return Model(inputs = base_model.input, outputs = y_pred)\n\n\nmodel = create_model()\nmodel.load_weights('../input/rsna-models/stage2/model_stage2.h5')\n\nmodel.compile(optimizer = Adam(learning_rate = 1e-4), \n                  loss = 'binary_crossentropy',\n                  metrics = ['acc', tf.keras.metrics.AUC()])\n\ny_pred3 = predictions(test_df, model)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"K.clear_session()\ndel model","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Easy Blending"},{"metadata":{"trusted":true},"cell_type":"code","source":"np.save('y_pred1', y_pred1)\nnp.save('y_pred2', y_pred2)\nnp.save('y_pred3', y_pred3) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Easy Blending\nfrom scipy import stats\n\n# Geom Mean\ny_test = 0.4*y_pred1 + 0.3*y_pred2 + 0.3*y_pred3\n\n\nprint(y_test.shape)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 6. Submit test predictions"},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df.iloc[:, :] = y_test\n\ntest_df = test_df.stack().reset_index()\n\ntest_df.insert(loc=0, column='ID', value=test_df['Image'].astype(str) + \"_\" + test_df['Diagnosis'])\n\ntest_df = test_df.drop([\"Image\", \"Diagnosis\"], axis=1)\n\ntest_df.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 7. Improvements\n\nSome improvements that could possibly be made:<br>\n* Image augmentation (which can be put in `_read()`)\n* Different learning rate and learning rate schedule\n* Increased input size\n* Train longer\n* Add more dense layers and regularization (e.g. `keras.layers.Dropout()` before the output layer)\n* Adding some optimal windowing\n<br>\n<br>\n*Feel free to comment!*\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.6.6"}},"nbformat":4,"nbformat_minor":1}