{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Keras model Stacking\nPreprocessing of data is an important process in machine learning and better handling of data always help you grab higher score in Machine Learning hackathons.  \nHowever when we talk of it's subset: **Deep Learning**, preprocessing is not extensively used.  \nThis is because neural nodes are themselves good at extracting relevant information and adjust parameters. It's amazing how a set of parallel linear computation attains so good results.  \n\nNever the less, data scientits always try to gather more data in Deep learning, in order to diversify the dataset.  \n\nA series of models are also used at times to increase score. For example in medical imaging, classification is usually done after segmentation.  \n\nAt times we observe certain pretrained model outperform other and after achieving similar score with maybe 2 or 3 models we try to implement one with less time consumed ( talking about models in pipeline )  \n\nHowever Much like my notebook explaining boosting my score with [Stacking Classifier](https://www.kaggle.com/kabirnagpal/feature-selection-and-stacking-f1-score-99), I've tried to explain implementation of a Stacking of several models into one.  \n**This will surely help for you to grab those final few points that might change your rank.**\n\n**Let's Get Started**","metadata":{}},{"cell_type":"code","source":"from numba import cuda\nimport tensorflow as tf, re, math\n#Activate the line below when there are memory issues and GPU is activated\n# device = cuda.get_current_device()\n# print(device)\n# device.reset()\n# cuda.close()\n# #torch.cuda.isavailable()\n# tf.test.is_gpu_available()","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:29.805694Z","iopub.execute_input":"2022-12-28T09:21:29.805933Z","iopub.status.idle":"2022-12-28T09:21:31.124553Z","shell.execute_reply.started":"2022-12-28T09:21:29.80591Z","shell.execute_reply":"2022-12-28T09:21:31.124064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import h5py\nfrom glob import glob\nfrom pathlib import Path","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:35.017446Z","iopub.execute_input":"2022-12-28T09:21:35.018383Z","iopub.status.idle":"2022-12-28T09:21:35.023044Z","shell.execute_reply.started":"2022-12-28T09:21:35.018332Z","shell.execute_reply":"2022-12-28T09:21:35.021654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:35.791446Z","iopub.execute_input":"2022-12-28T09:21:35.792092Z","iopub.status.idle":"2022-12-28T09:21:35.797242Z","shell.execute_reply.started":"2022-12-28T09:21:35.792063Z","shell.execute_reply":"2022-12-28T09:21:35.796002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:36.48319Z","iopub.execute_input":"2022-12-28T09:21:36.483617Z","iopub.status.idle":"2022-12-28T09:21:36.655758Z","shell.execute_reply.started":"2022-12-28T09:21:36.483577Z","shell.execute_reply":"2022-12-28T09:21:36.654208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"# DEVICE = \"TPU\" #or \"GPU\"\n\n# # USE DIFFERENT SEED FOR DIFFERENT STRATIFIED KFOLD\n# SEED = 42\n\n# FOLDS = 5\n# IMG_SIZE = [360,128]\n\n# BATCH_SIZE = 12\n# EPOCH = 3\n\n# # TEST TIME AUGMENTATION STEPS\n# TTA = 1","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:37.674749Z","iopub.execute_input":"2022-12-28T09:21:37.674997Z","iopub.status.idle":"2022-12-28T09:21:37.679877Z","shell.execute_reply.started":"2022-12-28T09:21:37.674973Z","shell.execute_reply":"2022-12-28T09:21:37.678766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DEVICE = \"TPU\" #or \"GPU\"\n\n# USE DIFFERENT SEED FOR DIFFERENT STRATIFIED KFOLD\nSEED = 42\n\nFOLDS = 5\nIMG_SIZE = [360,128]\n\nBATCH_SIZE = 32\nEPOCH = 35\n\n# TEST TIME AUGMENTATION STEPS\nTTA = 1","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:38.428188Z","iopub.execute_input":"2022-12-28T09:21:38.429071Z","iopub.status.idle":"2022-12-28T09:21:38.434755Z","shell.execute_reply.started":"2022-12-28T09:21:38.429003Z","shell.execute_reply":"2022-12-28T09:21:38.433569Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Importing Packages\nI've decided to go with Keras for now, however you can soon expect to view similar code in PyTorch","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.layers import *\nfrom tensorflow.keras import backend as K\nfrom tensorflow.keras.models import load_model\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:39.594752Z","iopub.execute_input":"2022-12-28T09:21:39.594984Z","iopub.status.idle":"2022-12-28T09:21:39.599097Z","shell.execute_reply.started":"2022-12-28T09:21:39.594958Z","shell.execute_reply":"2022-12-28T09:21:39.598441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras import regularizers\n#from tensorflow.keras.layers import Input,Flatten,Conv3D,MaxPooling3D, Conv2D, Concatenate, Dense, Lambda, BatchNormalization, GlobalAveragePooling2D, Activation,MaxPooling2D\nfrom tensorflow.keras import layers, Model, Input, losses, metrics, optimizers, callbacks","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:40.130689Z","iopub.execute_input":"2022-12-28T09:21:40.131016Z","iopub.status.idle":"2022-12-28T09:21:40.134507Z","shell.execute_reply.started":"2022-12-28T09:21:40.130993Z","shell.execute_reply":"2022-12-28T09:21:40.134066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#!pip install -q efficientnet >> /dev/null # custom effcient net TF model is available","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:40.639646Z","iopub.execute_input":"2022-12-28T09:21:40.639881Z","iopub.status.idle":"2022-12-28T09:21:40.643563Z","shell.execute_reply.started":"2022-12-28T09:21:40.639859Z","shell.execute_reply":"2022-12-28T09:21:40.642727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd, numpy as np\nfrom kaggle_datasets import KaggleDatasets\n\n#import efficientnet.tfkeras as efn\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import roc_auc_score\n\ntf.__version__","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:41.172393Z","iopub.execute_input":"2022-12-28T09:21:41.172696Z","iopub.status.idle":"2022-12-28T09:21:42.042569Z","shell.execute_reply.started":"2022-12-28T09:21:41.172669Z","shell.execute_reply":"2022-12-28T09:21:42.041404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"kaggle kernels output crischir/ensemble-xception-efficientnet-lea -p /path/to/dest","metadata":{}},{"cell_type":"code","source":"#GCS_PATH = KaggleDatasets().get_gcs_path('ensemble-xception-efficientnet-lea')","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:42.18302Z","iopub.execute_input":"2022-12-28T09:21:42.183285Z","iopub.status.idle":"2022-12-28T09:21:42.187075Z","shell.execute_reply.started":"2022-12-28T09:21:42.183261Z","shell.execute_reply":"2022-12-28T09:21:42.186019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# #Current\n# GCS_DS_PATH = KaggleDatasets().get_gcs_path(\"/kaggle/working\")\n# print(GCS_DS_PATH)","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:42.603514Z","iopub.execute_input":"2022-12-28T09:21:42.603793Z","iopub.status.idle":"2022-12-28T09:21:42.607571Z","shell.execute_reply.started":"2022-12-28T09:21:42.603765Z","shell.execute_reply":"2022-12-28T09:21:42.606853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# GCS_PATH_STRATIFICATED_modified=KaggleDatasets().get_gcs_path('traintfrecg2netdataset1')\n# # GCS_PATH_STRATIFICATED = '/kaggle/input/g2net-tfr-spectrogram-datasets/'\n# # GCS_PATH_STRATIFICATED = '/kaggle/input/g2net-tfr-spectrogram-datasets/'\n# files_train = tf.io.gfile.glob(GCS_PATH_STRATIFICATED_modified + '/train*.tfrec')\n# files_test  = tf.io.gfile.glob(GCS_PATH_STRATIFICATED_modified + '/test*.tfrec')\n# # files_test_compozit  = tf.io.gfile.glob(GCS_PATH_STRATIFICATED_modified + '/**/test*.tfrec')","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:42.863688Z","iopub.execute_input":"2022-12-28T09:21:42.863941Z","iopub.status.idle":"2022-12-28T09:21:42.869207Z","shell.execute_reply.started":"2022-12-28T09:21:42.863915Z","shell.execute_reply":"2022-12-28T09:21:42.867825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function to get hardware strategy\ndef get_hardware_strategy():\n    try:\n        # TPU detection. No parameters necessary if TPU_NAME environment variable is\n        # set: this is always the case on Kaggle.\n        tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n        print('Running on TPU ', tpu.master())\n    except ValueError:\n        tpu = None\n\n    if tpu:\n        tf.config.experimental_connect_to_cluster(tpu)\n        tf.tpu.experimental.initialize_tpu_system(tpu)\n        strategy = tf.distribute.experimental.TPUStrategy(tpu)\n        #policy = mixed_precision.Policy('mixed_bfloat16')\n        #mixed_precision.set_global_policy(policy)\n        tf.config.optimizer.set_jit(True)\n    else:\n        # Default distribution strategy in Tensorflow. Works on CPU and single GPU.\n        strategy = tf.distribute.get_strategy()\n\n    print(\"REPLICAS: \", strategy.num_replicas_in_sync)\n    return tpu, strategy\n\ntpu, strategy = get_hardware_strategy()","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:43.221675Z","iopub.execute_input":"2022-12-28T09:21:43.221928Z","iopub.status.idle":"2022-12-28T09:21:48.393646Z","shell.execute_reply.started":"2022-12-28T09:21:43.221905Z","shell.execute_reply":"2022-12-28T09:21:48.392263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"AUTO = tf.data.experimental.AUTOTUNE\nBATCH_SIZE *= strategy.num_replicas_in_sync","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:48.395423Z","iopub.execute_input":"2022-12-28T09:21:48.395693Z","iopub.status.idle":"2022-12-28T09:21:48.401468Z","shell.execute_reply.started":"2022-12-28T09:21:48.395668Z","shell.execute_reply":"2022-12-28T09:21:48.399838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pydantic import BaseModel","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:48.403479Z","iopub.execute_input":"2022-12-28T09:21:48.403889Z","iopub.status.idle":"2022-12-28T09:21:48.517706Z","shell.execute_reply.started":"2022-12-28T09:21:48.403863Z","shell.execute_reply":"2022-12-28T09:21:48.516636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Artefacts from previous notebooks -not quite needed\nclass Config(BaseModel):\n    seed = 887\n    model_name = \"enetb1_v1\"\n    model_dir = \"/kaggle/working/model\"\n    # data\n \n    path_submission = \"/kaggle/input/g2net-detecting-continuous-gravitational-waves/sample_submission.csv\"\n\n    img_size = (360, 360)\n    channels = 3\n    img_shape = (*img_size, channels)\n    # model\n    base_model_weights = \"imagenet\"\n    dropout = 0.3\n    # training\n    shuffle_size = 128\n    epochs = 150\n#     batch_size = 16 * strategy.num_replicas_in_sync\n    batch_size = 32 * strategy.num_replicas_in_sync\n    test_batch_size = 64\n    lr = 2e-5\n    patience = 12\n    \ncfg = Config()\ncfg.dict()","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:48.519136Z","iopub.execute_input":"2022-12-28T09:21:48.519311Z","iopub.status.idle":"2022-12-28T09:21:48.535841Z","shell.execute_reply.started":"2022-12-28T09:21:48.519288Z","shell.execute_reply":"2022-12-28T09:21:48.534955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### F1-Score\nF1 score is great way to learn about your model on an unbalanced dataset. It is given as:  \n\nRecall = TruePositives / (TruePositives + FalseNegatives)\n\nPrecision = TruePositives / (TruePositives + FalsePositives)\n\nF1 = 2 (precision recall) / (precision + recall)\n\n### An F1 score is considered perfect when it’s 1, while the model is a total failure when it’s 0.\n#### 0.5 is coin flipping :)\n","metadata":{}},{"cell_type":"code","source":"def recall_m(y_true, y_pred):\n\n    y_pred = K.cast(K.greater(K.clip(y_pred, 0, 1), 0.5),K.floatx())\n    true_positives = K.round(K.sum(K.clip(y_true * y_pred, 0, 1)))\n    possible_positives = K.sum(K.clip(y_true, 0, 1))\n    recall_ratio = true_positives / (possible_positives + K.epsilon())\n    return recall_ratio\n\ndef precision_m(y_true, y_pred):\n\n    y_pred = K.cast(K.greater(K.clip(y_pred, 0, 1), 0.5), K.floatx())\n    true_positives = K.round(K.sum(K.clip(y_true * y_pred, 0, 1)))\n    predicted_positives = K.sum(y_pred)\n    precision_ratio = true_positives / (predicted_positives + K.epsilon())\n    return precision_ratio\n\ndef f1_m(y_true, y_pred):\n    \n    precision = precision_m(y_true, y_pred)\n    recall = recall_m(y_true, y_pred)\n    return 2*((precision*recall)/(precision+recall+K.epsilon()))","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:48.537274Z","iopub.execute_input":"2022-12-28T09:21:48.537468Z","iopub.status.idle":"2022-12-28T09:21:48.549232Z","shell.execute_reply.started":"2022-12-28T09:21:48.537433Z","shell.execute_reply":"2022-12-28T09:21:48.548242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading Dataset\n**Note** My focus in this notebook is on model stacking. For Image augmentation refer [this](https://medium.com/data-science-community-srm/from-50-to-5000-an-image-augmentation-story-1fc30111e39) blog","metadata":{}},{"cell_type":"code","source":"from path import Path","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:48.550999Z","iopub.execute_input":"2022-12-28T09:21:48.552134Z","iopub.status.idle":"2022-12-28T09:21:48.571286Z","shell.execute_reply.started":"2022-12-28T09:21:48.552095Z","shell.execute_reply":"2022-12-28T09:21:48.569895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# GCS_PATH_STRATIFICATED = KaggleDatasets().get_gcs_path('tfrecg2netdataset2')\nGCS_PATH_STRATIFICATED = KaggleDatasets().get_gcs_path('traintfrecg2netdataset2')\nGCS_PATH_STRATIFICATED_modified=GCS_PATH_STRATIFICATED\n# GCS_PATH_STRATIFICATED = '/kaggle/input/g2net-tfr-spectrogram-datasets/'\n# GCS_PATH_STRATIFICATED = '/kaggle/input/g2net-tfr-spectrogram-datasets/'\nfiles_train = tf.io.gfile.glob(GCS_PATH_STRATIFICATED + '/train*.tfrec')\nfiles_test  = tf.io.gfile.glob(GCS_PATH_STRATIFICATED + '/test*.tfrec')","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:48.572915Z","iopub.execute_input":"2022-12-28T09:21:48.573178Z","iopub.status.idle":"2022-12-28T09:21:49.07054Z","shell.execute_reply.started":"2022-12-28T09:21:48.573152Z","shell.execute_reply":"2022-12-28T09:21:49.069703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_augumentation=True","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:49.07171Z","iopub.execute_input":"2022-12-28T09:21:49.071967Z","iopub.status.idle":"2022-12-28T09:21:49.077398Z","shell.execute_reply.started":"2022-12-28T09:21:49.071934Z","shell.execute_reply":"2022-12-28T09:21:49.076502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def freq_mask(image, DIM=IMG_SIZE, PROBABILITY = 0.66, SZ = 0.05):\n    # input image - is one image of size [dim,dim,3] not a batch of [b,dim,dim,3]\n    # output - image with CT squares of side size SZ*DIM removed\n    \n    # DO DROPOUT WITH PROBABILITY DEFINED ABOVE\n    P = tf.cast( tf.random.uniform([],0,1)<PROBABILITY, tf.int32)\n    if (P==0)|(SZ==0): return image\n    \n    SZ = SZ * tf.random.uniform([],minval=0.25, maxval=1, dtype='float32')\n    \n    # CHOOSE RANDOM LOCATION\n    y = tf.cast( tf.random.uniform([],0,DIM[0]),tf.int32)\n    # COMPUTE SQUARE \n    WIDTH = tf.cast( SZ*DIM[0],tf.int32) * P\n    ya = tf.math.maximum(0,y-WIDTH//2)\n    yb = tf.math.minimum(DIM[0],y+WIDTH//2)\n    xa = 0\n    xb = DIM[1]\n    # DROPOUT IMAGE\n    one = image[ya:yb,0:xa,:]\n\n    two = tf.zeros([yb-ya,xb-xa,2]) \n    three = image[ya:yb,xb:DIM[1],:]\n    middle = tf.concat([one,two,three],axis=1)\n    image = tf.concat([image[0:ya,:,:],middle,image[yb:DIM[0],:,:]],axis=0)\n\n    # RESHAPE HACK SO TPU COMPILER KNOWS SHAPE OF OUTPUT TENSOR \n    image = tf.reshape(image,[DIM[0],DIM[1],2])\n    return image","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:49.079536Z","iopub.execute_input":"2022-12-28T09:21:49.079792Z","iopub.status.idle":"2022-12-28T09:21:49.095079Z","shell.execute_reply.started":"2022-12-28T09:21:49.079767Z","shell.execute_reply":"2022-12-28T09:21:49.093888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def time_mask(image, DIM=IMG_SIZE, PROBABILITY = 0.66, SZ = 0.05):\n    # input image - is one image of size [dim,dim,3] not a batch of [b,dim,dim,3]\n    # output - image with CT squares of side size SZ*DIM removed\n    \n    # DO DROPOUT WITH PROBABILITY DEFINED ABOVE\n    P = tf.cast( tf.random.uniform([],0,1)<PROBABILITY, tf.int32)\n    if (P==0)|(SZ==0): return image\n    \n    SZ = SZ * tf.random.uniform([],minval=0.25, maxval=1, dtype='float32')\n    \n    # CHOOSE RANDOM LOCATION\n    x = tf.cast( tf.random.uniform([],0,DIM[1]),tf.int32) \n    # COMPUTE SQUARE \n    WIDTH = tf.cast( SZ*DIM[1],tf.int32) * P\n    ya = 0\n    yb = DIM[0]\n    xa = tf.math.maximum(0,x-WIDTH//2)\n    xb = tf.math.minimum(DIM[1],x+WIDTH//2)\n    # DROPOUT IMAGE\n    one = image[ya:yb,0:xa,:]\n    two = tf.zeros([yb-ya,xb-xa,2]) \n    three = image[ya:yb,xb:DIM[1],:]\n    middle = tf.concat([one,two,three],axis=1)\n    image = tf.concat([image[0:ya,:,:],middle,image[yb:DIM[0],:,:]],axis=0)\n\n    # RESHAPE HACK SO TPU COMPILER KNOWS SHAPE OF OUTPUT TENSOR \n    image = tf.reshape(image,[DIM[0], DIM[1], 2])\n    return image","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:49.503816Z","iopub.execute_input":"2022-12-28T09:21:49.504077Z","iopub.status.idle":"2022-12-28T09:21:49.513192Z","shell.execute_reply.started":"2022-12-28T09:21:49.504042Z","shell.execute_reply":"2022-12-28T09:21:49.511894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def time_shuffle(image, DIM=IMG_SIZE, PROBABILITY = 0.4):\n    # input image - is one image of size [dim,dim,3] not a batch of [b,dim,dim,3]\n    # output - image with CT squares of side size SZ*DIM removed\n    \n    # DO DROPOUT WITH PROBABILITY DEFINED ABOVE\n    P = tf.cast( tf.random.uniform([],0,1)<PROBABILITY, tf.int32)\n    if (P==0): return image\n    \n\n    # CHOOSE RANDOM LOCATION\n    x = tf.cast( tf.random.uniform([],0,DIM[1]),tf.int32)\n    y = tf.cast( tf.random.uniform([],0,DIM[0]),tf.int32)\n\n    ya = 0\n    yb = DIM[0]\n    xa = tf.math.maximum(0,x)\n    xb = DIM[1]\n   \n    image = tf.concat([image[:,xa:DIM[1],:],image[:,0:xa,:]],axis=1)\n            \n    # RESHAPE HACK SO TPU COMPILER KNOWS SHAPE OF OUTPUT TENSOR \n    image = tf.reshape(image,[DIM[0], DIM[1], 2])\n    return image","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:49.899279Z","iopub.execute_input":"2022-12-28T09:21:49.899517Z","iopub.status.idle":"2022-12-28T09:21:49.90708Z","shell.execute_reply.started":"2022-12-28T09:21:49.899492Z","shell.execute_reply":"2022-12-28T09:21:49.906081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def augmentation(img):\n    img = freq_mask(img, SZ = 0.05)\n    img = time_mask(img, SZ = 0.02)\n    img = tf.image.random_flip_left_right(img) # To be tested\n    img = freq_mask(img, SZ = 0.05)\n    img = time_mask(img, SZ = 0.02)\n    img = freq_mask(img, SZ = 0.07)\n    img = time_mask(img, SZ = 0.07)\n    img = freq_mask(img, SZ = 0.1)\n    img = time_mask(img, SZ = 0.1)\n    img = freq_mask(img, SZ = 0.15)\n    img = time_mask(img, SZ = 0.03)\n    img = time_shuffle(img, PROBABILITY = 0.75)\n    return img","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:50.207559Z","iopub.execute_input":"2022-12-28T09:21:50.207834Z","iopub.status.idle":"2022-12-28T09:21:50.214676Z","shell.execute_reply.started":"2022-12-28T09:21:50.207803Z","shell.execute_reply":"2022-12-28T09:21:50.213303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Read Tfrecs","metadata":{}},{"cell_type":"code","source":"def read_labeled_tfrecord(example):\n    tfrec_format = {\n        'spectrogram'          : tf.io.FixedLenFeature([], tf.string),\n        'target'               : tf.io.FixedLenFeature([], tf.float32),\n        'id'                   : tf.io.FixedLenFeature([], tf.string),\n    }           \n    example = tf.io.parse_single_example(example, tfrec_format)\n    example['spectrogram'] = tf.io.parse_tensor(example['spectrogram'],out_type=tf.float32)\n    example['spectrogram'] = tf.reshape(example['spectrogram'], [*IMG_SIZE,2])\n    #print(example['spectrogram'].shape)\n    return example['spectrogram'], example['target']\n\n\ndef read_unlabeled_tfrecord(example, return_image_name):\n    tfrec_format = {\n        'spectrogram'          : tf.io.FixedLenFeature([], tf.string),\n        'id'                   : tf.io.FixedLenFeature([], tf.string)\n    }\n    example = tf.io.parse_single_example(example, tfrec_format)\n    example['spectrogram'] = tf.io.parse_tensor(example['spectrogram'],out_type=tf.float32)\n    example['spectrogram'] = tf.reshape(example['spectrogram'], [*IMG_SIZE,2])\n    return example['spectrogram'], example['id'] if return_image_name else 0\n\n \ndef prepare_image(img, augment=False):    \n    if augment:\n        img = augmentation(img)\n# for testing withou a dummy function should be created: \n                                            # def augmentation(img):\n                                            #     img = img\n                                            #     return img\n    #The tf.reshape does not change the order of or the total number of elements in the tensor, \n    #and so it can reuse the underlying data buffer. \n    #This makes it a fast operation independent of how big of a tensor it is operating on.\n    img = tf.reshape(img, [*IMG_SIZE, 2])\n            \n    return img\n\ndef count_data_items(filenames):\n    n = [int(re.compile(r\"-([0-9]*)\\.\").search(filename).group(1)) \n         for filename in filenames]\n    return np.sum(n)","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:50.772501Z","iopub.execute_input":"2022-12-28T09:21:50.773327Z","iopub.status.idle":"2022-12-28T09:21:50.785153Z","shell.execute_reply.started":"2022-12-28T09:21:50.773299Z","shell.execute_reply":"2022-12-28T09:21:50.783453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_dataset(files, augment = False, shuffle = False, repeat = False, \n                labeled=True, return_image_names=True, batch_size=32):\n    \n    ds = tf.data.TFRecordDataset(files, num_parallel_reads=AUTO)\n    ds = ds.cache()\n    \n    if repeat:\n        ds = ds.repeat()\n    \n    if shuffle: \n        ds = ds.shuffle(1024*8 if tpu else 128)\n        opt = tf.data.Options()\n        opt.experimental_deterministic = False\n        ds = ds.with_options(opt)\n        \n    if labeled: \n        ds = ds.map(read_labeled_tfrecord, num_parallel_calls=AUTO)\n    else:\n        ds = ds.map(lambda example: read_unlabeled_tfrecord(example, return_image_names), \n                    num_parallel_calls=AUTO)      \n    #Augumentation here\n    ds = ds.map(lambda img, imgname_or_label: (prepare_image(img, augment=augment,), \n                                               imgname_or_label), \n                num_parallel_calls=AUTO)\n    \n    ds = ds.batch(batch_size)\n    ds = ds.prefetch(AUTO)\n    return ds","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:51.102089Z","iopub.execute_input":"2022-12-28T09:21:51.102355Z","iopub.status.idle":"2022-12-28T09:21:51.110048Z","shell.execute_reply.started":"2022-12-28T09:21:51.102324Z","shell.execute_reply":"2022-12-28T09:21:51.10902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plotting some images","metadata":{}},{"cell_type":"code","source":"FNAME = tf.io.gfile.glob([GCS_PATH_STRATIFICATED + '/train*.tfrec'])\nrow = 16; col = 2;\nrow = min(row,16//col)\n\n# all_elements = get_dataset(FNAME, augment=False, batch_size=32).unbatch()\n# augmented_element = all_elements.repeat().batch(32)\nall_elements = get_dataset(FNAME, augment=image_augumentation, batch_size=32)\n# for (img,mask) in augmented_element:\nfor (img,mask) in all_elements:\n\n    img = (img-np.min(img))/(np.max(img)-np.min(img)+1e-5)\n    plt.figure(figsize=(15,int(15*row/col)))\n\n    i=0\n    j=1\n    while j<=row*col:\n        plt.subplot(row,col,j)\n        plt.axis('on')\n\n        plt.imshow(img[i,:,:,0])\n        j += 1\n        \n        plt.subplot(row,col,j)\n        plt.axis('on')\n        plt.imshow(img[i,:,:,1])\n        j += 1\n\n        i += 1\n\n    plt.show()\n    break","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:21:52.238452Z","iopub.execute_input":"2022-12-28T09:21:52.238715Z","iopub.status.idle":"2022-12-28T09:22:03.117833Z","shell.execute_reply.started":"2022-12-28T09:21:52.23869Z","shell.execute_reply":"2022-12-28T09:22:03.114922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:03.119384Z","iopub.execute_input":"2022-12-28T09:22:03.11966Z","iopub.status.idle":"2022-12-28T09:22:03.398143Z","shell.execute_reply.started":"2022-12-28T09:22:03.119627Z","shell.execute_reply":"2022-12-28T09:22:03.397181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# img_width, img_height = 160, 160\n\n# if K.image_data_format() == 'channels_first':\n#     input_shape = (3, img_width, img_height)\n# else:\n#     input_shape = (img_width, img_height, 3)\n    \n# batch_size = 32\n# epochs = 25","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:03.399134Z","iopub.execute_input":"2022-12-28T09:22:03.400005Z","iopub.status.idle":"2022-12-28T09:22:03.410319Z","shell.execute_reply.started":"2022-12-28T09:22:03.399969Z","shell.execute_reply":"2022-12-28T09:22:03.409185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from keras.preprocessing.image import ImageDataGenerator\n\n# train_datagen = ImageDataGenerator(rescale = 1./255,\n#                                    validation_split=0.15)\n\n# training_set = train_datagen.flow_from_directory('../input/apparel-images-dataset',\n#                                                  target_size = (img_width, img_height),\n#                                                  batch_size = batch_size,\n#                                                  subset='training')\n\n# val_set = train_datagen.flow_from_directory('../input/apparel-images-dataset',\n#                                                  target_size = (img_width, img_height),\n#                                                  batch_size = batch_size,\n#                                                  subset='validation')","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:03.413092Z","iopub.execute_input":"2022-12-28T09:22:03.413407Z","iopub.status.idle":"2022-12-28T09:22:03.43225Z","shell.execute_reply.started":"2022-12-28T09:22:03.413376Z","shell.execute_reply":"2022-12-28T09:22:03.431267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Callbacks\nCallbacks are an easy way to stop training in case of Over fitting.  \nThere are many other Callbacks as well which can reduce Learning rate, or stop if error reaches NaN etc.  \nRefer [here](https://keras.io/api/callbacks/) for more details","metadata":{}},{"cell_type":"code","source":"def get_lr_callback():\n    lr_start   = 5e-5\n    lr_max     = 5e-4\n    lr_min     = 1e-5\n    lr_ramp_ep = 4\n    lr_sus_ep  = 4\n    lr_decay   = 0.9\n   \n    def lrfn(epoch):\n        if epoch < lr_ramp_ep:\n            lr = (lr_max - lr_start) / lr_ramp_ep * epoch + lr_start\n            \n        elif epoch < lr_ramp_ep + lr_sus_ep:\n            lr = lr_max\n            \n        else:\n            lr = (lr_max - lr_min) * lr_decay**(epoch - lr_ramp_ep - lr_sus_ep) + lr_min\n            \n        return lr\n\n    lr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose=False)\n    return lr_callback","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:03.433619Z","iopub.execute_input":"2022-12-28T09:22:03.433874Z","iopub.status.idle":"2022-12-28T09:22:03.445506Z","shell.execute_reply.started":"2022-12-28T09:22:03.433842Z","shell.execute_reply":"2022-12-28T09:22:03.444346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"files_train = tf.io.gfile.glob(GCS_PATH_STRATIFICATED_modified + '/train*.tfrec')\nfiles_test  = tf.io.gfile.glob(GCS_PATH_STRATIFICATED_modified + '/test*.tfrec')\n# files_test_compozit  = tf.io.gfile.glob(GCS_PATH_STRATIFICATED_modified + '/**/test*.tfrec'","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:03.446893Z","iopub.execute_input":"2022-12-28T09:22:03.447126Z","iopub.status.idle":"2022-12-28T09:22:03.533338Z","shell.execute_reply.started":"2022-12-28T09:22:03.447102Z","shell.execute_reply":"2022-12-28T09:22:03.532624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"valid_files=files_train[0]\ntrain_files=files_train[1:]","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:16.35499Z","iopub.execute_input":"2022-12-28T09:22:16.355237Z","iopub.status.idle":"2022-12-28T09:22:16.360214Z","shell.execute_reply.started":"2022-12-28T09:22:16.355211Z","shell.execute_reply":"2022-12-28T09:22:16.358746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"valid_files","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:20.505322Z","iopub.execute_input":"2022-12-28T09:22:20.505554Z","iopub.status.idle":"2022-12-28T09:22:20.51156Z","shell.execute_reply.started":"2022-12-28T09:22:20.505531Z","shell.execute_reply":"2022-12-28T09:22:20.510503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_files","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:21.712362Z","iopub.execute_input":"2022-12-28T09:22:21.712631Z","iopub.status.idle":"2022-12-28T09:22:21.719811Z","shell.execute_reply.started":"2022-12-28T09:22:21.712582Z","shell.execute_reply":"2022-12-28T09:22:21.718128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Stacking","metadata":{}},{"cell_type":"code","source":"import tensorflow.keras.applications as apps\nhelp(apps)","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:24.044637Z","iopub.execute_input":"2022-12-28T09:22:24.045003Z","iopub.status.idle":"2022-12-28T09:22:24.051396Z","shell.execute_reply.started":"2022-12-28T09:22:24.044973Z","shell.execute_reply":"2022-12-28T09:22:24.050716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.applications import Xception\n#Kaagle does not have the last version of tensorflow\n#from tensorflow.keras.applications import EfficientNetV2S\nfrom tensorflow.keras.applications import EfficientNetB5","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:16:46.049117Z","iopub.status.idle":"2022-12-28T09:16:46.049386Z","shell.execute_reply.started":"2022-12-28T09:16:46.049247Z","shell.execute_reply":"2022-12-28T09:16:46.049264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.applications import MobileNetV2","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:34.21934Z","iopub.execute_input":"2022-12-28T09:22:34.219571Z","iopub.status.idle":"2022-12-28T09:22:34.225439Z","shell.execute_reply.started":"2022-12-28T09:22:34.219549Z","shell.execute_reply":"2022-12-28T09:22:34.223916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.applications import DenseNet121","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:34.399323Z","iopub.execute_input":"2022-12-28T09:22:34.399805Z","iopub.status.idle":"2022-12-28T09:22:34.402964Z","shell.execute_reply.started":"2022-12-28T09:22:34.399781Z","shell.execute_reply":"2022-12-28T09:22:34.402362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"input_shape=[360,128,2]","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:34.611522Z","iopub.execute_input":"2022-12-28T09:22:34.611853Z","iopub.status.idle":"2022-12-28T09:22:34.616231Z","shell.execute_reply.started":"2022-12-28T09:22:34.611811Z","shell.execute_reply":"2022-12-28T09:22:34.615453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#K.clear_session()","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:34.956352Z","iopub.execute_input":"2022-12-28T09:22:34.95675Z","iopub.status.idle":"2022-12-28T09:22:34.960779Z","shell.execute_reply.started":"2022-12-28T09:22:34.956725Z","shell.execute_reply":"2022-12-28T09:22:34.959172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow.keras.metrics import AUC","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:35.311428Z","iopub.execute_input":"2022-12-28T09:22:35.311812Z","iopub.status.idle":"2022-12-28T09:22:35.315332Z","shell.execute_reply.started":"2022-12-28T09:22:35.311786Z","shell.execute_reply":"2022-12-28T09:22:35.314814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def models(input_size = input_shape):\n\n    inputs = tf.keras.layers.Input(input_size)\n    first_conv = tf.keras.layers.Conv2D(3, 7, strides=(1, 1), padding='same')(inputs)\n    second_conv = tf.keras.layers.Conv2D(3, 2, strides=(1, 1), padding='same')(inputs)\n    base_model1 =     tf.keras.applications.EfficientNetB5(\n        include_top = False,\n        weights     = 'imagenet',#initially None\n        #pooling     = 'avg'\n        )(first_conv)\n    base_model1 =tf.keras.layers.GlobalAveragePooling2D()(base_model1)\n    base_model2 = tf.keras.layers.AveragePooling2D(pool_size=(2, 2),strides=(1, 1), padding='valid')(second_conv)\n    base_model2 =     tf.keras.applications.Xception(\n        include_top = False,\n        weights     = 'imagenet',#initially None\n        pooling     = 'max'\n        )(base_model2)\n    base_model3 = tf.keras.layers.MaxPooling2D(pool_size=(2, 2),strides=(1, 1), padding='valid')(second_conv)\n    base_model3 =     tf.keras.applications.MobileNetV2(\n        include_top = False,\n        weights     = 'imagenet',#initially None\n        pooling     = 'max'\n        )(base_model3)\n    base_model4 =     tf.keras.applications.DenseNet121(\n    include_top = False,\n    weights     = 'imagenet',#initially None\n    pooling     = 'avg'\n    )(second_conv)\n    #base_model2 =tf.keras.layers.GlobalAveragePooling2D()(base_model2)  \n    model = tf.keras.layers.Concatenate()([base_model1,base_model2,base_model3,base_model4])\n    model = tf.keras.layers.Dense(1024, activation='relu')(model)\n    model = tf.keras.layers.Dense(512, activation='relu')(model)\n    model = tf.keras.layers.Dense(256, activation='relu')(model)\n    model = tf.keras.layers.Dense(128, activation='relu')(model)\n    model = tf.keras.layers.Dense(1, activation='sigmoid')(model)\n\n    M = tf.keras.Model(inputs=inputs, outputs=model)\n    M.compile(\n        optimizer = 'adam', \n        loss = 'BinaryCrossentropy',               \n        metrics=[AUC(),f1_m]\n    )\n    \n    return M\n\nmodel = models()\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:35.882068Z","iopub.execute_input":"2022-12-28T09:22:35.882438Z","iopub.status.idle":"2022-12-28T09:22:53.694047Z","shell.execute_reply.started":"2022-12-28T09:22:35.882414Z","shell.execute_reply":"2022-12-28T09:22:53.693057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.keras.utils.plot_model(model, to_file=\"my_model.png\", show_shapes=True)","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:22:57.189335Z","iopub.execute_input":"2022-12-28T09:22:57.189567Z","iopub.status.idle":"2022-12-28T09:22:58.769075Z","shell.execute_reply.started":"2022-12-28T09:22:57.189545Z","shell.execute_reply":"2022-12-28T09:22:58.76833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"files_train = tf.io.gfile.glob(train_files)\nfiles_valid = tf.io.gfile.glob(valid_files)\nfiles_test = tf.io.gfile.glob(GCS_PATH_STRATIFICATED_modified + '/test*.tfrec')","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:23:08.268534Z","iopub.execute_input":"2022-12-28T09:23:08.268816Z","iopub.status.idle":"2022-12-28T09:23:08.499195Z","shell.execute_reply.started":"2022-12-28T09:23:08.268789Z","shell.execute_reply":"2022-12-28T09:23:08.498066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# #test_df=pd.DataFrame()\n# test_df = pd.read_csv('../input/g2net-detecting-continuous-gravitational-waves/sample_submission.csv')\n# #test_files = glob(f\"{ROOT_DIR}/test/*.hdf5\")\n# test_df['cale_fisier']='/kaggle/input/g2net-detecting-continuous-gravitational-waves/test/'+test_df['id']+'.hdf5'\n# test_df=test_df.sample(10)\n# test_df=test_df.reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:16:54.439189Z","iopub.status.idle":"2022-12-28T09:16:54.43954Z","shell.execute_reply.started":"2022-12-28T09:16:54.439356Z","shell.execute_reply":"2022-12-28T09:16:54.439374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from tqdm.auto import tqdm","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:23:15.462845Z","iopub.execute_input":"2022-12-28T09:23:15.463132Z","iopub.status.idle":"2022-12-28T09:23:15.468515Z","shell.execute_reply.started":"2022-12-28T09:23:15.463095Z","shell.execute_reply":"2022-12-28T09:23:15.467309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Utility to read hdf5 file\n# def read_data(file: Path):\n#     with h5py.File(file, \"r\") as f:\n#         file=Path(file)\n#         filename = file.stem\n#         f = f[filename]\n#         h1 = f[\"H1\"]\n#         l1 = f[\"L1\"]\n#         #freq_hz = list(f[\"frequency_Hz\"])\n#         ###\n#         h1_stft = h1[\"SFTs\"][()]\n#         ###\n#         #h1_timestamp = h1[\"timestamps_GPS\"][()]\n#         h1_timestamp =1\n#         # H2 data\n#         l1_stft = l1[\"SFTs\"][()]\n#         #l1_timestamp = l1[\"timestamps_GPS\"][()]\n#         l1_timestamp =2\n#         return {\n#             \"H1\": [h1_stft, h1_timestamp],\n#             \"L1\": [l1_stft, l1_timestamp],\n#             #\"freq_hz\": freq_hz\n#         }\n# def read_data_c(file):\n#     file = Path(file)\n#     with h5py.File(file, \"r\") as f:\n#         filename = file.stem\n#         f = f[filename]\n#         h1 = f[\"H1\"]\n#         l1 = f[\"L1\"]\n#         freq_hz = f[\"frequency_Hz\"]\n        \n#         h1_stft = h1[\"SFTs\"][()]\n#         h1_timestamp = h1[\"timestamps_GPS\"][()]\n#         # H2 data\n#         l1_stft = l1[\"SFTs\"][()]\n#         l1_timestamp = l1[\"timestamps_GPS\"][()]\n        \n#         return [h1_stft, h1_timestamp],            [l1_stft, l1_timestamp], np.array(freq_hz)\n# def power_spectrogram2(h1_sft,l1_sft):\n#     i=0\n    \n#     img = np.empty((360,360,3), dtype=np.float32)\n#     try:\n#         a=h1_sft[:360, :4320]*1e22\n#         p = a.real**2 + a.imag**2\n#         p /= np.mean(p)  # normalize\n#         q=(a.real*2)*(np.angle(a))\n#         q /= np.mean(q)  # normalize\n#         try:\n#             p = np.mean(p.reshape(360, 360,12), axis=2)\n#             q = np.mean(q.reshape(360, 360,12), axis=2)\n#         except:\n#             a=h1_sft[:360, :3960]*1e22\n#             p = a.real**2 + a.imag**2\n#             q=(a.real*2)*(np.angle(a))\n#             p /= np.mean(p)  # normalize\n#             q /= np.mean(q)  # normalize\n#             p = np.mean(p.reshape(360, 360,11), axis=2)\n#             q = np.mean(q.reshape(360, 360,11), axis=2)\n\n#         #normalized=(255*(p - np.min(p))/np.ptp(p)).astype(int)\n\n#     #     normalized=(255*(p - np.min(p))/np.ptp(p)).astype(int)\n#         img[:,:,0]= p*255\n# #         img[:,:,2]= np.ones((360,360), dtype=np.float32)\n# #         img[:,:,2]=img[:,:,2]*255\n\n\n#         img[:,:,2]= q*255\n#         a=l1_sft[:360, :4320]*1e22\n#         p = a.real**2 + a.imag**2\n#         p /= np.mean(p)  # normalize\n#         try:\n#             p = np.mean(p.reshape(360, 360,12), axis=2)\n#         except:\n#             a=l1_sft[:360, :3960]*1e22\n#             p = a.real**2 + a.imag**2\n#             p /= np.mean(p)  # normalize\n#             p = np.mean(p.reshape(360, 360,11), axis=2)\n\n#         #normalized=(255*(p - np.min(p))/np.ptp(p)).astype(int)\n#         img[:,:,1]= p*255\n\n# #         img = np.moveaxis(img, 0, -1)\n#         #return np.asarray([img[0],img[1]]).astype('float32')\n#     except:\n#         print(\"empty image\")\n# #         img = np.moveaxis(img, 0, -1)\n\n#     return img\n# def show_power_spectrogram2(filename):\n#     f = h5py.File(filename, 'r')\n\n#   # read fourier transform coefficients\n#     (h1_sfts, h1_ts), (l1_sfts, l1_ts), freq=read_data_c(filename)\n#     img=power_spectrogram2(h1_sfts,l1_sfts)\n#     plt.imshow(img)\n#     plt.show()\n# def test_dataset(filename):\n#     filename=Path(filename)\n#     with h5py.File(filename) as f:\n#                 filename=filename.stem\n#                 g = f[filename]\n#                 print(filename)\n#                 try:\n#                     for ch, s in enumerate(['H1', 'L1']):\n#                             x=g[s]['SFTs'].shape\n#                 except:\n#                     print('bad dataset')\n# def _bytes_feature(value):\n#   \"\"\"Returns a bytes_list from a string / byte.\"\"\"\n#   if isinstance(value, type(tf.constant(0))):\n#     value = value.numpy() # BytesList won't unpack a string from an EagerTensor.\n#   return tf.train.Feature(bytes_list=tf.train.BytesList(value=[value]))\n# def _float_feature(value):\n#   \"\"\"Returns a float_list from a float / double.\"\"\"\n#   return tf.train.Feature(float_list=tf.train.FloatList(value=[value]))\n# def _int64_feature(value):\n#   \"\"\"Returns an int64_list from a bool / enum / int / uint.\"\"\"\n#   return tf.train.Feature(int64_list=tf.train.Int64List(value=[value]))\n# def serialize_example_train(img, tgt, name):\n#   feature = {\n#       'spectrogram': _bytes_feature(img),\n#       'target': _float_feature(tgt),\n#       'id': _bytes_feature(name),\n#   }\n#   example_proto = tf.train.Example(features=tf.train.Features(feature=feature))\n#   return example_proto.SerializeToString()\n\n# def serialize_example_test(img, name):\n#   feature = {\n#       'spectrogram': _bytes_feature(img),\n#       'id': _bytes_feature(name),\n#   }\n#   example_proto = tf.train.Example(features=tf.train.Features(feature=feature))\n#   return example_proto.SerializeToString()\n# heigh=360\n# # width=360\n# # first_cut=4320 #360\n# # compress_1=12\n# # second_cut=3960 #360\n# # compress_2=11\n# width=128\n# first_cut=4096\n# compress_1=32\n# second_cut=4096-128\n# compress_2=31","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:23:15.954735Z","iopub.execute_input":"2022-12-28T09:23:15.95498Z","iopub.status.idle":"2022-12-28T09:23:15.963158Z","shell.execute_reply.started":"2022-12-28T09:23:15.954955Z","shell.execute_reply":"2022-12-28T09:23:15.962022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # %%time\n# #test_df = pd.read_csv(di + '/sample_submission.csv')\n# train_only=False\n# if train_only:\n#     print('skip the nonstationary model dataset save')\n# else:\n#     ct = len(test_df)\n#     idx = test_df.index\n#     f = 0\n#     print(ct)\n#     print('Writing TFRecord %i of %i...'%(f,ct))\n#     xi=0\n#     with tf.io.TFRecordWriter('test%.2i-%i.tfrec'%(f,ct)) as writer:\n#         for _, i in enumerate((tqdm(idx))):\n#             r = test_df.iloc[i]\n#             file_id = r.id\n#             filename=r.cale_fisier\n#             img = np.empty((360, width, 2), dtype=np.float32)\n\n#             #filename = '%s/test/%s.hdf5' % (di, file_id)\n#             with h5py.File(filename, 'r') as f:\n#                 g = f[file_id]\n#                 #print('unu')\n#                 for ch, s in enumerate(['H1', 'L1']):\n#                     print(xi)\n#                     xi=xi+1\n#                     try:\n#                         a = g[s]['SFTs'][:360, :first_cut] * 1e22  # Fourier coefficient complex64\n#                         p = 2*(a.real**2 + a.imag**2)  # power\n#                         p /= np.mean(p)  # normalize\n#                         p = np.mean(p.reshape(360, width,compress_1), axis=2)\n#                         #print('doi')\n#                     except:\n#                         a = g[s]['SFTs'][:360, :second_cut]* 1e22  # Fourier coefficient complex64\n#                         p = 2*(a.real**2 + a.imag**2)  # power\n#                         p /= np.mean(p)  # normalize\n#                         p = np.mean(p.reshape(360, width,compress_2), axis=2)\n#                     img[...,ch] = p\n\n#     #             try:\n#     #                 a=g[\"H1\"]\n#     #                 a=a['SFTs'][:360, :4320]*1e22\n#     #                 #a = g['H1']['SFTs'][:360, :4320]# Fourier coefficient complex64\n#     #                 q = np.sin(np.angle(a))*(np.abs(a)**2)  # power\n#     #                 q /= np.mean(q)  # normalize\n#     #                 q = np.mean(q.reshape(360, 360,12), axis=2)\n#     #                 print(np.min(q),np.max(q))\n#     #                 scaler = MinMaxScaler(feature_range=(0, 255))\n#     #                 scaler = scaler.fit(q)\n#     #                 img[...,2] = q\n#     #                 #cv2.imwrite(\"blabla.jpg\",img)\n#     #             except:\n#     #                 a=g[\"H1\"]\n#     #                 a=a['SFTs'][:360, :3960]*1e22\n#     #                 p = np.sin(np.angle(a))*(np.abs(a)**2)  # power\n#     #                 p /= np.mean(p)  # normalize\n#     #                 p = np.mean(p.reshape(360, 360,11), axis=2)\n#     #                 scaler = scaler.fit(p)\n#     #                 img[...,2] = p  \n#             serialized_img = tf.io.serialize_tensor(img)\n#             example = serialize_example_test(serialized_img, str.encode(file_id))\n#             writer.write(example)\n# train_only=False","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:23:16.770398Z","iopub.execute_input":"2022-12-28T09:23:16.77067Z","iopub.status.idle":"2022-12-28T09:23:16.77652Z","shell.execute_reply.started":"2022-12-28T09:23:16.770646Z","shell.execute_reply":"2022-12-28T09:23:16.77533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# files_test = tf.io.gfile.glob(\"/kaggle/working\" + '/test*.tfrec')","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:23:17.317647Z","iopub.execute_input":"2022-12-28T09:23:17.317897Z","iopub.status.idle":"2022-12-28T09:23:17.322561Z","shell.execute_reply.started":"2022-12-28T09:23:17.31787Z","shell.execute_reply":"2022-12-28T09:23:17.321328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#files_test1","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:23:17.727664Z","iopub.execute_input":"2022-12-28T09:23:17.727965Z","iopub.status.idle":"2022-12-28T09:23:17.73226Z","shell.execute_reply.started":"2022-12-28T09:23:17.727941Z","shell.execute_reply":"2022-12-28T09:23:17.731088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# files_test","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:23:18.007577Z","iopub.execute_input":"2022-12-28T09:23:18.007912Z","iopub.status.idle":"2022-12-28T09:23:18.012582Z","shell.execute_reply.started":"2022-12-28T09:23:18.007887Z","shell.execute_reply":"2022-12-28T09:23:18.011569Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## train","metadata":{}},{"cell_type":"code","source":"# DEVICE = \"TPU\" #or \"GPU\"\n\n# # USE DIFFERENT SEED FOR DIFFERENT STRATIFIED KFOLD\n# SEED = 42\n\n# FOLDS = 5\n# IMG_SIZE = [360,128]\n\n# BATCH_SIZE = 32\n# EPOCH = 1\n\n# # TEST TIME AUGMENTATION STEPS\n# TTA = 1","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:23:18.939174Z","iopub.execute_input":"2022-12-28T09:23:18.939435Z","iopub.status.idle":"2022-12-28T09:23:18.943436Z","shell.execute_reply.started":"2022-12-28T09:23:18.93941Z","shell.execute_reply":"2022-12-28T09:23:18.942269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#files_train = tf.io.gfile.glob([GCS_PATH_STRATIFICATED_modified + '/train%.2i*.tfrec'])","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:23:19.381327Z","iopub.execute_input":"2022-12-28T09:23:19.38158Z","iopub.status.idle":"2022-12-28T09:23:19.385629Z","shell.execute_reply.started":"2022-12-28T09:23:19.38155Z","shell.execute_reply":"2022-12-28T09:23:19.384717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# BUILD MODEL\n# K.clear_session()\nwith strategy.scope():\n    model = models()","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:23:20.297112Z","iopub.execute_input":"2022-12-28T09:23:20.297397Z","iopub.status.idle":"2022-12-28T09:24:24.587012Z","shell.execute_reply.started":"2022-12-28T09:23:20.297368Z","shell.execute_reply":"2022-12-28T09:24:24.586087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# SAVE BEST MODEL \n\nsv = tf.keras.callbacks.ModelCheckpoint(\n    \"/kaggle/working/agregatted_model.h5\", monitor='val_loss', verbose=0, save_best_only=True,\n    save_weights_only=True, mode='min', save_freq='epoch')\nes =tf.keras.callbacks.EarlyStopping(\n    monitor='val_loss',\n    min_delta=0,\n    patience=12,\n    verbose=1,\n    mode='auto',\n    )\n   ","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:24:24.589078Z","iopub.execute_input":"2022-12-28T09:24:24.589366Z","iopub.status.idle":"2022-12-28T09:24:24.595763Z","shell.execute_reply.started":"2022-12-28T09:24:24.589334Z","shell.execute_reply":"2022-12-28T09:24:24.594414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# def connect_to_tpu(tpu_address: str = None):\n#     if tpu_address is not None:  # When using GCP\n#         cluster_resolver = tf.distribute.cluster_resolver.TPUClusterResolver(\n#             tpu=tpu_address)\n#         if tpu_address not in (\"\", \"local\"):\n#             tf.config.experimental_connect_to_cluster(cluster_resolver)\n#         tf.tpu.experimental.initialize_tpu_system(cluster_resolver)\n#         strategy = tf.distribute.experimental.TPUStrategy(cluster_resolver)\n#         print(\"Running on TPU \", cluster_resolver.master())\n#         print(\"REPLICAS: \", strategy.num_replicas_in_sync)\n#         return cluster_resolver, strategy\n#     else:                           # When using Colab or Kaggle\n#         try:\n#             cluster_resolver = tf.distribute.cluster_resolver.TPUClusterResolver.connect()\n#             strategy = tf.distribute.experimental.TPUStrategy(cluster_resolver)\n#             print(\"Running on TPU \", cluster_resolver.master())\n#             print(\"REPLICAS: \", strategy.num_replicas_in_sync)\n#             return cluster_resolver, strategy\n#         except:\n#             print(\"WARNING: No TPU detected.\")\n#             mirrored_strategy = tf.distribute.MirroredStrategy()\n#             return None, mirrored_strategy","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:24:24.597101Z","iopub.execute_input":"2022-12-28T09:24:24.59736Z","iopub.status.idle":"2022-12-28T09:24:24.611003Z","shell.execute_reply.started":"2022-12-28T09:24:24.597328Z","shell.execute_reply":"2022-12-28T09:24:24.609844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# files_train = tf.io.gfile.glob(train_files)\n# files_valid = tf.io.gfile.glob(valid_files)\n# #files_test = tf.io.gfile.glob(GCS_PATH_STRATIFICATED_modified + '/test*.tfrec')\n# # files_test = tf.io.gfile.glob(\"/kaggle/working\" + '/test*.tfrec')","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:16:54.45954Z","iopub.status.idle":"2022-12-28T09:16:54.459906Z","shell.execute_reply.started":"2022-12-28T09:16:54.459745Z","shell.execute_reply":"2022-12-28T09:16:54.45976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"files_valid ","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:25:54.431895Z","iopub.execute_input":"2022-12-28T09:25:54.432165Z","iopub.status.idle":"2022-12-28T09:25:54.439071Z","shell.execute_reply.started":"2022-12-28T09:25:54.432134Z","shell.execute_reply":"2022-12-28T09:25:54.43781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# TRAIN\nprint('Training...')\nhistory = model.fit(\n    get_dataset(files_train, augment=image_augumentation, shuffle=True, repeat=True,\n            batch_size=BATCH_SIZE), \n    epochs=EPOCH, callbacks = [sv,get_lr_callback()], \n    steps_per_epoch=count_data_items(files_train)/BATCH_SIZE,\n    validation_data=get_dataset(files_valid,augment=False,shuffle=False,\n            repeat=False),\n#         verbose=VERBOSE\n)","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:25:58.792433Z","iopub.execute_input":"2022-12-28T09:25:58.792736Z","iopub.status.idle":"2022-12-28T09:49:43.286237Z","shell.execute_reply.started":"2022-12-28T09:25:58.792702Z","shell.execute_reply":"2022-12-28T09:49:43.284865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Loading best model...')\nmodel.load_weights(\"agregatted_model.h5\")","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:55:54.781043Z","iopub.execute_input":"2022-12-28T09:55:54.781665Z","iopub.status.idle":"2022-12-28T09:56:09.510462Z","shell.execute_reply.started":"2022-12-28T09:55:54.781631Z","shell.execute_reply":"2022-12-28T09:56:09.509211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"oof_pred = []; oof_tar = []; oof_val = []; oof_names = []; oof_folds = [] \npreds = np.zeros((count_data_items(files_test),1))","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:56:09.513814Z","iopub.execute_input":"2022-12-28T09:56:09.514272Z","iopub.status.idle":"2022-12-28T09:56:09.521353Z","shell.execute_reply.started":"2022-12-28T09:56:09.514242Z","shell.execute_reply":"2022-12-28T09:56:09.519541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # USE VERBOSE=0 for silent, VERBOSE=1 for interactive, VERBOSE=2 for commit\nVERBOSE = 2 if tpu else 1\nDISPLAY_PLOT = True","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:56:09.524093Z","iopub.execute_input":"2022-12-28T09:56:09.524329Z","iopub.status.idle":"2022-12-28T09:56:09.536911Z","shell.execute_reply.started":"2022-12-28T09:56:09.524305Z","shell.execute_reply":"2022-12-28T09:56:09.535399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# PREDICT OOF USING TTA\nprint('Predicting OOF ...')\nds_valid = get_dataset(files_valid,labeled=False,return_image_names=False,augment=True,\n        repeat=True,shuffle=False,batch_size=BATCH_SIZE*4)\nct_valid = count_data_items(files_valid); STEPS = ct_valid/BATCH_SIZE/4\npred = model.predict(ds_valid,steps=STEPS,verbose=VERBOSE)[:ct_valid,] \noof_pred.append( np.mean(pred.reshape((ct_valid,TTA),order='F'),axis=1) ) ","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:56:09.54012Z","iopub.execute_input":"2022-12-28T09:56:09.540425Z","iopub.status.idle":"2022-12-28T09:56:48.389971Z","shell.execute_reply.started":"2022-12-28T09:56:09.540386Z","shell.execute_reply":"2022-12-28T09:56:48.38901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# GET OOF TARGETS AND NAMES\nds_valid = get_dataset(files_valid, augment=False, repeat=False, \n        labeled=True, return_image_names=True)\noof_tar.append( np.array([target.numpy() for img, target in iter(ds_valid.unbatch())]) )\noof_folds.append( np.ones_like(oof_tar[-1],dtype='int8')*1 )\nds = get_dataset(files_valid, augment=False, repeat=False,\n            labeled=False, return_image_names=True)\noof_names.append( np.array([img_name.numpy().decode(\"utf-8\") for img, img_name in iter(ds.unbatch())]))","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:56:48.391128Z","iopub.execute_input":"2022-12-28T09:56:48.391373Z","iopub.status.idle":"2022-12-28T09:57:06.728511Z","shell.execute_reply.started":"2022-12-28T09:56:48.39134Z","shell.execute_reply":"2022-12-28T09:57:06.72785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# PREDICT TEST USING TTA\nprint('Predicting Test')\nds_test = get_dataset(files_test,labeled=False,return_image_names=False,augment=False,\n        repeat=True,shuffle=False,batch_size=BATCH_SIZE*4)\nct_test = count_data_items(files_test); STEPS = TTA * ct_test/BATCH_SIZE/4\npred = model.predict(ds_test,steps=STEPS,verbose=VERBOSE)[:TTA*ct_test,] \npreds[:,0] += np.mean(pred.reshape((ct_test,TTA),order='F'),axis=1) ","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:57:06.729981Z","iopub.execute_input":"2022-12-28T09:57:06.730446Z","iopub.status.idle":"2022-12-28T09:57:27.712267Z","shell.execute_reply.started":"2022-12-28T09:57:06.73041Z","shell.execute_reply":"2022-12-28T09:57:27.710861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"aucx= list(history.history)[1]\n# print(aucx)\nval_aucx= list(history.history)[4]\n# print(val_auc)","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:57:27.7138Z","iopub.execute_input":"2022-12-28T09:57:27.714097Z","iopub.status.idle":"2022-12-28T09:57:27.720484Z","shell.execute_reply.started":"2022-12-28T09:57:27.714062Z","shell.execute_reply":"2022-12-28T09:57:27.718914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" \n# REPORT RESULTS\n#auc = roc_auc_score(oof_tar[-1],oof_pred[-1])\n#oof_val.append(np.max( history.history['val_auc_1'] ))\n#print('#### FOLD %i OOF AUC without TTA = %.3f, with TTA = %.3f'%(fold+1,oof_val[-1],auc))\n#print(auc)","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:57:27.722274Z","iopub.execute_input":"2022-12-28T09:57:27.722503Z","iopub.status.idle":"2022-12-28T09:57:27.733731Z","shell.execute_reply.started":"2022-12-28T09:57:27.722477Z","shell.execute_reply":"2022-12-28T09:57:27.732975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"try:\n    # PLOT TRAINING\n    if DISPLAY_PLOT:\n        plt.figure(figsize=(15,5))\n        plt.plot(np.arange(EPOCH),history.history[aucx],'-o',label='Train AUC',color='#ff7f0e')\n        plt.plot(np.arange(EPOCH),history.history[val_aucx],'-o',label='Val AUC',color='#1f77b4')\n        x = np.argmax( history.history[val_aucx] ); y = np.max( history.history[val_aucx] )\n        xdist = plt.xlim()[1] - plt.xlim()[0]; ydist = plt.ylim()[1] - plt.ylim()[0]\n        plt.scatter(x,y,s=200,color='#1f77b4'); plt.text(x-0.03*xdist,y-0.13*ydist,'max auc\\n%.2f'%y,size=14)\n        plt.ylabel('AUC',size=14); plt.xlabel('Epoch',size=14)\n        plt.legend(loc=2)\n        plt2 = plt.gca().twinx()\n        plt2.plot(np.arange(EPOCH),history.history['loss'],'-o',label='Train Loss',color='#2ca02c')\n        plt2.plot(np.arange(EPOCH),history.history['val_loss'],'-o',label='Val Loss',color='#d62728')\n        x = np.argmin( history.history['val_loss'] ); y = np.min( history.history['val_loss'] )\n        ydist = plt.ylim()[1] - plt.ylim()[0]\n        plt.scatter(x,y,s=200,color='#d62728'); plt.text(x-0.03*xdist,y+0.05*ydist,'min loss',size=14)\n        plt.ylabel('Loss',size=14)\n        plt.title(\"loss and auc\",size=18)\n        plt.legend(loc=3)\n        plt.show()\nexcept:\n    print('no plot today')","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:57:27.734964Z","iopub.execute_input":"2022-12-28T09:57:27.735131Z","iopub.status.idle":"2022-12-28T09:57:28.093741Z","shell.execute_reply.started":"2022-12-28T09:57:27.735109Z","shell.execute_reply":"2022-12-28T09:57:28.092883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # USE VERBOSE=0 for silent, VERBOSE=1 for interactive, VERBOSE=2 for commit\n# VERBOSE = 2 if tpu else 1\n# DISPLAY_PLOT = True\n\n# skf = KFold(n_splits=FOLDS,shuffle=False)\n# oof_pred = []; oof_tar = []; oof_val = []; oof_names = []; oof_folds = [] \n# preds = np.zeros((count_data_items(files_test),1))\n\n# for fold,(idxT,idxV) in enumerate(skf.split(np.arange(5))):\n    \n#     # DISPLAY FOLD INFO\n#     if DEVICE=='TPU':\n#         if tpu: tf.tpu.experimental.initialize_tpu_system(tpu)\n#     print('#'*25); print('#### FOLD',fold+1)\n    \n#     # CREATE TRAIN AND VALIDATION SUBSETS\n#     #files_train = tf.io.gfile.glob(GCS_PATH_STRATIFICATED_modified + '/train*.tfrec')\n#     files_train = tf.io.gfile.glob([GCS_PATH_STRATIFICATED_modified + '/train%.2i*.tfrec'%x for x in idxT])\n#     files_valid = tf.io.gfile.glob([GCS_PATH_STRATIFICATED_modified + '/train%.2i*.tfrec'%x for x in idxV])\n#     files_test = tf.io.gfile.glob(GCS_PATH_STRATIFICATED_modified + '/test*.tfrec')\n    \n#     # BUILD MODEL\n#     K.clear_session()\n#     with strategy.scope():\n#         model = models()\n        \n#     # SAVE BEST MODEL EACH FOLD\n#     sv = tf.keras.callbacks.ModelCheckpoint(\n#         'fold-%i.h5'%fold, monitor='val_loss', verbose=0, save_best_only=True,\n#         save_weights_only=True, mode='min', save_freq='epoch')\n    \n#     es =tf.keras.callbacks.EarlyStopping(\n#         monitor='val_loss',\n#         min_delta=0,\n#         patience=12,\n#         verbose=1,\n#         mode='auto',\n#         )\n   \n#     # TRAIN\n#     print('Training...')\n#     history = model.fit(\n#         get_dataset(files_train, augment=image_augumentation, shuffle=True, repeat=True,\n#                 batch_size=BATCH_SIZE), \n#         epochs=EPOCH, callbacks = [sv,get_lr_callback()], \n#         steps_per_epoch=count_data_items(files_train)/BATCH_SIZE,\n#         validation_data=get_dataset(files_valid,augment=False,shuffle=False,\n#                 repeat=False),\n# #         verbose=VERBOSE\n#     )\n    \n#     print('Loading best model...')\n#     model.load_weights('fold-%i.h5'%fold)\n    \n#     # PREDICT OOF USING TTA\n#     print('Predicting OOF ...')\n#     ds_valid = get_dataset(files_valid,labeled=False,return_image_names=False,augment=True,\n#             repeat=True,shuffle=False,batch_size=BATCH_SIZE*4)\n#     ct_valid = count_data_items(files_valid); STEPS = ct_valid/BATCH_SIZE/4\n#     pred = model.predict(ds_valid,steps=STEPS,verbose=VERBOSE)[:ct_valid,] \n#     oof_pred.append( np.mean(pred.reshape((ct_valid,TTA),order='F'),axis=1) )                 \n   \n    \n#     # GET OOF TARGETS AND NAMES\n#     ds_valid = get_dataset(files_valid, augment=False, repeat=False, \n#             labeled=True, return_image_names=True)\n#     oof_tar.append( np.array([target.numpy() for img, target in iter(ds_valid.unbatch())]) )\n#     oof_folds.append( np.ones_like(oof_tar[-1],dtype='int8')*fold )\n#     ds = get_dataset(files_valid, augment=False, repeat=False,\n#                 labeled=False, return_image_names=True)\n#     oof_names.append( np.array([img_name.numpy().decode(\"utf-8\") for img, img_name in iter(ds.unbatch())]))\n    \n#     # PREDICT TEST USING TTA\n#     print('Predicting Test')\n#     ds_test = get_dataset(files_test,labeled=False,return_image_names=False,augment=False,\n#             repeat=True,shuffle=False,batch_size=BATCH_SIZE*4)\n#     ct_test = count_data_items(files_test); STEPS = TTA * ct_test/BATCH_SIZE/4\n#     pred = model.predict(ds_test,steps=STEPS,verbose=VERBOSE)[:TTA*ct_test,] \n#     preds[:,0] += np.mean(pred.reshape((ct_test,TTA),order='F'),axis=1) / FOLDS\n    \n#     # REPORT RESULTS\n#     auc = roc_auc_score(oof_tar[-1],oof_pred[-1])\n#     oof_val.append(np.max( history.history['val_auc'] ))\n#     print('#### FOLD %i OOF AUC without TTA = %.3f, with TTA = %.3f'%(fold+1,oof_val[-1],auc))\n    \n#     # PLOT TRAINING\n#     if DISPLAY_PLOT:\n#         plt.figure(figsize=(15,5))\n#         plt.plot(np.arange(EPOCH),history.history['auc'],'-o',label='Train AUC',color='#ff7f0e')\n#         plt.plot(np.arange(EPOCH),history.history['val_auc'],'-o',label='Val AUC',color='#1f77b4')\n#         x = np.argmax( history.history['val_auc'] ); y = np.max( history.history['val_auc'] )\n#         xdist = plt.xlim()[1] - plt.xlim()[0]; ydist = plt.ylim()[1] - plt.ylim()[0]\n#         plt.scatter(x,y,s=200,color='#1f77b4'); plt.text(x-0.03*xdist,y-0.13*ydist,'max auc\\n%.2f'%y,size=14)\n#         plt.ylabel('AUC',size=14); plt.xlabel('Epoch',size=14)\n#         plt.legend(loc=2)\n#         plt2 = plt.gca().twinx()\n#         plt2.plot(np.arange(EPOCH),history.history['loss'],'-o',label='Train Loss',color='#2ca02c')\n#         plt2.plot(np.arange(EPOCH),history.history['val_loss'],'-o',label='Val Loss',color='#d62728')\n#         x = np.argmin( history.history['val_loss'] ); y = np.min( history.history['val_loss'] )\n#         ydist = plt.ylim()[1] - plt.ylim()[0]\n#         plt.scatter(x,y,s=200,color='#d62728'); plt.text(x-0.03*xdist,y+0.05*ydist,'min loss',size=14)\n#         plt.ylabel('Loss',size=14)\n#         plt.title('FOLD %i'%(fold+1),size=18)\n#         plt.legend(loc=3)\n#         plt.show()  ","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:57:28.095515Z","iopub.execute_input":"2022-12-28T09:57:28.095754Z","iopub.status.idle":"2022-12-28T09:57:28.103274Z","shell.execute_reply.started":"2022-12-28T09:57:28.095732Z","shell.execute_reply":"2022-12-28T09:57:28.10213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#########################\n#### FOLD 1\nTraining...\nEpoch 1/20\n2022-12-24 13:24:51.187513: I tensorflow/stream_executor/cuda/cuda_dnn.cc:369] Loaded cuDNN version 8005\n\n","metadata":{}},{"cell_type":"code","source":"# # # COMPUTE OVERALL OOF AUC\n# oof = np.concatenate(oof_pred)\n# # true = np.concatenate(oof_tar)\n# # names = np.concatenate(oof_names)\n# # #folds = np.concatenate(oof_folds)\n# # #auc = roc_auc_score(true,oof)\n# # print('Overall OOF AUC with TTA = %.4f'%auc)\n\n# # # SAVE OOF TO DISK\n# df_oof = pd.DataFrame(dict(Id = names, target=true, pred = oof, fold=1))\n# # df_oof.to_csv('oof.csv',index=False)\n# # df_oof.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:57:28.104444Z","iopub.execute_input":"2022-12-28T09:57:28.104643Z","iopub.status.idle":"2022-12-28T09:57:28.119542Z","shell.execute_reply.started":"2022-12-28T09:57:28.104619Z","shell.execute_reply":"2022-12-28T09:57:28.118426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# selectie=df_oof[df_oof.Id.apply(lambda x: len(str(x))<=10)]\n# selectie['diference']=df_oof.target-df_oof.pred\n# selectie.describe()","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:57:28.121026Z","iopub.execute_input":"2022-12-28T09:57:28.121215Z","iopub.status.idle":"2022-12-28T09:57:28.13307Z","shell.execute_reply.started":"2022-12-28T09:57:28.121192Z","shell.execute_reply":"2022-12-28T09:57:28.131626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.read_csv('../input/g2net-detecting-continuous-gravitational-waves/sample_submission.csv')\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:57:28.134405Z","iopub.execute_input":"2022-12-28T09:57:28.134682Z","iopub.status.idle":"2022-12-28T09:57:28.20342Z","shell.execute_reply.started":"2022-12-28T09:57:28.134648Z","shell.execute_reply":"2022-12-28T09:57:28.202399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission['target'] = preds[:,0]\nsubmission = submission.sort_values('id') \nsubmission.to_csv('submission.csv', index=False)\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:57:28.204914Z","iopub.execute_input":"2022-12-28T09:57:28.20512Z","iopub.status.idle":"2022-12-28T09:57:28.25566Z","shell.execute_reply.started":"2022-12-28T09:57:28.205097Z","shell.execute_reply":"2022-12-28T09:57:28.254733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.hist(submission.target,bins=100)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:57:28.257Z","iopub.execute_input":"2022-12-28T09:57:28.257173Z","iopub.status.idle":"2022-12-28T09:57:28.543713Z","shell.execute_reply.started":"2022-12-28T09:57:28.257151Z","shell.execute_reply":"2022-12-28T09:57:28.542811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds = np.zeros((count_data_items(files_test),1))\nlen(preds)","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:57:28.545173Z","iopub.execute_input":"2022-12-28T09:57:28.54541Z","iopub.status.idle":"2022-12-28T09:57:28.552419Z","shell.execute_reply.started":"2022-12-28T09:57:28.545378Z","shell.execute_reply":"2022-12-28T09:57:28.551331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print(f\"Max F1-score in validation dataset = {max(history.history['val_f1_m'])}\")","metadata":{"execution":{"iopub.status.busy":"2022-12-28T09:57:28.554107Z","iopub.execute_input":"2022-12-28T09:57:28.554426Z","iopub.status.idle":"2022-12-28T09:57:28.565244Z","shell.execute_reply.started":"2022-12-28T09:57:28.554392Z","shell.execute_reply":"2022-12-28T09:57:28.563671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you see there's a clear increase in F1-Score from 87% to 91% and can be increased with further fine-tuning and Image Augmenattion.  \nMy motive was to get you familiarised with this new concept.  \nI hoped you liked it and learned from it.  \n**Happy Learning**","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}