{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Update (version 9):\n\nThis version (9) incorporates the latest updates to [ROHITH's notebbok](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place) (version 5)","metadata":{"id":"yn6s3ZLRh6k6"}},{"cell_type":"markdown","source":"# CTC on TPU\n\nThis modification of [ROHITH INGILELA's CTC notebook](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place) can run on TPU. I used [another implementation](https://github.com/alexeytochin/tf_seq2seq_losses) for CTC since TensorFlow implementation does not work on Kaggle's TPU. This implementation also does not run on Kaggle's TPU but was much easier to debug since 'This is a pure Python/TensorFlow implementation. We do not have to build or compile any C++/CUDA stuff.' as stated on the project page. This makes a HUGE difference, as the original TensorFlow implementation has some messy code that relies directly on TensorFlow's ops behind the curtains. When you dig deep enough, it does not work on Kaggle TPU due to a bug related to some TensorFlow ops. Then you have to start debugging some ugly C-related code (don't ask how I know 😉)...well, this is why I eventually searched for another route, haha. Anyway, As I said, this implementation also did not work at first on Kaggle's TPU, but after some debugging that resulted in two small changes to the source code, I could make it work. All this work was NOT easy by any means...took me about two days of trial and error. I would appreciate your votes🙏\n\n**RUNTIME:** With Kaggle TPU and some optimization to data loading (namely, load all the data to the RAM before converting it to a dataset instead of loading in batches from scratch each iteration, since TPU notebooks have more than enough rum), result in a MUCH faster runtime. ~47 min for the entire notebook (50 epochs) as opposed to the original notebook with ~6 hours for the whole notebook. More than seven times faster.\n\n**SUBMITTING:** To submit the TPU model, save it to h5 with model.save_weights, then build the model on a GPU notebook and load the weights with model.load_weights. If you submit directly from a TPU notebook, you will have to solve some problems that are not worth the time (if possible to solve at all).\n\n**P.S.** The original notebook accidentally trained on the validation score, which is, of course, a big NO-NO...and this is why it achieved a very low CTC validation score (~6). I fixed it and also added two callback functions, one for calculating the validation set Levenshtein distance, since this is our metric, and one for saving the models during the training (I saved the weights every five epochs, you can change that)","metadata":{"id":"HqzYGsc5h6k8"}},{"cell_type":"code","source":"###---- Environment config ----###\n\n# MACHINE = \"COLAB\"\n# device = \"TPU\"\n\nMACHINE = \"KAGGLE\"\n# device = \"TPU-VM\"\ndevice = \"GPU\"\n\n# MACHINE = \"JAYOO_PC\"\n# device = \"GPU\"\n\n\n# DEBUG = True\nDEBUG = False\nif DEBUG == True:\n    print(\"IN DEBUG MODE\")\n    device = \"CPU\"\n\n\n# Set root directory\nif MACHINE == \"JAYOO_PC\":\n    ROOT = '/jayoo'  # local\nelif MACHINE == \"COLAB\":\n    ROOT = './drive/MyDrive/colab_env'\n    from google.colab import drive\n    drive.mount('/content/drive')\n    !pwd\nelse:\n    ROOT = ''  # Kaggle\n\nprint(f\"Machine: {MACHINE}, device: {device}, root: {ROOT}\")\n\nimport multiprocessing\nprint(multiprocessing.cpu_count())","metadata":{"executionInfo":{"elapsed":13297,"status":"ok","timestamp":1692563992703,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"},"user_tz":360},"id":"4qFVN4vxh6k9","outputId":"ef0cbd17-bb40-46d1-82db-d30b8c632412","execution":{"iopub.status.busy":"2023-08-24T21:10:34.907997Z","iopub.execute_input":"2023-08-24T21:10:34.90837Z","iopub.status.idle":"2023-08-24T21:10:34.925535Z","shell.execute_reply.started":"2023-08-24T21:10:34.908337Z","shell.execute_reply":"2023-08-24T21:10:34.924313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\nimport json\nimport math\nimport pickle\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\n!pip install -q tensorflow-addons\nimport tensorflow_addons as tfa\nimport tensorflow.keras.mixed_precision as mixed_precision\n\nimport os\nimport time\nimport sys\nimport random\nimport cv2\nimport glob\nimport gc\nimport datetime\n\n!pip install cached-property\nfrom cached_property import cached_property\nfrom shutil import copyfile\n\n!pip install fastparquet\nimport fastparquet\n\n!pip install Levenshtein\nimport Levenshtein as lev","metadata":{"executionInfo":{"elapsed":52317,"status":"ok","timestamp":1692564045018,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"},"user_tz":360},"id":"xmpgv9Fjh6k-","outputId":"749444a5-5a79-4e7d-def0-6591370c91eb","execution":{"iopub.status.busy":"2023-08-24T21:10:34.928126Z","iopub.execute_input":"2023-08-24T21:10:34.928488Z","iopub.status.idle":"2023-08-24T21:11:31.162163Z","shell.execute_reply.started":"2023-08-24T21:10:34.928449Z","shell.execute_reply":"2023-08-24T21:11:31.160895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Configure Strategy. Assume TPU...if not set default for GPU\nif \"TPU\" in device:\n        print(\"connecting to TPU...\")\n        if device == 'TPU-VM':  # kaggle\n            tpu = 'local'\n            tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=tpu)\n            strategy = tf.distribute.TPUStrategy(tpu)\n        if device == 'TPU':  # colab\n            tpu = tf.distribute.cluster_resolver.TPUClusterResolver()  # TPU detection\n            tf.config.experimental_connect_to_cluster(tpu)\n            tf.tpu.experimental.initialize_tpu_system(tpu)\n            strategy = tf.distribute.TPUStrategy(tpu)\n\n        IS_TPU = True\n\nif device == \"GPU\"  or device==\"CPU\":\n        IS_TPU = False\n        gpus = tf.config.experimental.list_physical_devices('GPU')\n        try:\n            for gpu in gpus:\n                tf.config.experimental.set_memory_growth(gpu, True)\n        except:\n            pass\n        ngpu = len(gpus)\n        if ngpu>1:\n            print(\"Using multi GPU\")\n            strategy = tf.distribute.MirroredStrategy()\n        elif ngpu==1:\n            print(\"Using single GPU\")\n            strategy = tf.distribute.get_strategy()\n        else:\n            print(\"Using CPU\")\n            strategy = tf.distribute.get_strategy()\n\nif device == \"GPU\":\n        print(\"Num GPUs Available: \", ngpu)\n\nAUTO = tf.data.experimental.AUTOTUNE\nREPLICAS = strategy.num_replicas_in_sync\nprint(f'REPLICAS: {REPLICAS}')\n\n# tpu = None\n# try:\n#     tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=\"local\") # \"local\" for 1VM TPU\n#     strategy = tf.distribute.TPUStrategy(tpu)\n#     print(\"on TPU\")\n#     print(\"REPLICAS: \", strategy.num_replicas_in_sync)\n# except:\n#     strategy = tf.distribute.get_strategy()","metadata":{"executionInfo":{"elapsed":11766,"status":"ok","timestamp":1692564056782,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"},"user_tz":360},"id":"LEhtNKrXh6k_","outputId":"e41d47db-4d1a-4de2-d805-a10971961b06","execution":{"iopub.status.busy":"2023-08-24T21:11:31.165153Z","iopub.execute_input":"2023-08-24T21:11:31.166848Z","iopub.status.idle":"2023-08-24T21:11:31.365842Z","shell.execute_reply.started":"2023-08-24T21:11:31.166817Z","shell.execute_reply":"2023-08-24T21:11:31.364879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if MACHINE == \"KAGGLE\":\n    # copy our file into the working directory (make sure it has .py suffix)\n    copyfile(src = \"/kaggle/input/ctc-tpu/CTC_TPU.py\", dst = \"/kaggle/working//CTC_TPU.py\")\nelif MACHINE == \"COLAB\":\n    copyfile(src = ROOT+\"/CTC_TPU.py\", dst = \"/content/CTC_TPU.py\")\n\n# import all our functions\nfrom CTC_TPU import classic_ctc_loss","metadata":{"id":"ylnbeZpVh6k_","executionInfo":{"status":"ok","timestamp":1692564057000,"user_tz":360,"elapsed":227,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"execution":{"iopub.status.busy":"2023-08-24T21:11:31.367161Z","iopub.execute_input":"2023-08-24T21:11:31.367562Z","iopub.status.idle":"2023-08-24T21:11:33.632742Z","shell.execute_reply.started":"2023-08-24T21:11:31.367526Z","shell.execute_reply":"2023-08-24T21:11:33.631767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# If on Colab, get files from GCS\nif MACHINE == \"COLAB\":\n    GCS_PATH = {\n            'ASLFR': 'gs://kds-fccc385146c7904b48b83ddab53511fdafd22d8306eab412927f000d',\n            'TFR': 'gs://kds-c3aa6fb0c44cdeb15556c4920117ded3e78198299e00daa21d7bc2ad',\n            'MINE': 'gs://kds-185cc50c32a5ca0d5ab0f1f6449486d1afe849a17c6ad6c602c8de5f',\n            }\n\n    COMPETITION_PATH = GCS_PATH['ASLFR']\n    TFR_DATA_PATH = GCS_PATH['TFR']\n\n    # copy files from GCS to colab folders\n    !gsutil cp {COMPETITION_PATH}/train.csv {ROOT}/kaggle/input/asl-fingerspelling\n    !gsutil cp {COMPETITION_PATH}/character_to_prediction_index.json {ROOT}/kaggle/input/asl-fingerspelling\n\n    !gsutil cp {TFR_DATA_PATH}/mean_std/rh_mean.npy {ROOT}/kaggle/input/aslfr-dataset-tfrecords/mean_std\n    !gsutil cp {TFR_DATA_PATH}/mean_std/lh_mean.npy {ROOT}/kaggle/input/aslfr-dataset-tfrecords/mean_std\n    !gsutil cp {TFR_DATA_PATH}/mean_std/rp_mean.npy {ROOT}/kaggle/input/aslfr-dataset-tfrecords/mean_std\n    !gsutil cp {TFR_DATA_PATH}/mean_std/lp_mean.npy {ROOT}/kaggle/input/aslfr-dataset-tfrecords/mean_std\n    !gsutil cp {TFR_DATA_PATH}/mean_std/lip_mean.npy {ROOT}/kaggle/input/aslfr-dataset-tfrecords/mean_std\n\n    !gsutil cp {TFR_DATA_PATH}/mean_std/rh_std.npy {ROOT}/kaggle/input/aslfr-dataset-tfrecords/mean_std\n    !gsutil cp {TFR_DATA_PATH}/mean_std/lh_std.npy {ROOT}/kaggle/input/aslfr-dataset-tfrecords/mean_std\n    !gsutil cp {TFR_DATA_PATH}/mean_std/rp_std.npy {ROOT}/kaggle/input/aslfr-dataset-tfrecords/mean_std\n    !gsutil cp {TFR_DATA_PATH}/mean_std/lp_std.npy {ROOT}/kaggle/input/aslfr-dataset-tfrecords/mean_std\n    !gsutil cp {TFR_DATA_PATH}/mean_std/lip_std.npy {ROOT}/kaggle/input/aslfr-dataset-tfrecords/mean_std\n\n# RHM = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/rh_mean.npy\")\n# LHM = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/lh_mean.npy\")\n# RPM = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/rp_mean.npy\")\n# LPM = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/lp_mean.npy\")\n# LIPM = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/lip_mean.npy\")\n\n# RHS = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/rh_std.npy\")\n# LHS = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/lh_std.npy\")\n# RPS = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/rp_std.npy\")\n# LPS = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/lp_std.npy\")\n# LIPS = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/lip_std.npy\")\n","metadata":{"executionInfo":{"elapsed":50949,"status":"ok","timestamp":1692564107948,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"},"user_tz":360},"id":"POrEOVP4k4Mo","outputId":"a5991ebc-bcad-4e79-d080-6533678221aa","execution":{"iopub.status.busy":"2023-08-24T21:11:33.635982Z","iopub.execute_input":"2023-08-24T21:11:33.63634Z","iopub.status.idle":"2023-08-24T21:11:33.687724Z","shell.execute_reply.started":"2023-08-24T21:11:33.636306Z","shell.execute_reply":"2023-08-24T21:11:33.686718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with open (ROOT+\"/kaggle/input/asl-fingerspelling/character_to_prediction_index.json\", \"r\") as f:\n    char_to_num = json.load(f)\n\npad_token = '^'\npad_token_idx = 59\n\nchar_to_num[pad_token] = pad_token_idx\n\nnum_to_char = {j:i for i,j in char_to_num.items()}\ndf = pd.read_csv(ROOT+'/kaggle/input/asl-fingerspelling/train.csv')\n\nLIP = [\n    61, 185, 40, 39, 37, 0, 267, 269, 270, 409,\n    291, 146, 91, 181, 84, 17, 314, 405, 321, 375,\n    78, 191, 80, 81, 82, 13, 312, 311, 310, 415,\n    95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n]\nLPOSE = [13, 15, 17, 19, 21]\nRPOSE = [14, 16, 18, 20, 22]\nPOSE = LPOSE + RPOSE\n\nX = [f'x_right_hand_{i}' for i in range(21)] + [f'x_left_hand_{i}' for i in range(21)] + [f'x_pose_{i}' for i in POSE] + [f'x_face_{i}' for i in LIP]\nY = [f'y_right_hand_{i}' for i in range(21)] + [f'y_left_hand_{i}' for i in range(21)] + [f'y_pose_{i}' for i in POSE] + [f'y_face_{i}' for i in LIP]\nZ = [f'z_right_hand_{i}' for i in range(21)] + [f'z_left_hand_{i}' for i in range(21)] + [f'z_pose_{i}' for i in POSE] + [f'z_face_{i}' for i in LIP]\n\nSEL_COLS = X + Y + Z\n# FRAME_LEN = 256\nMAX_PHRASE_LENGTH = 64\n\nLIP_IDX_X   = [i for i, col in enumerate(SEL_COLS)  if  \"face\" in col and \"x\" in col]\nRHAND_IDX_X = [i for i, col in enumerate(SEL_COLS)  if \"right\" in col and \"x\" in col]\nLHAND_IDX_X = [i for i, col in enumerate(SEL_COLS)  if  \"left\" in col and \"x\" in col]\nRPOSE_IDX_X = [i for i, col in enumerate(SEL_COLS)  if  \"pose\" in col and int(col[-2:]) in RPOSE and \"x\" in col]\nLPOSE_IDX_X = [i for i, col in enumerate(SEL_COLS)  if  \"pose\" in col and int(col[-2:]) in LPOSE and \"x\" in col]\n\nLIP_IDX_Y   = [i for i, col in enumerate(SEL_COLS)  if  \"face\" in col and \"y\" in col]\nRHAND_IDX_Y = [i for i, col in enumerate(SEL_COLS)  if \"right\" in col and \"y\" in col]\nLHAND_IDX_Y = [i for i, col in enumerate(SEL_COLS)  if  \"left\" in col and \"y\" in col]\nRPOSE_IDX_Y = [i for i, col in enumerate(SEL_COLS)  if  \"pose\" in col and int(col[-2:]) in RPOSE and \"y\" in col]\nLPOSE_IDX_Y = [i for i, col in enumerate(SEL_COLS)  if  \"pose\" in col and int(col[-2:]) in LPOSE and \"y\" in col]\n\nLIP_IDX_Z   = [i for i, col in enumerate(SEL_COLS)  if  \"face\" in col and \"z\" in col]\nRHAND_IDX_Z = [i for i, col in enumerate(SEL_COLS)  if \"right\" in col and \"z\" in col]\nLHAND_IDX_Z = [i for i, col in enumerate(SEL_COLS)  if  \"left\" in col and \"z\" in col]\nRPOSE_IDX_Z = [i for i, col in enumerate(SEL_COLS)  if  \"pose\" in col and int(col[-2:]) in RPOSE and \"z\" in col]\nLPOSE_IDX_Z = [i for i, col in enumerate(SEL_COLS)  if  \"pose\" in col and int(col[-2:]) in LPOSE and \"z\" in col]\n\nRHM = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/rh_mean.npy\")\nLHM = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/lh_mean.npy\")\nRPM = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/rp_mean.npy\")\nLPM = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/lp_mean.npy\")\nLIPM = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/lip_mean.npy\")\n\nRHS = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/rh_std.npy\")\nLHS = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/lh_std.npy\")\nRPS = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/rp_std.npy\")\nLPS = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/lp_std.npy\")\nLIPS = np.load(ROOT+\"/kaggle/input/aslfr-dataset-tfrecords/mean_std/lip_std.npy\")","metadata":{"id":"vEmE5paUh6k_","executionInfo":{"status":"ok","timestamp":1692564110130,"user_tz":360,"elapsed":2191,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"execution":{"iopub.status.busy":"2023-08-24T21:11:33.690372Z","iopub.execute_input":"2023-08-24T21:11:33.690704Z","iopub.status.idle":"2023-08-24T21:11:33.903527Z","shell.execute_reply.started":"2023-08-24T21:11:33.69068Z","shell.execute_reply":"2023-08-24T21:11:33.902512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_relevant_data_subset(pq_path):\n    return pd.read_parquet(pq_path, columns=SEL_COLS)\n\nfile_id = df.file_id.iloc[0]\ninpdir = ROOT+\"/kaggle/input/asl-fingerspelling/train_landmarks\"\npqfile = f\"{inpdir}/{file_id}.parquet\"\nseq_refs = df.loc[df.file_id == file_id]\nseqs = load_relevant_data_subset(pqfile)\n\nseq_id = seq_refs.sequence_id.iloc[0]\nframes = seqs.iloc[seqs.index == seq_id]\nphrase = str(df.loc[df.sequence_id == seq_id].phrase.iloc[0])","metadata":{"id":"xZq3dxU3h6lA","executionInfo":{"status":"ok","timestamp":1692564110131,"user_tz":360,"elapsed":5,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"execution":{"iopub.status.busy":"2023-08-24T21:11:33.905197Z","iopub.execute_input":"2023-08-24T21:11:33.905919Z","iopub.status.idle":"2023-08-24T21:11:38.272824Z","shell.execute_reply.started":"2023-08-24T21:11:33.905883Z","shell.execute_reply":"2023-08-24T21:11:38.271843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"@tf.function()\ndef resize_pad(x):\n    if tf.shape(x)[0] < FRAME_LEN:\n        x = tf.pad(x, ([[0, FRAME_LEN-tf.shape(x)[0]], [0, 0], [0, 0]]), constant_values=float(\"NaN\"))\n    else:\n        x = tf.image.resize(x, (FRAME_LEN, tf.shape(x)[1]))\n    return x\n\n@tf.function(jit_compile=True)\ndef pre_process0(x):\n    lip_x = tf.gather(x, LIP_IDX_X, axis=1)\n    lip_y = tf.gather(x, LIP_IDX_Y, axis=1)\n    lip_z = tf.gather(x, LIP_IDX_Z, axis=1)\n\n    rhand_x = tf.gather(x, RHAND_IDX_X, axis=1)\n    rhand_y = tf.gather(x, RHAND_IDX_Y, axis=1)\n    rhand_z = tf.gather(x, RHAND_IDX_Z, axis=1)\n\n    lhand_x = tf.gather(x, LHAND_IDX_X, axis=1)\n    lhand_y = tf.gather(x, LHAND_IDX_Y, axis=1)\n    lhand_z = tf.gather(x, LHAND_IDX_Z, axis=1)\n\n    rpose_x = tf.gather(x, RPOSE_IDX_X, axis=1)\n    rpose_y = tf.gather(x, RPOSE_IDX_Y, axis=1)\n    rpose_z = tf.gather(x, RPOSE_IDX_Z, axis=1)\n\n    lpose_x = tf.gather(x, LPOSE_IDX_X, axis=1)\n    lpose_y = tf.gather(x, LPOSE_IDX_Y, axis=1)\n    lpose_z = tf.gather(x, LPOSE_IDX_Z, axis=1)\n\n    lip   = tf.concat([lip_x[..., tf.newaxis], lip_y[..., tf.newaxis], lip_z[..., tf.newaxis]], axis=-1)\n    rhand = tf.concat([rhand_x[..., tf.newaxis], rhand_y[..., tf.newaxis], rhand_z[..., tf.newaxis]], axis=-1)\n    lhand = tf.concat([lhand_x[..., tf.newaxis], lhand_y[..., tf.newaxis], lhand_z[..., tf.newaxis]], axis=-1)\n    rpose = tf.concat([rpose_x[..., tf.newaxis], rpose_y[..., tf.newaxis], rpose_z[..., tf.newaxis]], axis=-1)\n    lpose = tf.concat([lpose_x[..., tf.newaxis], lpose_y[..., tf.newaxis], lpose_z[..., tf.newaxis]], axis=-1)\n\n    # Don't remove nan hand frames\n    # hand = tf.concat([rhand, lhand], axis=1)\n    # hand = tf.where(tf.math.is_nan(hand), 0.0, hand)\n    # mask = tf.math.not_equal(tf.reduce_sum(hand, axis=[1, 2]), 0.0)\n    # lip = lip[mask]\n    # rhand = rhand[mask]\n    # lhand = lhand[mask]\n    # rpose = rpose[mask]\n    # lpose = lpose[mask]\n\n    return lip, rhand, lhand, rpose, lpose\n\n\n# motion features:\ndef get_dx(x):\n    dx = tf.cond(tf.shape(x)[0]>1, lambda:tf.pad(x[1:] - x[:-1], [[0,1],[0,0],[0,0]]), lambda:tf.zeros_like(x))\n    return dx\n\ndef get_dx2(x):\n    dx2 = tf.cond(tf.shape(x)[0]>2, lambda:tf.pad(x[2:] - x[:-2], [[0,2],[0,0],[0,0]]), lambda:tf.zeros_like(x))\n    return dx2\n\n@tf.function()\ndef pre_process1(lip, rhand, lhand, rpose, lpose):\n    lip   = (resize_pad(lip) - LIPM) / LIPS\n    rhand = (resize_pad(rhand) - RHM) / RHS\n    lhand = (resize_pad(lhand) - LHM) / LHS\n    rpose = (resize_pad(rpose) - RPM) / RPS\n    lpose = (resize_pad(lpose) - LPM) / LPS\n\n    # Try without normalization\n    # lip   = resize_pad(lip)\n    # rhand = resize_pad(rhand)\n    # lhand = resize_pad(lhand)\n    # rpose = resize_pad(rpose)\n    # lpose = resize_pad(lpose)\n\n    x = tf.concat([lip, rhand, lhand, rpose, lpose], axis=1)\n    \n     # add motion features\n#     x = x[...,:2]  # use xy points only\n#     dx = get_dx(x)\n#     dx2 = get_dx2(x)\n#     x = tf.concat([x, dx, dx2], axis=1)\n    \n    s = tf.shape(x)\n    x = tf.reshape(x, (s[0], s[1]*s[2]))\n    x = tf.where(tf.math.is_nan(x), 0.0, x)\n    return x\n\nCHANNELS = 276\n# CHANNELS = 552 #362\n\n#This fail on TPU\n'''\npre0 = pre_process0(frames)\npre1 = pre_process1(*pre0)\nINPUT_SHAPE = list(pre1.shape)\nprint(INPUT_SHAPE)\npre1\n'''\n","metadata":{"executionInfo":{"elapsed":5,"status":"ok","timestamp":1692564110131,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"},"user_tz":360},"id":"2nT560iIh6lA","outputId":"1fcefd96-07dd-4ebe-fec4-9607726d5923","execution":{"iopub.status.busy":"2023-08-24T21:11:38.275304Z","iopub.execute_input":"2023-08-24T21:11:38.276Z","iopub.status.idle":"2023-08-24T21:11:38.307028Z","shell.execute_reply.started":"2023-08-24T21:11:38.275965Z","shell.execute_reply":"2023-08-24T21:11:38.306293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Augmentation\ndef resample(x, rate=1.0):\n    length = tf.shape(x)[0]\n    new_size = tf.cast(rate*tf.cast(length,tf.float32), tf.int32)\n    new_size = tf.math.maximum(new_size, 1)\n    x = tf.image.resize(x, (new_size, tf.shape(x)[1]),'bilinear')\n    return x\n\ndef resample_all(lip, rhand, lhand, rpose, lpose, size=(0.8,1.2)):\n    rate = tf.random.uniform((), size[0], size[1])\n    length = tf.shape(rhand)[0]\n    new_size = tf.cast(rate*tf.cast(length,tf.float32), tf.int32)\n    new_size = tf.math.maximum(new_size, 1)\n\n    lip = tf.image.resize(lip, (new_size, tf.shape(lip)[1]), 'bilinear')  #resample(lip, rate)\n    rhand = tf.image.resize(rhand, (new_size, tf.shape(rhand)[1]), 'bilinear')  #resample(rhand, rate)\n    lhand = tf.image.resize(lhand, (new_size, tf.shape(lhand)[1]), 'bilinear')  #resample(lhand, rate)\n    rpose = tf.image.resize(rpose, (new_size, tf.shape(rpose)[1]), 'bilinear')  #resample(rpose, rate)\n    lpose = tf.image.resize(lpose, (new_size, tf.shape(lpose)[1]), 'bilinear')  #resample(lpose, rate)\n\n    return lip, rhand, lhand, rpose, lpose\n\ndef flip_x(x):\n    x,y,z = tf.unstack(x, axis=-1)\n    x = 1-x\n    new_x = tf.stack([x,y,z], -1)\n#     new_x = tf.transpose(new_x, [1,0,2])  # point, frame, xyz\n    return new_x\n\ndef flip_lr(lip, rhand, lhand, rpose, lpose):\n    lip = flip_x(lip)\n    rhand = flip_x(rhand)\n    lhand = flip_x(lhand)\n    rpose = flip_x(rpose)\n    lpose = flip_x(lpose)\n\n    # swap left and right hand\n    temp_hand = rhand\n    rhand = lhand\n    lhand = temp_hand\n\n    # swap left and right pose\n    temp_pose = rpose\n    rpose = lpose\n    lpose = temp_pose\n\n    return lip, rhand, lhand, rpose, lpose\n\ndef augment_fn(lip, rhand, lhand, rpose, lpose, always=False):\n    # random resize\n    lip, rhand, lhand, rpose, lpose = resample_all(lip, rhand, lhand, rpose, lpose, size=(0.8, 1.2))\n\n    # horizontal flip\n    if tf.random.uniform(())<0.5 or always:\n        lip, rhand, lhand, rpose, lpose = flip_lr(lip, rhand, lhand, rpose, lpose)\n\n    return lip, rhand, lhand, rpose, lpose\n","metadata":{"id":"aIY-1ACBlda7","executionInfo":{"status":"ok","timestamp":1692564110360,"user_tz":360,"elapsed":232,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"execution":{"iopub.status.busy":"2023-08-24T21:11:38.308948Z","iopub.execute_input":"2023-08-24T21:11:38.309574Z","iopub.status.idle":"2023-08-24T21:11:38.325138Z","shell.execute_reply.started":"2023-08-24T21:11:38.309541Z","shell.execute_reply":"2023-08-24T21:11:38.324526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def decode_fn(record_bytes):\n    schema = {\n        \"lip\": tf.io.VarLenFeature(tf.float32),\n        \"rhand\": tf.io.VarLenFeature(tf.float32),\n        \"lhand\": tf.io.VarLenFeature(tf.float32),\n        \"rpose\": tf.io.VarLenFeature(tf.float32),\n        \"lpose\": tf.io.VarLenFeature(tf.float32),\n        \"phrase\": tf.io.VarLenFeature(tf.int64)\n    }\n    x = tf.io.parse_single_example(record_bytes, schema)\n\n    lip = tf.reshape(tf.sparse.to_dense(x[\"lip\"]), (-1, 40, 3))\n    rhand = tf.reshape(tf.sparse.to_dense(x[\"rhand\"]), (-1, 21, 3))\n    lhand = tf.reshape(tf.sparse.to_dense(x[\"lhand\"]), (-1, 21, 3))\n    rpose = tf.reshape(tf.sparse.to_dense(x[\"rpose\"]), (-1, 5, 3))\n    lpose = tf.reshape(tf.sparse.to_dense(x[\"lpose\"]), (-1, 5, 3))\n    phrase = tf.sparse.to_dense(x[\"phrase\"])\n\n    return lip, rhand, lhand, rpose, lpose, phrase\n\ndef pre_process_fn(lip, rhand, lhand, rpose, lpose, phrase, augment=False):\n    phrase = tf.cast(phrase, dtype=tf.int32)\n    phrase = tf.pad(phrase, [[0, MAX_PHRASE_LENGTH-tf.shape(phrase)[0]]], constant_values=pad_token_idx)\n\n    if augment==True:\n        lip, rhand, lhand, rpose, lpose = augment_fn(lip, rhand, lhand, rpose, lpose)\n\n    return pre_process1(lip, rhand, lhand, rpose, lpose), phrase\n\n\nif MACHINE == \"COLAB\":\n    # tffiles = [f\"{GCS_PATH['TFR']}/tfds/{file_id}.tfrecord\" for file_id in df.file_id.unique()]\n    data_dir = f\"{GCS_PATH['MINE']}/tfrecords/train/tfds\"\nelif MACHINE == \"KAGGLE\":\n    data_dir = f\"/kaggle/input/my-dataset/tfrecords/tfrecords/train/tfds\"\nelse: # GPU\n    # tffiles = [f\"{ROOT}/kaggle/input/aslfr-dataset-tfrecords/tfds/{file_id}.tfrecord\" for file_id in df.file_id.unique()]\n    data_dir = ROOT+'/kaggle/input/my-dataset/train/tfds'\n\n","metadata":{"id":"FLEaZmEvh6lB","executionInfo":{"status":"ok","timestamp":1692564110360,"user_tz":360,"elapsed":3,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"execution":{"iopub.status.busy":"2023-08-24T21:11:38.32639Z","iopub.execute_input":"2023-08-24T21:11:38.3269Z","iopub.status.idle":"2023-08-24T21:11:38.340899Z","shell.execute_reply.started":"2023-08-24T21:11:38.326869Z","shell.execute_reply":"2023-08-24T21:11:38.340297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model","metadata":{"id":"q87kVwzndMuU"}},{"cell_type":"code","source":"#Copied from previous comp 1st place model: https://www.kaggle.com/code/hoyso48/1st-place-solution-training\nclass ECA(tf.keras.layers.Layer):\n    def __init__(self, kernel_size=5, **kwargs):\n        super().__init__(**kwargs)\n        self.supports_masking = True\n        self.kernel_size = kernel_size\n        self.conv = tf.keras.layers.Conv1D(1, kernel_size=kernel_size, strides=1, padding=\"same\", use_bias=False)\n\n    def call(self, inputs, mask=None):\n        nn = tf.keras.layers.GlobalAveragePooling1D()(inputs, mask=mask)\n        nn = tf.expand_dims(nn, -1)\n        nn = self.conv(nn)\n        nn = tf.squeeze(nn, -1)\n        nn = tf.nn.sigmoid(nn)\n        nn = nn[:,None,:]\n        return inputs * nn\n\nclass LateDropout(tf.keras.layers.Layer):\n    def __init__(self, rate, noise_shape=None, start_step=0, **kwargs):\n        super().__init__(**kwargs)\n        self.supports_masking = True\n        self.rate = rate\n        self.start_step = start_step\n        self.dropout = tf.keras.layers.Dropout(rate, noise_shape=noise_shape)\n\n    def build(self, input_shape):\n        super().build(input_shape)\n        agg = tf.VariableAggregation.ONLY_FIRST_REPLICA\n        self._train_counter = tf.Variable(0, dtype=\"int64\", aggregation=agg, trainable=False)\n\n    def call(self, inputs, training=False):\n        x = tf.cond(self._train_counter < self.start_step, lambda:inputs, lambda:self.dropout(inputs, training=training))\n        if training:\n            self._train_counter.assign_add(1)\n        return x\n\nclass CausalDWConv1D(tf.keras.layers.Layer):\n    def __init__(self,\n        kernel_size=17,\n        dilation_rate=1,\n        use_bias=False,\n        depthwise_initializer='glorot_uniform',\n        name='', **kwargs):\n        super().__init__(name=name,**kwargs)\n        self.causal_pad = tf.keras.layers.ZeroPadding1D((dilation_rate*(kernel_size-1),0),name=name + '_pad')\n        self.dw_conv = tf.keras.layers.DepthwiseConv1D(\n                            kernel_size,\n                            strides=1,\n                            dilation_rate=dilation_rate,\n                            padding='valid',\n                            use_bias=use_bias,\n                            depthwise_initializer=depthwise_initializer,\n                            name=name + '_dwconv')\n        self.supports_masking = True\n\n    def call(self, inputs):\n        x = self.causal_pad(inputs)\n        x = self.dw_conv(x)\n        return x\n\ndef Conv1DBlock(channel_size,\n          kernel_size,\n          dilation_rate=1,\n          drop_rate=0.0,\n          expand_ratio=2,\n          se_ratio=0.25,\n          activation='swish',\n          name=None):\n    '''\n    efficient conv1d block, @hoyso48\n    '''\n    if name is None:\n        name = str(tf.keras.backend.get_uid(\"mbblock\"))\n    # Expansion phase\n    def apply(inputs):\n        channels_in = tf.keras.backend.int_shape(inputs)[-1]\n        channels_expand = channels_in * expand_ratio\n\n        skip = inputs\n\n        x = tf.keras.layers.Dense(\n            channels_expand,\n            use_bias=True,\n            activation=activation,\n            name=name + '_expand_conv')(inputs)\n\n        # Depthwise Convolution\n        x = CausalDWConv1D(kernel_size,\n            dilation_rate=dilation_rate,\n            use_bias=False,\n            name=name + '_dwconv')(x)\n\n        x = tf.keras.layers.BatchNormalization(momentum=0.95, name=name + '_bn')(x)\n\n        x  = ECA()(x)\n\n        x = tf.keras.layers.Dense(\n            channel_size,\n            use_bias=True,\n            name=name + '_project_conv')(x)\n\n        if drop_rate > 0:\n            x = tf.keras.layers.Dropout(drop_rate, noise_shape=(None,1,1), name=name + '_drop')(x)\n\n        if (channels_in == channel_size):\n            x = tf.keras.layers.add([x, skip], name=name + '_add')\n        return x\n\n    return apply\n\nclass MultiHeadSelfAttention(tf.keras.layers.Layer):\n    def __init__(self, dim=256, num_heads=4, dropout=0, **kwargs):\n        super().__init__(**kwargs)\n        self.dim = dim\n        self.scale = self.dim ** -0.5\n        self.num_heads = num_heads\n        self.qkv = tf.keras.layers.Dense(3 * dim, use_bias=False)\n        self.drop1 = tf.keras.layers.Dropout(dropout)\n        self.proj = tf.keras.layers.Dense(dim, use_bias=False)\n        self.supports_masking = True\n\n    def call(self, inputs, mask=None):\n        qkv = self.qkv(inputs)\n        qkv = tf.keras.layers.Permute((2, 1, 3))(tf.keras.layers.Reshape((-1, self.num_heads, self.dim * 3 // self.num_heads))(qkv))\n        q, k, v = tf.split(qkv, [self.dim // self.num_heads] * 3, axis=-1)\n\n        attn = tf.matmul(q, k, transpose_b=True) * self.scale\n\n        if mask is not None:\n            mask = mask[:, None, None, :]\n\n        attn = tf.keras.layers.Softmax(axis=-1)(attn, mask=mask)\n        attn = self.drop1(attn)\n\n        x = attn @ v\n        x = tf.keras.layers.Reshape((-1, self.dim))(tf.keras.layers.Permute((2, 1, 3))(x))\n        x = self.proj(x)\n        return x\n\n\ndef TransformerBlock(dim=256, num_heads=6, expand=4, attn_dropout=0.2, drop_rate=0.2, activation='swish'):\n    def apply(inputs):\n        x = inputs\n        x = tf.keras.layers.LayerNormalization(epsilon=1e-6)(x)\n        x = MultiHeadSelfAttention(dim=dim,num_heads=num_heads,dropout=attn_dropout)(x)\n        x = tf.keras.layers.Dropout(drop_rate, noise_shape=(None,1,1))(x)\n        x = tf.keras.layers.Add()([inputs, x])\n        attn_out = x\n\n        x = tf.keras.layers.LayerNormalization(epsilon=1e-6)(x)\n        x = tf.keras.layers.Dense(dim*expand, use_bias=False, activation=activation)(x)\n        x = tf.keras.layers.Dense(dim, use_bias=False)(x)\n        x = tf.keras.layers.Dropout(drop_rate, noise_shape=(None,1,1))(x)\n        x = tf.keras.layers.Add()([attn_out, x])\n        return x\n    return apply\n\ndef positional_encoding(maxlen, num_hid):\n        depth = num_hid/2\n        positions = tf.range(maxlen, dtype = tf.float32)[..., tf.newaxis]\n        depths = tf.range(depth, dtype = tf.float32)[np.newaxis, :]/depth\n        angle_rates = tf.math.divide(1, tf.math.pow(tf.cast(10000, tf.float32), depths))\n        angle_rads = tf.linalg.matmul(positions, angle_rates)\n        pos_encoding = tf.concat(\n          [tf.math.sin(angle_rads), tf.math.cos(angle_rads)],\n          axis=-1)\n        return pos_encoding","metadata":{"id":"IbHPsNTmh6lB","executionInfo":{"status":"ok","timestamp":1692564110361,"user_tz":360,"elapsed":4,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"execution":{"iopub.status.busy":"2023-08-24T21:11:38.342392Z","iopub.execute_input":"2023-08-24T21:11:38.342968Z","iopub.status.idle":"2023-08-24T21:11:38.621875Z","shell.execute_reply.started":"2023-08-24T21:11:38.342935Z","shell.execute_reply":"2023-08-24T21:11:38.620793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def CTCLoss(labels, logits):\n    label_length = tf.reduce_sum(tf.cast(labels != pad_token_idx, tf.int32), axis=-1)\n    logit_length = tf.ones(tf.shape(logits)[0], dtype=tf.int32) * tf.shape(logits)[1]\n\n    loss = classic_ctc_loss(\n            labels=labels,\n            logits=logits,\n            label_length=label_length,\n            logit_length=logit_length,\n            blank_index=pad_token_idx,\n        )\n    '''\n    loss = tf.nn.ctc_loss(\n            labels=labels,\n            logits=logits,\n            label_length=label_length,\n            logit_length=logit_length,\n            blank_index=pad_token_idx,\n            logits_time_major=False\n        )\n    '''\n    loss = tf.reduce_mean(loss)\n    return loss","metadata":{"id":"wYrHcCeQh6lB","executionInfo":{"status":"ok","timestamp":1692564110361,"user_tz":360,"elapsed":4,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"execution":{"iopub.status.busy":"2023-08-24T21:11:38.623805Z","iopub.execute_input":"2023-08-24T21:11:38.624292Z","iopub.status.idle":"2023-08-24T21:11:38.635791Z","shell.execute_reply.started":"2023-08-24T21:11:38.624258Z","shell.execute_reply":"2023-08-24T21:11:38.634781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# k_size = list of kernel sizes, if len(k_size) == 1, for all Conv1DBlocks\ndef get_model(input_shape, dim = 384, num_blocks=3, drop_rate=0.4, dropout_step=0, use_pe=False, k_sizes=None, extra_t=0):\n    inp = tf.keras.Input(input_shape)\n    x = tf.keras.layers.Masking(mask_value=0.0)(inp)\n    x = tf.keras.layers.Dense(dim, use_bias=False,name='stem_conv')(x)\n    # positional encoding\n    if use_pe:\n        pe = tf.cast(positional_encoding(input_shape[0], dim), dtype=x.dtype)\n        x = x + pe\n    x = tf.keras.layers.BatchNormalization(momentum=0.95,name='stem_bn')(x)\n\n    for i in range(num_blocks):\n        # Convolutional layers\n        if len(k_sizes) > 1:\n            for j in range(len(k_sizes)):\n                x = Conv1DBlock(dim, kernel_size=k_sizes[j], drop_rate=drop_rate)(x)\n        else:  #len(k_sizes) == 1\n            x = Conv1DBlock(dim, kernel_size=k_sizes[0], drop_rate=drop_rate)(x)\n            x = Conv1DBlock(dim, kernel_size=k_sizes[0], drop_rate=drop_rate)(x)\n            x = Conv1DBlock(dim, kernel_size=k_sizes[0], drop_rate=drop_rate)(x)\n        # Transformer\n        x = TransformerBlock(dim, expand=2)(x)\n\n#     for i in range(extra_t):\n#         x = TransformerBlock(dim, expand=2)(x)\n\n\n    x = tf.keras.layers.Dense(dim*2,activation='relu',name='top_conv')(x)\n    x = tf.keras.layers.Dropout(0.4)(x)\n    # x = LateDropout(0.6, start_step=dropout_step)(x)\n    x = tf.keras.layers.Dense(len(char_to_num), dtype='float32')(x)\n\n    return tf.keras.Model(inp, x)\n\n\n# tf.keras.backend.clear_session()\n\n# model = get_model()\n# model(batch[0])\n# model.summary()","metadata":{"id":"6_2Oy-GJh6lC","executionInfo":{"status":"ok","timestamp":1692564110361,"user_tz":360,"elapsed":3,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"execution":{"iopub.status.busy":"2023-08-24T21:11:38.641125Z","iopub.execute_input":"2023-08-24T21:11:38.641467Z","iopub.status.idle":"2023-08-24T21:11:38.654982Z","shell.execute_reply.started":"2023-08-24T21:11:38.641431Z","shell.execute_reply":"2023-08-24T21:11:38.654052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Callbacks","metadata":{"id":"uPo18tZRh6lC"}},{"cell_type":"code","source":"def num_to_char_fn(y):\n    return [num_to_char.get(x, \"\") for x in y]\n\n@tf.function()\ndef decode_phrase(pred):\n    x = tf.argmax(pred, axis=1)\n    diff = tf.not_equal(x[:-1], x[1:])\n    adjacent_indices = tf.where(diff)[:, 0]\n    x = tf.gather(x, adjacent_indices)\n    mask = x != pad_token_idx\n    x = tf.boolean_mask(x, mask, axis=0)\n    return x\n\n# A utility function to decode the output of the network\ndef decode_batch_predictions(pred):\n    output_text = []\n    for result in pred:\n        result = \"\".join(num_to_char_fn(decode_phrase(result).numpy()))\n        output_text.append(result)\n    return output_text\n\ndef print_divider():\n    print()\n    print()\n    print('----------------------------------------------------------------------------------------------')","metadata":{"id":"Q7_kRMXTh6lC","executionInfo":{"status":"ok","timestamp":1692564110362,"user_tz":360,"elapsed":4,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"execution":{"iopub.status.busy":"2023-08-24T21:11:38.656531Z","iopub.execute_input":"2023-08-24T21:11:38.657325Z","iopub.status.idle":"2023-08-24T21:11:38.667115Z","shell.execute_reply.started":"2023-08-24T21:11:38.657292Z","shell.execute_reply":"2023-08-24T21:11:38.666167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# A callback class to output a few transcriptions during training\nclass CallbackEval(tf.keras.callbacks.Callback):\n    \"\"\"Displays a batch of outputs after every epoch.\"\"\"\n\n    def __init__(self, dataset):\n        super().__init__()\n        self.dataset = dataset\n\n    def on_epoch_end(self, epoch: int, logs=None):\n        model.save_weights(\"model.h5\")\n        predictions = []\n        targets = []\n        for batch in self.dataset:\n            X, y = batch\n            batch_predictions = model(X)\n            batch_predictions = decode_batch_predictions(batch_predictions)\n            predictions.extend(batch_predictions)\n            for label in y:\n                label = \"\".join(num_to_char_fn(label.numpy()))\n                targets.append(label)\n        print(\"-\" * 100)\n        # for i in np.random.randint(0, len(predictions), 2):\n        for i in range(32):\n            print(f\"Target    : {targets[i]}\")\n            print(f\"Prediction: {predictions[i]}, len: {len(predictions[i])}\")\n            print(\"-\" * 100)\n\n# Callback function to check transcription on the val set.\n# validation_callback = CallbackEval(val_dataset.take(1))","metadata":{"id":"u07d05Twh6lC","executionInfo":{"status":"ok","timestamp":1692564110362,"user_tz":360,"elapsed":4,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"execution":{"iopub.status.busy":"2023-08-24T21:11:38.668612Z","iopub.execute_input":"2023-08-24T21:11:38.668963Z","iopub.status.idle":"2023-08-24T21:11:38.680931Z","shell.execute_reply.started":"2023-08-24T21:11:38.668931Z","shell.execute_reply":"2023-08-24T21:11:38.679786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nwith open (ROOT+\"/kaggle/input/asl-fingerspelling/character_to_prediction_index.json\", \"r\") as f:\n    character_map = json.load(f)\nrev_character_map = {j:i for i,j in character_map.items()}\n\n\nclass val_lev_callback(tf.keras.callbacks.Callback):\n    def __init__(self, model):\n        super().__init__()\n        # self.val_set=val_set,\n        self.model=model,\n    def on_epoch_end(self, epoch: int, logs=None):\n        calculate_val_lev(self.model, print_lev=True)\n\n\ndef calculate_val_lev(model, val_split=None, print_lev=False, ret=False):\n    if val_split == 'AB':\n        val_set = AB_val_set\n    elif val_split == 'CD':\n        val_set = CD_val_set\n    elif val_split == 'EF':\n        val_set = EF_val_set\n    else: # None\n        val_set = full_val_set\n\n    preds = []\n    targets = []\n    scores = []\n    for batch_idx in range(len(val_set)):\n        preds_batch = model.predict(val_set[batch_idx][0], verbose = 0)\n        targets_batch = val_set[batch_idx][1]\n        for pred_idx in range(len(preds_batch)):\n            preds.append(\"\".join([rev_character_map.get(s, \"\") for s in decode_phrase(preds_batch[pred_idx]).numpy()]))\n            targets.append(\"\".join([rev_character_map.get(s, \"\") for s in targets_batch[pred_idx].numpy()]))\n\n    N = [len(phrase) for phrase in targets]\n    lev_dist = [lev.distance(preds[i], targets[i]) for i in range(len(targets))]\n    lev_total = (np.sum(N) - np.sum(lev_dist)) / np.sum(N)\n\n    if print_lev:\n        print('Lev distance: '+str(lev_total))\n    if ret:\n        return lev_total","metadata":{"id":"0uu0WWPQh6lC","executionInfo":{"status":"ok","timestamp":1692564110362,"user_tz":360,"elapsed":4,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"execution":{"iopub.status.busy":"2023-08-24T21:11:38.682813Z","iopub.execute_input":"2023-08-24T21:11:38.683169Z","iopub.status.idle":"2023-08-24T21:11:38.698442Z","shell.execute_reply.started":"2023-08-24T21:11:38.683138Z","shell.execute_reply":"2023-08-24T21:11:38.697477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def lrfn(current_step, num_warmup_steps, lr_max, num_cycles=0.50, num_training_steps=50):\n    WARMUP_METHOD = \"lin\"\n    if current_step < num_warmup_steps:\n        if WARMUP_METHOD == 'log':\n            return lr_max * 0.10 ** (num_warmup_steps - current_step)\n        elif WARMUP_METHOD == 'exp':\n            return lr_max * 2 ** -(num_warmup_steps - current_step)\n        else:  # linear\n            return lr_max * ((current_step + 0.01) / num_warmup_steps)\n    else:\n        progress = float(current_step - num_warmup_steps) / float(max(1, num_training_steps - num_warmup_steps))\n\n        return max(0.0, 0.5 * (1.0 + math.cos(math.pi * float(num_cycles) * 2.0 * progress))) * lr_max\n\ndef plot_lr_schedule(lr_schedule, epochs):\n    fig = plt.figure(figsize=(20, 10))\n    plt.plot([None] + lr_schedule + [None])\n    # X Labels\n    x = np.arange(1, epochs + 1)\n    x_axis_labels = [i if epochs <= 40 or i % 5 == 0 or i == 1 else None for i in range(1, epochs + 1)]\n    plt.xlim([1, epochs])\n    plt.xticks(x, x_axis_labels) # set tick step to 1 and let x axis start at 1\n\n    # Increase y-limit for better readability\n    plt.ylim([0, max(lr_schedule) * 1.1])\n\n    # Title\n    schedule_info = f'start: {lr_schedule[0]:.1E}, max: {max(lr_schedule):.1E}, final: {lr_schedule[-1]:.1E}'\n    plt.title(f'Step Learning Rate Schedule, {schedule_info}', size=18, pad=12)\n\n    # Plot Learning Rates\n    for x, val in enumerate(lr_schedule):\n        if epochs <= 40 or x % 5 == 0 or x is epochs - 1:\n            if x < len(lr_schedule) - 1:\n                if lr_schedule[x - 1] < val:\n                    ha = 'right'\n                else:\n                    ha = 'left'\n            elif x == 0:\n                ha = 'right'\n            else:\n                ha = 'left'\n            plt.plot(x + 1, val, 'o', color='black');\n            offset_y = (max(lr_schedule) - min(lr_schedule)) * 0.02\n            plt.annotate(f'{val:.1E}', xy=(x + 1, val + offset_y), size=12, ha=ha)\n\n    plt.xlabel('Epoch', size=16, labelpad=5)\n    plt.ylabel('Learning Rate', size=16, labelpad=5)\n    plt.grid()\n    plt.show()\n\n# Custom callback to update weight decay with learning rate\nclass WeightDecayCallback(tf.keras.callbacks.Callback):\n    def __init__(self, model, wd_ratio=0.05):\n        self.model = model\n        self.step_counter = 0\n        self.wd_ratio = wd_ratio\n\n    def on_epoch_begin(self, epoch, logs=None):\n        self.model.optimizer.weight_decay = self.model.optimizer.learning_rate * self.wd_ratio\n        print(f'learning rate: {self.model.optimizer.learning_rate.numpy():.2e}, weight decay: {self.model.optimizer.weight_decay.numpy():.2e}')\n\nclass save_model_callback(tf.keras.callbacks.Callback):\n    def __init__(self, label, freq):\n        super().__init__()\n        self.label = label\n        self.freq = freq\n    def on_epoch_end(self, epoch: int, logs=None):\n        if (epoch+1)%self.freq == 0:\n            self.model.save_weights(f\"{self.label}_e{epoch}.h5\")","metadata":{"id":"GG7qWZwIh6lC","executionInfo":{"status":"ok","timestamp":1692564110520,"user_tz":360,"elapsed":162,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"execution":{"iopub.status.busy":"2023-08-24T21:11:38.700255Z","iopub.execute_input":"2023-08-24T21:11:38.700703Z","iopub.status.idle":"2023-08-24T21:11:38.720879Z","shell.execute_reply.started":"2023-08-24T21:11:38.700672Z","shell.execute_reply":"2023-08-24T21:11:38.719751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# save training log in text file\ndef save_history(history, save_path, comment):\n    # Convert the history logs to a text file\n    with open(f\"{save_path}/{comment}.txt\", \"w\") as file:\n        # Write the headers (column names)\n        file.write(\"Epoch\\t\" + \"\\t\".join([key for key in history.history.keys()]) + \"\\n\")\n\n        # Write the values for each epoch\n        for i in range(len(history.history['loss'])):  # Assuming 'loss' is always present\n            file.write(f\"{i+1}\\t\" + \"\\t\".join([str(history.history[key][i]) for key in history.history.keys()]) + \"\\n\")","metadata":{"id":"kD7ZtacsJgRT","executionInfo":{"status":"ok","timestamp":1692564110521,"user_tz":360,"elapsed":2,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"execution":{"iopub.status.busy":"2023-08-24T21:11:38.721867Z","iopub.execute_input":"2023-08-24T21:11:38.722111Z","iopub.status.idle":"2023-08-24T21:11:38.735491Z","shell.execute_reply.started":"2023-08-24T21:11:38.72209Z","shell.execute_reply":"2023-08-24T21:11:38.734428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class cfg:\n    def __init__(self, name):\n        # data\n        self.max_len = 256  # 128\n        self.augment = True\n        self.preload_ds = True\n\n        # model\n        self.dim = 384  # 192\n        self.num_blocks = 6 # 6\n        self.dropout = 0.3\n        self.pe = False\n        self.k_sizes = [11,5,3]  # list of kernel sizes\n        self.extra_t = 0\n\n        # hyperparameters\n        self.lr = 2.5e-2  # 5e-4 * replicas\n        self.weight_decay = 0.1  # 0.1\n        self.lr_min = 1e-6\n        self.epoch = 100  # 300\n        self.warmup = 10  # 0\n        self.batch_size = 1024  # 64 * replicas\n\n        self.fp16 = True\n        #     fgm = False\n        #     awp = True\n        #     awp_lambda = 0.2\n        #     awp_start_epoch = 15\n        self.dropout_start_epoch = 15\n        #     resume = 0\n        #     decay_type = 'cosine'\n\n        self.save_output=True\n        self.save_freq = self.epoch  # save last epoch only\n        self.name = name\n\n    def set_param(self, save_output=None, save_freq=None, name=None,\n                  dim=None, num_blocks=None, dropout=None, pe=None, k_sizes=None, extra_t=None,\n                  lr=None, weight_decay=None, epoch=None, warmup=None, batch_size=None,\n                  ):\n        if save_output:\n            self.save_output = save_output\n        if save_freq:\n            self.save_freq = save_freq\n        if name:\n            self.name = name\n        if dim:\n            self.dim = dim\n        if num_blocks:\n            self.num_blocks = num_blocks\n        if dropout:\n            self.dropout = dropout\n        if pe:\n            self.pe = pe\n        if k_sizes:\n            self.k_sizes = k_sizes\n        if extra_t:\n            self.extra_t = extra_t\n        if lr:\n            self.lr = lr\n        if weight_decay:\n            self.weight_decay = weight_decay\n        if epoch:\n            self.epoch = epoch\n        if warmup:\n            self.warmup = warmup\n        if batch_size:\n            self.batch_size = batch_size\n\n    def get_comment(self):\n        comment = f\"{self.name}_rs{self.max_len}_b{self.num_blocks}_d{self.dropout}_wd{self.weight_decay}_bs{self.batch_size}_lr{self.lr}\"\n        return comment\n\n    def get_name(self):\n        return self.name\n\n\n# Create cfg object\nCFG = cfg(name = 'NewDS')\nif DEBUG:\n    CFG.set_param(epoch=2)\n    CFG.set_param(batch_size=64)","metadata":{"id":"ihlZ7xnEh6lC","executionInfo":{"status":"ok","timestamp":1692564110521,"user_tz":360,"elapsed":2,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"execution":{"iopub.status.busy":"2023-08-24T21:11:38.738945Z","iopub.execute_input":"2023-08-24T21:11:38.739199Z","iopub.status.idle":"2023-08-24T21:11:38.753148Z","shell.execute_reply.started":"2023-08-24T21:11:38.739177Z","shell.execute_reply":"2023-08-24T21:11:38.752194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def print_gsfiles(file_list):\n    for i in range(len(file_list)):\n        print(f\"{i}: {file_list[i][-33:]}\")\n\n# Key: grade ('A', 'B', 'C',..), Value: list of files\ntrain_splits = dict()\ndata_grades = ['A', 'B', 'C', 'D', 'E', 'F']\nfor grade in data_grades:\n    pattern = f\"{data_dir}/{grade}_train/*.tfrecord\"\n    files = tf.io.gfile.glob(pattern)\n    files.sort()\n    train_splits[grade] = files\n\n# list of indices of validition files in each split\n# val_indices = {'A':{66}, 'B':{}, 'C':{2,43}, 'D':{7}, 'E':{}, 'F':{}}  # first val_set (640)\nval_indices = {'A':{22,11}, 'B':{7,34}, 'C':{13,17}, 'D':{32,33}, 'E':{24,25}, 'F':{56,60}}\n\n# subset val_files\nAB_val_files = []\nCD_val_files = []\nEF_val_files = []\n\n# Create train val splits\ntrain_files = []\nval_files = []\nfor grade in data_grades:\n    split_files = train_splits[grade]\n    val_split = []\n    for i in val_indices[grade]:\n        val_split.append(split_files[i])\n    val_files.extend(val_split)\n\n    if grade == 'A' or grade == 'B':\n        AB_val_files.extend(val_split)\n    elif grade == 'C' or grade == 'D':\n        CD_val_files.extend(val_split)\n    elif grade == 'E' or grade == 'F':\n        EF_val_files.extend(val_split)\n\n#     if grade == 'F':\n#         train_split = [file for file in split_files if file not in val_split]\n#         train_files.extend(train_split)\n\n\nif DEBUG==True:\n    # tffiles = tffiles[:10]\n    train_files = train_files[:50]\n    val_files = val_files[:50]\n\nprint(f\"train files:{len(train_files)}, val_files: {len(val_files)}\")\nprint_gsfiles(val_files)\nprint_gsfiles(train_files[-20:])","metadata":{"id":"8vLfrNJEaqbj","executionInfo":{"status":"ok","timestamp":1692564110811,"user_tz":360,"elapsed":292,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"outputId":"6bbc7934-01d3-4a3c-f48e-65374b1e3317","execution":{"iopub.status.busy":"2023-08-24T21:11:38.754554Z","iopub.execute_input":"2023-08-24T21:11:38.755082Z","iopub.status.idle":"2023-08-24T21:11:38.94823Z","shell.execute_reply.started":"2023-08-24T21:11:38.755051Z","shell.execute_reply":"2023-08-24T21:11:38.947209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# def print_gsfiles(file_list):\n#     for i in range(len(file_list)):\n#         print(f\"{i}: {file_list[i][-33:]}\")\n\n# # Key: grade ('A', 'B', 'C',..), Value: list of files\n# train_splits = dict()\n# data_grades = ['A', 'B', 'C', 'D', 'E'] #removed 'F'\n# for grade in data_grades:\n#     pattern = f\"{data_dir}/{grade}_train/*.tfrecord\"\n#     files = tf.io.gfile.glob(pattern)\n#     files.sort()\n#     train_splits[grade] = files\n\n# # list of indices of validition files in each split\n# # val_indices = {'A':{66}, 'B':{}, 'C':{2,43}, 'D':{7}, 'E':{}, 'F':{}}  # first val_set (640)\n# val_indices = {'A':{22,11}, 'B':{7,34}, 'C':{13,17}, 'D':{32,33}, 'E':{24,25}}  #'F'{56,60}\n\n\n# # Create train val splits\n# train_files = []\n# val_files = []\n# for grade in data_grades:\n#     split_files = train_splits[grade]\n#     val_split = []\n#     for i in val_indices[grade]:\n#         val_split.append(split_files[i])\n#     val_files.extend(val_split)\n\n#     if grade != 'F':\n#         train_split = [file for file in split_files if file not in val_split]\n#         train_files.extend(train_split)\n#     # Double A and B\n#     # if grade == 'A' or grade == 'B':\n#     #     train_files.extend(train_split)\n\n\n# # Better F data\n# pattern = f\"/kaggle/input/my-dataset/F2/tfds/F2/*.tfrecord\"\n# files = tf.io.gfile.glob(pattern)\n# for file in files:\n#     if file == f\"/kaggle/input/my-dataset/F2/tfds/F2/127_n0.12_g0.94.tfrecord\" or file == f\"/kaggle/input/my-dataset/F2/tfds/F2/129_n0.12_g0.95.tfrecord\":\n#         val_files.append(file)\n#     else:\n#         train_files.append(file)\n\n# if DEBUG==True:\n#     # tffiles = tffiles[:10]\n#     train_files = train_files[:50]\n#     val_files = val_files[:50]\n\n# print(f\"train files:{len(train_files)}, val_files: {len(val_files)}\")\n# print_gsfiles(val_files)\n# print()\n# print_gsfiles(train_files[:10])\n# print_gsfiles(train_files[-10:])","metadata":{"execution":{"iopub.status.busy":"2023-08-24T21:11:38.949633Z","iopub.execute_input":"2023-08-24T21:11:38.949967Z","iopub.status.idle":"2023-08-24T21:11:38.956155Z","shell.execute_reply.started":"2023-08-24T21:11:38.949936Z","shell.execute_reply":"2023-08-24T21:11:38.955272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# EF_val_files = []\n# EF_val_files.append('/kaggle/input/my-dataset/tfrecords/tfrecords/train/tfds/E_train/80_n0.26_g2.16.tfrecord')\n# EF_val_files.append('/kaggle/input/my-dataset/tfrecords/tfrecords/train/tfds/E_train/80_n0.28_g2.18.tfrecord')\n# EF_val_files.append('/kaggle/input/my-dataset/F2/tfds/F2/129_n0.12_g0.95.tfrecord')\n# EF_val_files.append('/kaggle/input/my-dataset/F2/tfds/F2/127_n0.12_g0.94.tfrecord')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T21:11:38.957731Z","iopub.execute_input":"2023-08-24T21:11:38.958328Z","iopub.status.idle":"2023-08-24T21:11:38.969548Z","shell.execute_reply.started":"2023-08-24T21:11:38.958296Z","shell.execute_reply":"2023-08-24T21:11:38.968568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create datasets\nglobal FRAME_LEN\nFRAME_LEN = CFG.max_len\ntrain_batch_size = CFG.batch_size\nval_batch_size = 992  #640? 1920\n\n# no augment\n# train_dataset_pre =  tf.data.TFRecordDataset(tffiles[val_len:]).prefetch(tf.data.AUTOTUNE).map(decode_fn, num_parallel_calls=tf.data.AUTOTUNE).map(pre_process_fn, num_parallel_calls=tf.data.AUTOTUNE)\n# augment\n# train_dataset_pre =  tf.data.TFRecordDataset(train_files).prefetch(tf.data.AUTOTUNE).map(decode_fn, num_parallel_calls=tf.data.AUTOTUNE).map(\n#     lambda lip, rhand, lhand, rpose, lpose, phrase: pre_process_fn(lip, rhand, lhand, rpose, lpose, phrase, CFG.augment), num_parallel_calls=tf.data.AUTOTUNE)\nval_dataset_pre =  tf.data.TFRecordDataset(val_files).prefetch(tf.data.AUTOTUNE).map(\n    decode_fn, num_parallel_calls=tf.data.AUTOTUNE).map(pre_process_fn, num_parallel_calls=tf.data.AUTOTUNE)\n\n\nval_items = [x for x in val_dataset_pre]\nval_items_X = [x[0] for x in val_items]\nval_items_y = [tf.cast(x[1], dtype = tf.int32) for x in val_items]\nval_dataset = tf.data.Dataset.from_tensor_slices((val_items_X, val_items_y)).prefetch(tf.data.AUTOTUNE).batch(\n    val_batch_size, drop_remainder=True).prefetch(tf.data.AUTOTUNE)\n\n# train_items = [x for x in train_dataset_pre]\n# train_items_X = [x[0] for x in train_items]\n# train_items_y = [tf.cast(x[1], dtype = tf.int32) for x in train_items]\n# train_dataset = tf.data.Dataset.from_tensor_slices((train_items_X, train_items_y))\n# train_dataset = train_dataset.prefetch(tf.data.AUTOTUNE).repeat().shuffle(\n#     buffer_size=60000, reshuffle_each_iteration = True).batch(train_batch_size, drop_remainder=True).prefetch(\n#     tf.data.AUTOTUNE)\n\n\nINPUT_SHAPE = [CFG.max_len, CHANNELS]\nfull_val_set = [x for x in val_dataset]\n","metadata":{"id":"i_GBUnTrk-m_","executionInfo":{"status":"ok","timestamp":1692564124152,"user_tz":360,"elapsed":13342,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"outputId":"2ac72b7f-49c2-4ada-fb78-8e88b37be9eb","execution":{"iopub.status.busy":"2023-08-24T21:11:38.971029Z","iopub.execute_input":"2023-08-24T21:11:38.971539Z","iopub.status.idle":"2023-08-24T21:11:45.718883Z","shell.execute_reply.started":"2023-08-24T21:11:38.971506Z","shell.execute_reply":"2023-08-24T21:11:45.717778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# specialized validation datasets\nAB_val_dataset_pre =  tf.data.TFRecordDataset(AB_val_files).prefetch(tf.data.AUTOTUNE).map(\n    decode_fn, num_parallel_calls=tf.data.AUTOTUNE).map(pre_process_fn, num_parallel_calls=tf.data.AUTOTUNE)\nAB_val_items = [x for x in AB_val_dataset_pre]\nAB_val_items_X = [x[0] for x in AB_val_items]\nAB_val_items_y = [tf.cast(x[1], dtype = tf.int32) for x in AB_val_items]\nAB_val_dataset = tf.data.Dataset.from_tensor_slices((AB_val_items_X, AB_val_items_y)).prefetch(tf.data.AUTOTUNE).batch(\n    64, drop_remainder=True).prefetch(tf.data.AUTOTUNE)\nAB_val_set = [x for x in AB_val_dataset]\n\nCD_val_dataset_pre =  tf.data.TFRecordDataset(CD_val_files).prefetch(tf.data.AUTOTUNE).map(\n    decode_fn, num_parallel_calls=tf.data.AUTOTUNE).map(pre_process_fn, num_parallel_calls=tf.data.AUTOTUNE)\nCD_val_items = [x for x in CD_val_dataset_pre]\nCD_val_items_X = [x[0] for x in CD_val_items]\nCD_val_items_y = [tf.cast(x[1], dtype = tf.int32) for x in CD_val_items]\nCD_val_dataset = tf.data.Dataset.from_tensor_slices((CD_val_items_X, CD_val_items_y)).prefetch(tf.data.AUTOTUNE).batch(\n    64, drop_remainder=True).prefetch(tf.data.AUTOTUNE)\nCD_val_set = [x for x in CD_val_dataset]\n\nEF_val_dataset_pre =  tf.data.TFRecordDataset(EF_val_files).prefetch(tf.data.AUTOTUNE).map(\n    decode_fn, num_parallel_calls=tf.data.AUTOTUNE).map(pre_process_fn, num_parallel_calls=tf.data.AUTOTUNE)\nEF_val_items = [x for x in EF_val_dataset_pre]\nEF_val_items_X = [x[0] for x in EF_val_items]\nEF_val_items_y = [tf.cast(x[1], dtype = tf.int32) for x in EF_val_items]\nEF_val_dataset = tf.data.Dataset.from_tensor_slices((EF_val_items_X, EF_val_items_y)).prefetch(tf.data.AUTOTUNE).batch(\n    64, drop_remainder=True).prefetch(tf.data.AUTOTUNE)\nEF_val_set = [x for x in EF_val_dataset]\n","metadata":{"id":"oUBbQjayClHx","executionInfo":{"status":"ok","timestamp":1692564127520,"user_tz":360,"elapsed":3376,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"outputId":"e98af09f-3277-4a39-9438-f4f4a8826c62","execution":{"iopub.status.busy":"2023-08-24T21:11:45.724211Z","iopub.execute_input":"2023-08-24T21:11:45.726564Z","iopub.status.idle":"2023-08-24T21:11:51.794384Z","shell.execute_reply.started":"2023-08-24T21:11:45.726526Z","shell.execute_reply":"2023-08-24T21:11:51.793185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"CFG.set_param(k_sizes=[11,5,3])\nCFG.set_param(num_blocks=6)\nCFG.set_param(extra_t=0)\n# Create model\nwith strategy.scope():\n    model = get_model(input_shape=INPUT_SHAPE, num_blocks=CFG.num_blocks, dim=CFG.dim, drop_rate=CFG.dropout,\n                      use_pe=CFG.pe, k_sizes=CFG.k_sizes, extra_t=CFG.extra_t)\n\n    # Adam Optimizer\n    optimizer = tfa.optimizers.RectifiedAdam(sma_threshold=4)\n    optimizer = tfa.optimizers.Lookahead(optimizer, sync_period=5)\n\n    model.compile(\n        loss=CTCLoss,\n        optimizer=optimizer,\n#         steps_per_execution=steps_per_epoch,\n    )\n\n# model.summary()","metadata":{"id":"oOS5VZJRb9xo","executionInfo":{"status":"ok","timestamp":1692567278409,"user_tz":360,"elapsed":16683,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"outputId":"255287a0-dcb5-46e6-c0ba-0c31f4e4612a","execution":{"iopub.status.busy":"2023-08-24T21:11:51.795818Z","iopub.execute_input":"2023-08-24T21:11:51.796849Z","iopub.status.idle":"2023-08-24T21:11:55.118582Z","shell.execute_reply.started":"2023-08-24T21:11:51.79682Z","shell.execute_reply":"2023-08-24T21:11:55.117589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.load_weights('/kaggle/input/trained4/0.762-Ftrain_supp_k1153_rs256_b6_d0.3_wd0.1_bs160_lr0.005_e79.h5')\nprint(\"Overall:\")\ncalculate_val_lev(model, print_lev=True)\nprint(\"AB val:\")\ncalculate_val_lev(model, val_split='AB', print_lev=True)\nprint(\"CD val:\")\ncalculate_val_lev(model, val_split='CD', print_lev=True)\nprint(\"EF val:\")\ncalculate_val_lev(model, val_split='EF', print_lev=True)","metadata":{"id":"acrSSAu2b-6i","executionInfo":{"status":"ok","timestamp":1692567348407,"user_tz":360,"elapsed":69999,"user":{"displayName":"Jayoo Hwang","userId":"15747763069873097775"}},"outputId":"15caf0cb-dba8-4b93-d7f9-aabd9c37b568","execution":{"iopub.status.busy":"2023-08-24T21:11:55.122708Z","iopub.execute_input":"2023-08-24T21:11:55.123013Z","iopub.status.idle":"2023-08-24T21:12:31.033947Z","shell.execute_reply.started":"2023-08-24T21:11:55.122987Z","shell.execute_reply":"2023-08-24T21:12:31.032929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class TFLiteModel(tf.Module):\n    def __init__(self, model):\n        super(TFLiteModel, self).__init__()\n        self.model = model\n    \n    @tf.function(input_signature=[tf.TensorSpec(shape=[None, len(SEL_COLS)], dtype=tf.float32, name='inputs')])\n    def __call__(self, inputs, training=False):\n        # Preprocess Data\n        x = tf.cast(inputs, tf.float32)\n        x = x[None]\n        x = tf.cond(tf.shape(x)[1] == 0, lambda: tf.zeros((1, 1, len(SEL_COLS))), lambda: tf.identity(x))\n        x = x[0]\n        x = pre_process0(x)\n        x = pre_process1(*x)\n        x = tf.reshape(x, INPUT_SHAPE)\n        x = x[None]\n        x = self.model(x, training=False)\n        x = x[0]\n        x = decode_phrase(x)\n        x = tf.cond(tf.shape(x)[0] == 0, lambda: tf.zeros(1, tf.int32), lambda: tf.cast(tf.identity(x), tf.int32))\n        x = tf.one_hot(x, 59)\n        return {'outputs': x}\n\ntflitemodel_base = TFLiteModel(model)\ntflitemodel_base(frames)[\"outputs\"].shape","metadata":{"execution":{"iopub.status.busy":"2023-08-24T21:16:23.450089Z","iopub.execute_input":"2023-08-24T21:16:23.450656Z","iopub.status.idle":"2023-08-24T21:16:28.853753Z","shell.execute_reply.started":"2023-08-24T21:16:23.450615Z","shell.execute_reply":"2023-08-24T21:16:28.852765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"keras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflitemodel_base)\nkeras_model_converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS]#, tf.lite.OpsSet.SELECT_TF_OPS]\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT]\nkeras_model_converter.target_spec.supported_types = [tf.float16]\ntflite_model = keras_model_converter.convert()\nwith open('model.tflite', 'wb') as f:\n    f.write(tflite_model)\n    \nwith open('inference_args.json', \"w\") as f:\n    json.dump({\"selected_columns\" : SEL_COLS}, f)\n    \n!zip submission.zip  './model.tflite' './inference_args.json'","metadata":{"execution":{"iopub.status.busy":"2023-08-24T21:16:28.858641Z","iopub.execute_input":"2023-08-24T21:16:28.860958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with open (\"inference_args.json\", \"r\") as f:\n    SEL_COLS = json.load(f)[\"selected_columns\"]\n    \ndef load_relevant_data_subset(pq_path):\n    return pd.read_parquet(pq_path, columns=SEL_COLS)\n\ndef create_data_gen(file_ids, y_mul=1):\n    def gen():\n        for file_id in file_ids:\n            pqfile = f\"{inpdir}/{file_id}.parquet\"\n            seq_refs = df.loc[df.file_id == file_id]\n            seqs = load_relevant_data_subset(pqfile)\n\n            for seq_id in seq_refs.sequence_id:\n                x = seqs.iloc[seqs.index == seq_id].to_numpy()\n                y = str(df.loc[df.sequence_id == seq_id].phrase.iloc[0])\n                \n                r_nonan = np.sum(np.sum(np.isnan(x[:, RHAND_IDX_X]), axis = 1) == 0)\n                l_nonan = np.sum(np.sum(np.isnan(x[:, LHAND_IDX_X]), axis = 1) == 0)\n                no_nan = max(r_nonan, l_nonan)\n                \n                if y_mul*len(y)<no_nan:\n                    yield x, y\n    return gen\n\npqfiles = df.file_id.unique()\nval_len = int(0.05 * len(pqfiles))\n\ntest_dataset = tf.data.Dataset.from_generator(create_data_gen(pqfiles[:val_len], 0),\n    output_signature=(tf.TensorSpec(shape=(None, len(SEL_COLS)), dtype=tf.float32), tf.TensorSpec(shape=(), dtype=tf.string))\n).prefetch(buffer_size=2000)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"interpreter = tf.lite.Interpreter(\"model.tflite\")\n\nREQUIRED_SIGNATURE = \"serving_default\"\nREQUIRED_OUTPUT = \"outputs\"\n\nwith open (\"/kaggle/input/asl-fingerspelling/character_to_prediction_index.json\", \"r\") as f:\n    character_map = json.load(f)\nrev_character_map = {j:i for i,j in character_map.items()}\n\nprediction_fn = interpreter.get_signature_runner(REQUIRED_SIGNATURE)\n\nfor frame, target in test_dataset.skip(100).take(10):\n    output = prediction_fn(inputs=frame)\n    prediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output[REQUIRED_OUTPUT], axis=1)])\n    target = target.numpy().decode(\"utf-8\")\n    print(\"pred =\", prediction_str, \"; target =\", target)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%timeit -n 10\noutput = prediction_fn(inputs=frame)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from Levenshtein import distance\n\nscores = []\n\nfor i, (frame, target) in tqdm(enumerate(test_dataset.take(1000))):\n    output = prediction_fn(inputs=frame)\n    prediction_str = \"\".join([rev_character_map.get(s, \"\") for s in np.argmax(output[REQUIRED_OUTPUT], axis=1)])\n    target = target.numpy().decode(\"utf-8\")\n    score = (len(target) - distance(prediction_str, target)) / len(target)\n    scores.append(score)\n    if i % 50 == 0:\n        print(np.sum(scores) / len(scores))\n    \nscores = np.array(scores)\nprint(np.sum(scores) / len(scores))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}