{"metadata":{"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":52950,"databundleVersionId":5973250,"sourceType":"competition"},{"sourceId":5835808,"sourceType":"datasetVersion","datasetId":3354626},{"sourceId":5939194,"sourceType":"datasetVersion","datasetId":3282038}],"dockerImageVersionId":30627,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"The model consists of a transformer embedding + encoder + decoder.\n\nInference is performed by starting with an SOS token and predicting one character at a time using the previous prediction.\n\nInference requires the encoder to encode the input frames and subsequently use that encoding to predict the 1st character by inputting the encoding and SOS (Start of Sentence) token. Next, the encoding, SOS token and 1st predicted token are used to predict the 2nd character. Inference thus requires 1 call to the encoder and multiple calls to the encoder. On average a phrase is 18 characters long, requiring 18+1(SOS token) calls to the decoder.\n\nSome inspiration is taken from the [1st place solution - training](https://www.kaggle.com/code/hoyso48/1st-place-solution-training) from the last [Google - Isolated Sign Language Recognition\n](https://www.kaggle.com/competitions/asl-signs) competition.\n\nSpecial thanks for all of these guys, Many many thanks to them: \n\nhttps://www.kaggle.com/competitions/asl-fingerspelling/discussion/434364\n\n[1st place solution] Improved Squeezeformer + TransformerDecoder + Clever augmentations: https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485\n\n[5th place solution] Vanilla Transformer, Data2vec Pretraining, CutMix, and KD: https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434415\n\nhttps://www.kaggle.com/code/gusthema/asl-fingerspelling-recognition-w-tensorflow\n\nThis man helps me alot: https://www.kaggle.com/competitions/asl-fingerspelling/discussion/411060","metadata":{}},{"cell_type":"markdown","source":"The processing is as follows:\n\n1) Select dominant hand based on most number of non empty hand frames\n\n2) Filter out all frames with missing dominant hand coordinates\n\n3) Resize video to 256 frames\n\n4) Excluding samples with low frames per character ratio\n\n5) Added phrase type","metadata":{}},{"cell_type":"markdown","source":"You Can Find the data and the competition here: https://www.kaggle.com/competitions/asl-fingerspelling/overview","metadata":{}},{"cell_type":"code","source":"#For MLOPS\n#You can find more about it here, special thanks for: https://medium.com/@darragh.hanley_94135/mastering-mlops-with-neptune-ai-84e635d36bf2\n!pip install neptune","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:02.877434Z","iopub.execute_input":"2024-01-01T08:14:02.878369Z","iopub.status.idle":"2024-01-01T08:14:25.29514Z","shell.execute_reply.started":"2024-01-01T08:14:02.878324Z","shell.execute_reply":"2024-01-01T08:14:25.293898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Uninstall the previous installed nltk library\n!pip install -U nltk\n\n# This upgraded nltkto version 3.5 in which meteor_score is there.\n!pip install nltk==3.5","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:49:30.042472Z","iopub.execute_input":"2024-01-01T12:49:30.043303Z","iopub.status.idle":"2024-01-01T12:50:02.868538Z","shell.execute_reply.started":"2024-01-01T12:49:30.043262Z","shell.execute_reply":"2024-01-01T12:50:02.867304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import sacrebleu\nfrom nltk.translate.bleu_score import sentence_bleu\nfrom rouge import Rouge\nfrom nltk.translate.meteor_score import meteor_score","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:51:38.431605Z","iopub.execute_input":"2024-01-01T12:51:38.432689Z","iopub.status.idle":"2024-01-01T12:51:38.438384Z","shell.execute_reply.started":"2024-01-01T12:51:38.432641Z","shell.execute_reply":"2024-01-01T12:51:38.437067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport matplotlib as mpl\nimport seaborn as sn\nimport tensorflow_addons as tfa\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\n\nfrom tqdm.notebook import tqdm\nfrom sklearn.model_selection import train_test_split, GroupShuffleSplit\nfrom pathlib import Path\nfrom leven import levenshtein\n\nimport glob\nimport sys\nimport os\nimport math\nimport gc\nimport sys\nimport sklearn\nimport time\nimport json\nimport re\n\nprint(f'Tensorflow Version {tf.__version__}')\nprint(f'Python Version: {sys.version}')\n# TQDM Progress Bar With Pandas Apply Function\ntqdm.pandas()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:25.297676Z","iopub.execute_input":"2024-01-01T08:14:25.298092Z","iopub.status.idle":"2024-01-01T08:14:38.105345Z","shell.execute_reply.started":"2024-01-01T08:14:25.29805Z","shell.execute_reply":"2024-01-01T08:14:38.104305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Character 2 Ordinal Encoding","metadata":{}},{"cell_type":"markdown","source":"You have first get the data from the competition, you can find it here: https://www.kaggle.com/competitions/asl-fingerspelling/data","metadata":{}},{"cell_type":"code","source":"# Read Character to Ordinal Encoding Mapping\nwith open('/kaggle/input/asl-fingerspelling/character_to_prediction_index.json') as json_file:\n    CHAR2ORD = json.load(json_file)\n    \n# Ordinal to Character Mapping\nORD2CHAR = {j:i for i,j in CHAR2ORD.items()}\n    \n# Character to Ordinal Encoding Mapping   \ndisplay(pd.Series(CHAR2ORD).to_frame('Ordinal Encoding'))","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:38.106889Z","iopub.execute_input":"2024-01-01T08:14:38.107505Z","iopub.status.idle":"2024-01-01T08:14:38.135059Z","shell.execute_reply.started":"2024-01-01T08:14:38.107473Z","shell.execute_reply":"2024-01-01T08:14:38.134177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Global Config","metadata":{}},{"cell_type":"code","source":"# If Notebook Is Run By Committing or In Interactive Mode For Development\nIS_INTERACTIVE = os.environ['KAGGLE_KERNEL_RUN_TYPE'] == 'Interactive'\n# Verbose Setting during training\nVERBOSE = 1 if IS_INTERACTIVE else 2\n# Describe Statistics Percentiles\nPERCENTILES = [0.01, 0.10, 0.05, 0.25, 0.50, 0.75, 0.90, 0.95, 0.99, 0.999]\n# Global Random Seed\nSEED = 42\n# Number of Frames to resize recording to\nN_TARGET_FRAMES = 128\n# Global debug flag, takes subset of train\nDEBUG = False\n# Fast Processing\nFAST= False\n# Number of Unique Characters To Predict + Pad Token + SOS Token + EOS Token\nN_UNIQUE_CHARACTERS0 = len(CHAR2ORD)\nN_UNIQUE_CHARACTERS = len(CHAR2ORD) + 1 + 1 + 1\nPAD_TOKEN = len(CHAR2ORD) # Padding\nSOS_TOKEN = len(CHAR2ORD) + 1 # Start Of Sentence\nEOS_TOKEN = len(CHAR2ORD) + 2 # End Of Sentence\n# Whether to use 10% of data for validation\nUSE_VAL = True\n# Batch Size\nBATCH_SIZE = 64\n# Number of Epochs to Train for\nN_EPOCHS = 100\n# Number of Warmup Epochs in Learning Rate Scheduler\nN_WARMUP_EPOCHS = 10\n# Maximum Learning Rate\nLR_MAX = 1e-3\n# Weight Decay Ratio as Ratio of Learning Rate\nWD_RATIO = 0.05\n# Length of Phrase + EOS Token\nMAX_PHRASE_LENGTH = 31 + 1\n# Whether to Train The model\nTRAIN_MODEL = True\n# Whether to Load Pretrained Weights\nLOAD_WEIGHTS = False\n# Learning Rate Warmup Method [log,exp]\nWARMUP_METHOD = 'exp'","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:38.136468Z","iopub.execute_input":"2024-01-01T08:14:38.137499Z","iopub.status.idle":"2024-01-01T08:14:38.149366Z","shell.execute_reply.started":"2024-01-01T08:14:38.137472Z","shell.execute_reply":"2024-01-01T08:14:38.148385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Neptune project, our MLOps stack.","metadata":{}},{"cell_type":"code","source":"# Get neptune Framework to log (loss, val_loss, and alot more) for each run.\nimport neptune\n\n#Connect your Neptune project with your code\nrun = neptune.init_run(\n    #Your Project Name in Neptune.ai\n    project=\"ASL-/ASL\",\n    #Your API_TOKEN, this will help you: https://medium.com/@darragh.hanley_94135/mastering-mlops-with-neptune-ai-84e635d36bf2\n    api_token=\"eyJhcGlfYWRkcmVzcyI6Imh0dHBzOi8vYXBwLm5lcHR1bmUuYWkiLCJhcGlfdXJsIjoiaHR0cHM6Ly9hcHAubmVwdHVuZS5haSIsImFwaV9rZXkiOiI1NWFkNWMwNi00MTMzLTRiZGMtYjIwZi1jNGI0ZTU1MDVjNDYifQ==\",\n)  ","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:38.151859Z","iopub.execute_input":"2024-01-01T08:14:38.152191Z","iopub.status.idle":"2024-01-01T08:14:40.779679Z","shell.execute_reply.started":"2024-01-01T08:14:38.152158Z","shell.execute_reply":"2024-01-01T08:14:40.778706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plot Config","metadata":{}},{"cell_type":"code","source":"# MatplotLib Global Settings\nmpl.rcParams.update(mpl.rcParamsDefault)\nmpl.rcParams['xtick.labelsize'] = 16\nmpl.rcParams['ytick.labelsize'] = 16\nmpl.rcParams['axes.labelsize'] = 18\nmpl.rcParams['axes.titlesize'] = 24","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:40.78111Z","iopub.execute_input":"2024-01-01T08:14:40.781452Z","iopub.status.idle":"2024-01-01T08:14:40.788189Z","shell.execute_reply.started":"2024-01-01T08:14:40.781416Z","shell.execute_reply":"2024-01-01T08:14:40.787276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA and PreProcessing","metadata":{}},{"cell_type":"markdown","source":"## Utils","metadata":{}},{"cell_type":"code","source":"# Prints Shape and Dtype For List Of Variables\ndef print_shape_dtype(l, names):\n    for e, n in zip(l, names):\n        print(f'{n} shape: {e.shape}, dtype: {e.dtype}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:40.789354Z","iopub.execute_input":"2024-01-01T08:14:40.789641Z","iopub.status.idle":"2024-01-01T08:14:40.819259Z","shell.execute_reply.started":"2024-01-01T08:14:40.789615Z","shell.execute_reply":"2024-01-01T08:14:40.818355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train","metadata":{}},{"cell_type":"code","source":"# Read Train DataFrame\nif DEBUG:\n    train = pd.read_csv('/kaggle/input/asl-fingerspelling/train.csv').head(5000)\nelse:\n    train = pd.read_csv('/kaggle/input/asl-fingerspelling/train.csv')\n    \n# Set Train Indexed By sqeuence_id\ntrain_sequence_id = train.set_index('sequence_id')\n\n# Number Of Train Samples\nN_SAMPLES = len(train)\nprint(f'N_SAMPLES: {N_SAMPLES}')\n\ndisplay(train.info())\ndisplay(train.head())","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:40.820412Z","iopub.execute_input":"2024-01-01T08:14:40.820667Z","iopub.status.idle":"2024-01-01T08:14:41.016813Z","shell.execute_reply.started":"2024-01-01T08:14:40.820644Z","shell.execute_reply":"2024-01-01T08:14:41.01589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Phrase Type","metadata":{}},{"cell_type":"code","source":"\"\"\"\nAttempt to retrieve phrase type\nCould be used for pretraining or type specific inference\n *) Phone Number\\\n *) URL\n *3) Addres\n\"\"\"\ndef get_phrase_type(phrase):\n    # Phone Number\n    if re.match(r'^[\\d+-]+$', phrase):\n        return 'phone_number'\n    # url\n    elif any([substr in phrase for substr in ['www', '.', '/']]) and ' ' not in phrase:\n        return 'url'\n    # Address\n    else:\n        return 'address'\n    \ntrain['phrase_type'] = train['phrase'].apply(get_phrase_type)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:41.018298Z","iopub.execute_input":"2024-01-01T08:14:41.018649Z","iopub.status.idle":"2024-01-01T08:14:41.188971Z","shell.execute_reply.started":"2024-01-01T08:14:41.01862Z","shell.execute_reply":"2024-01-01T08:14:41.188182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## File Path","metadata":{}},{"cell_type":"code","source":"# Get complete file path to file\ndef get_file_path(path):\n    return f'/kaggle/input/asl-fingerspelling/{path}'\n\ntrain['file_path'] = train['path'].apply(get_file_path)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:41.190235Z","iopub.execute_input":"2024-01-01T08:14:41.190609Z","iopub.status.idle":"2024-01-01T08:14:41.219566Z","shell.execute_reply.started":"2024-01-01T08:14:41.190578Z","shell.execute_reply":"2024-01-01T08:14:41.218565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Example File Paths","metadata":{}},{"cell_type":"code","source":"# Unique Parquet Files\nINFERENCE_FILE_PATHS = pd.Series(\n        glob.glob('/kaggle/input/aslfr-preprocessing-dataset/train_landmark_subsets/*')\n    )\n\nprint(f'Found {len(INFERENCE_FILE_PATHS)} Inference Pickle Files')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:41.220868Z","iopub.execute_input":"2024-01-01T08:14:41.221267Z","iopub.status.idle":"2024-01-01T08:14:41.243683Z","shell.execute_reply.started":"2024-01-01T08:14:41.22123Z","shell.execute_reply":"2024-01-01T08:14:41.242767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Phrase Processing","metadata":{}},{"cell_type":"code","source":"# Split Phrase To Char Tuple\ntrain['phrase_char'] = train['phrase'].apply(tuple)\n# Character Length of Phrase\ntrain['phrase_char_len'] = train['phrase_char'].apply(len)\n\n# Maximum Input Length\nMAX_PHRASE_LENGTH = train['phrase_char_len'].max()\nprint(f'MAX_PHRASE_LENGTH: {MAX_PHRASE_LENGTH}')\n\n# Train DataFrame indexed by sequence_id to convenientlyy lookup recording data\ntrain_sequence_id = train.set_index('sequence_id')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:41.244753Z","iopub.execute_input":"2024-01-01T08:14:41.24509Z","iopub.status.idle":"2024-01-01T08:14:41.351855Z","shell.execute_reply.started":"2024-01-01T08:14:41.245063Z","shell.execute_reply":"2024-01-01T08:14:41.350798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Phrase Character Length Statistics\ndisplay(train['phrase_char_len'].describe(percentiles=PERCENTILES).to_frame().round(1))","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:41.353011Z","iopub.execute_input":"2024-01-01T08:14:41.353337Z","iopub.status.idle":"2024-01-01T08:14:41.372739Z","shell.execute_reply.started":"2024-01-01T08:14:41.353309Z","shell.execute_reply":"2024-01-01T08:14:41.371885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Character Count Occurance\nplt.figure(figsize=(15,8))\nplt.title('Character Length Occurance of Phrases')\ntrain['phrase_char_len'].value_counts().sort_index().plot(kind='bar')\nplt.xlim(-0.50, train['phrase_char_len'].max() - 1.50)\nplt.xlabel('Pharse Character Length')\nplt.ylabel('Sample Count')\nplt.grid(axis='y')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:41.377156Z","iopub.execute_input":"2024-01-01T08:14:41.377474Z","iopub.status.idle":"2024-01-01T08:14:41.898494Z","shell.execute_reply.started":"2024-01-01T08:14:41.377448Z","shell.execute_reply":"2024-01-01T08:14:41.897565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Find Unique Character","metadata":{}},{"cell_type":"code","source":"# Use Set to keep track of unique characters in phrases\nUNIQUE_CHARACTERS = set()\n\nfor phrase in tqdm(train['phrase_char']):\n    for c in phrase:\n        UNIQUE_CHARACTERS.add(c)\n        \n# Sorted Unique Character\nUNIQUE_CHARACTERS = np.array(sorted(UNIQUE_CHARACTERS))\n# Number of Unique Characters\nprint(f'N_UNIQUE_CHARACTERS: {N_UNIQUE_CHARACTERS}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:41.899974Z","iopub.execute_input":"2024-01-01T08:14:41.900556Z","iopub.status.idle":"2024-01-01T08:14:42.120985Z","shell.execute_reply.started":"2024-01-01T08:14:41.900518Z","shell.execute_reply":"2024-01-01T08:14:42.120114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Example Parquet File","metadata":{}},{"cell_type":"code","source":"# Read First Parquet File\nexample_parquet_df = pd.read_parquet(train['file_path'][0])\n\n# Each parquet file contains 1000 recordings\nprint(f'# Unique Recording: {example_parquet_df.index.nunique()}')\n# Display DataFrame layout\ndisplay(example_parquet_df.head())","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:42.122116Z","iopub.execute_input":"2024-01-01T08:14:42.122441Z","iopub.status.idle":"2024-01-01T08:14:55.108746Z","shell.execute_reply.started":"2024-01-01T08:14:42.122414Z","shell.execute_reply":"2024-01-01T08:14:55.10784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Video Statistics","metadata":{}},{"cell_type":"code","source":"# Number of parquet chunks to analyse\nN = 5 if IS_INTERACTIVE else 25\n# Number of Unique Frames in Recording\nN_UNIQUE_FRAMES = []\n\nUNIQUE_FILE_PATHS = pd.Series(train['file_path'].unique())\n\nfor idx, file_path in enumerate(tqdm(UNIQUE_FILE_PATHS.sample(N, random_state=SEED))):\n    df = pd.read_parquet(file_path)\n    for group, group_df in df.groupby('sequence_id'):\n        N_UNIQUE_FRAMES.append(group_df['frame'].nunique())\n\n# Convert to Numpy Array\nN_UNIQUE_FRAMES = np.array(N_UNIQUE_FRAMES)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:14:55.11011Z","iopub.execute_input":"2024-01-01T08:14:55.11081Z","iopub.status.idle":"2024-01-01T08:16:27.685159Z","shell.execute_reply.started":"2024-01-01T08:14:55.110768Z","shell.execute_reply":"2024-01-01T08:16:27.68381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of unique frames in each video\ndisplay(pd.Series(N_UNIQUE_FRAMES).describe(percentiles=PERCENTILES).to_frame('Value').astype(int))\n\nplt.figure(figsize=(15,8))\nplt.title('Number of Unique Frames', size=24)\npd.Series(N_UNIQUE_FRAMES).plot(kind='hist', bins=128)\nplt.grid()\nxlim = math.ceil(plt.xlim()[1])\nplt.xlim(0, xlim)\nplt.xticks(np.arange(0, xlim+50, 50))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:16:27.689906Z","iopub.execute_input":"2024-01-01T08:16:27.691251Z","iopub.status.idle":"2024-01-01T08:16:28.344727Z","shell.execute_reply.started":"2024-01-01T08:16:27.691173Z","shell.execute_reply":"2024-01-01T08:16:28.343678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# With N_TARGET_FRAMES = 256 ~85% will be below\nN_UNIQUE_FRAMES_WATERFALL = []\n# Maximum Number of Unique Frames to use\nN_MAX_UNIQUE_FRAMES = 400\n# Compute Percentage\nfor n in tqdm(range(0,N_MAX_UNIQUE_FRAMES+1)):\n    N_UNIQUE_FRAMES_WATERFALL.append(sum(N_UNIQUE_FRAMES >= n) / len(N_UNIQUE_FRAMES) * 100)\n\nplt.figure(figsize=(18,10))\nplt.title('Waterfall Plot For Number Of Unique Frames')\npd.Series(N_UNIQUE_FRAMES_WATERFALL).plot(kind='bar')\nplt.grid(axis='y')\nplt.xticks([1] + np.arange(5, N_MAX_UNIQUE_FRAMES+5, 5).tolist(), size=8, rotation=45)\nplt.xlabel('Number of Unique Frames', size=16)\nplt.yticks(np.arange(0, 100+5, 5), [f'{i}%' for i in range(0,100+5,5)])\nplt.ylim(0, 100)\nplt.ylabel('Percentage of Samples With At Least N Unique Frames', size=16)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:16:28.346216Z","iopub.execute_input":"2024-01-01T08:16:28.346584Z","iopub.status.idle":"2024-01-01T08:16:30.990899Z","shell.execute_reply.started":"2024-01-01T08:16:28.346551Z","shell.execute_reply":"2024-01-01T08:16:30.989977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Landmark Indices\n","metadata":{}},{"cell_type":"code","source":"def get_idxs(df, words_pos, words_neg=[], ret_names=True, idxs_pos=None):\n    idxs = []\n    names = []\n    for w in words_pos:\n        for col_idx, col in enumerate(example_parquet_df.columns):\n            # Exclude Non Landmark Columns\n            if col in ['frame']:\n                continue\n                \n            col_idx = int(col.split('_')[-1])\n            # Check if column name contains all words\n            if (w in col) and (idxs_pos is None or col_idx in idxs_pos) and all([w not in col for w in words_neg]):\n                idxs.append(col_idx)\n                names.append(col)\n    # Convert to Numpy arrays\n    idxs = np.array(idxs)\n    names = np.array(names)\n    # Returns either both column indices and names\n    if ret_names:\n        return idxs, names\n    # Or only columns indices\n    else:\n        return idxs","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:16:30.992447Z","iopub.execute_input":"2024-01-01T08:16:30.992794Z","iopub.status.idle":"2024-01-01T08:16:31.001808Z","shell.execute_reply.started":"2024-01-01T08:16:30.992762Z","shell.execute_reply":"2024-01-01T08:16:31.000478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lips Landmark Face Ids\nLIPS_LANDMARK_IDXS = np.array([\n        61, 185, 40, 39, 37, 0, 267, 269, 270, 409,\n        291, 146, 91, 181, 84, 17, 314, 405, 321, 375,\n        78, 191, 80, 81, 82, 13, 312, 311, 310, 415,\n        95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n    ])\n\n# Landmark Indices for Left/Right hand without z axis in raw data\nLEFT_HAND_IDXS0, LEFT_HAND_NAMES0 = get_idxs(example_parquet_df, ['left_hand'], ['z'])\nRIGHT_HAND_IDXS0, RIGHT_HAND_NAMES0 = get_idxs(example_parquet_df, ['right_hand'], ['z'])\nLIPS_IDXS0, LIPS_NAMES0 = get_idxs(example_parquet_df, ['face'], ['z'], idxs_pos=LIPS_LANDMARK_IDXS)\nCOLUMNS0 = np.concatenate((LEFT_HAND_NAMES0, RIGHT_HAND_NAMES0, LIPS_NAMES0))\nN_COLS0 = len(COLUMNS0)\n# Only X/Y axes are used\nN_DIMS0 = 2\n\nprint(f'N_COLS0: {N_COLS0}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:16:31.003331Z","iopub.execute_input":"2024-01-01T08:16:31.003821Z","iopub.status.idle":"2024-01-01T08:16:31.028061Z","shell.execute_reply.started":"2024-01-01T08:16:31.003787Z","shell.execute_reply":"2024-01-01T08:16:31.027167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Landmark Indices in subset of dataframe with only COLUMNS selected\nLEFT_HAND_IDXS = np.argwhere(np.isin(COLUMNS0, LEFT_HAND_NAMES0)).squeeze()\nRIGHT_HAND_IDXS = np.argwhere(np.isin(COLUMNS0, RIGHT_HAND_NAMES0)).squeeze()\nLIPS_IDXS = np.argwhere(np.isin(COLUMNS0, LIPS_NAMES0)).squeeze()\nN_COLS = N_COLS0\n# Only X/Y axes are used\nN_DIMS = 2\n\nprint(f'N_COLS: {N_COLS}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:16:31.029271Z","iopub.execute_input":"2024-01-01T08:16:31.029604Z","iopub.status.idle":"2024-01-01T08:16:31.044425Z","shell.execute_reply.started":"2024-01-01T08:16:31.029576Z","shell.execute_reply":"2024-01-01T08:16:31.043489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Indices in processed data by axes with only dominant hand\nHAND_X_IDXS = np.array(\n        [idx for idx, name in enumerate(LEFT_HAND_NAMES0) if 'x' in name]\n    ).squeeze()\nHAND_Y_IDXS = np.array(\n        [idx for idx, name in enumerate(LEFT_HAND_NAMES0) if 'y' in name]\n    ).squeeze()\n# Names in processed data by axes\nHAND_X_NAMES = LEFT_HAND_NAMES0[HAND_X_IDXS]\nHAND_Y_NAMES = LEFT_HAND_NAMES0[HAND_Y_IDXS]","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:16:31.045629Z","iopub.execute_input":"2024-01-01T08:16:31.046153Z","iopub.status.idle":"2024-01-01T08:16:31.052659Z","shell.execute_reply.started":"2024-01-01T08:16:31.046099Z","shell.execute_reply":"2024-01-01T08:16:31.051939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Number Of Non-NaN Frames","metadata":{}},{"cell_type":"code","source":"\"\"\"\n    Tensorflow layer to process data in TFLite\n    Data needs to be processed in the model itself, so we can not use Python\n\"\"\" \nclass PreprocessLayerNonNaN(tf.keras.layers.Layer):\n    def __init__(self):\n        super(PreprocessLayerNonNaN, self).__init__()\n    \n    @tf.function(\n        input_signature=(tf.TensorSpec(shape=[None,N_COLS0], dtype=tf.float32),),\n    )\n    def call(self, data0):\n        # Fill NaN Values With 0\n        data = tf.where(tf.math.is_nan(data0), 0.0, data0)\n        \n        # Hacky\n        data = data[None]\n        \n        # Empty Hand Frame Filtering\n        hands = tf.slice(data, [0,0,0], [-1, -1, 84])\n        hands = tf.abs(hands)\n        mask = tf.reduce_sum(hands, axis=2)\n        mask = tf.not_equal(mask, 0)\n        data = data[mask][None]\n        data = tf.squeeze(data, axis=[0])\n        \n        return data\n    \npreprocess_layer_non_nan = PreprocessLayerNonNaN()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:16:31.053879Z","iopub.execute_input":"2024-01-01T08:16:31.054482Z","iopub.status.idle":"2024-01-01T08:16:31.085682Z","shell.execute_reply.started":"2024-01-01T08:16:31.054438Z","shell.execute_reply":"2024-01-01T08:16:31.08478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Unique Parquet Files\nUNIQUE_FILE_PATHS = pd.Series(train['file_path'].unique())\n# Number of parquet chunks to analyse\nN = 5 if (IS_INTERACTIVE or FAST) else len(UNIQUE_FILE_PATHS)\n# Number of Non Nan Frames in Recording\nN_NON_NAN_FRAMES = []\n\nfor idx, file_path in enumerate(tqdm(UNIQUE_FILE_PATHS.sample(N, random_state=SEED))):\n    df = pd.read_parquet(file_path)\n    for group, group_df in df.groupby('sequence_id'):\n        frames = preprocess_layer_non_nan(group_df[COLUMNS0].values).numpy()\n        N_NON_NAN_FRAMES.append(len(frames))\n\n# Convert to Numpy Array\nN_NON_NAN_FRAMES = pd.Series(N_NON_NAN_FRAMES).to_frame('# Frames')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:16:31.086736Z","iopub.execute_input":"2024-01-01T08:16:31.087035Z","iopub.status.idle":"2024-01-01T08:16:58.687873Z","shell.execute_reply.started":"2024-01-01T08:16:31.087009Z","shell.execute_reply":"2024-01-01T08:16:58.686919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of frames in each video with hand coordinates\ndisplay(N_NON_NAN_FRAMES.describe(percentiles=PERCENTILES).astype(int))\n\nN_NON_NAN_FRAMES.plot(kind='hist', bins=128, figsize=(15,8))\nplt.title('Number of Non NaN Frames', size=24)\nplt.grid()\nxlim = np.percentile(N_NON_NAN_FRAMES, 99)\nplt.xlim(0, xlim)\nplt.xticks(np.arange(0, xlim+32, 32))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:16:58.689093Z","iopub.execute_input":"2024-01-01T08:16:58.689414Z","iopub.status.idle":"2024-01-01T08:16:59.279622Z","shell.execute_reply.started":"2024-01-01T08:16:58.689387Z","shell.execute_reply":"2024-01-01T08:16:59.278577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Tensorflow Preprocess Layer","metadata":{}},{"cell_type":"code","source":"\"\"\"\n    Tensorflow layer to process data in TFLite\n    Data needs to be processed in the model itself, so we can not use Python\n\"\"\" \nclass PreprocessLayer(tf.keras.layers.Layer):\n    def __init__(self):\n        super(PreprocessLayer, self).__init__()\n        self.normalisation_correction = tf.constant(\n                    # Add 0.50 to x coordinates of left hand (original right hand) and substract 0.50 of right hand (original left hand)\n                     [0.50 if 'x' in name else 0.00 for name in LEFT_HAND_NAMES0],\n                dtype=tf.float32,\n            )\n    \n    @tf.function(\n        input_signature=(tf.TensorSpec(shape=[None,N_COLS0], dtype=tf.float32),),\n    )\n    def call(self, data0, resize=True):\n        # Fill NaN Values With 0\n        data = tf.where(tf.math.is_nan(data0), 0.0, data0)\n        \n        # Hacky\n        data = data[None]\n        \n        # Empty Hand Frame Filtering\n        hands = tf.slice(data, [0,0,0], [-1, -1, 84])\n        hands = tf.abs(hands)\n        mask = tf.reduce_sum(hands, axis=2)\n        mask = tf.not_equal(mask, 0)\n        data = data[mask][None]\n        \n        # Pad Zeros\n        N_FRAMES = len(data[0])\n        if N_FRAMES < N_TARGET_FRAMES:\n            data = tf.concat((\n                data,\n                tf.zeros([1,N_TARGET_FRAMES-N_FRAMES,N_COLS], dtype=tf.float32)\n            ), axis=1)\n        # Downsample\n        data = tf.image.resize(\n            data,\n            [1, N_TARGET_FRAMES],\n            method=tf.image.ResizeMethod.BILINEAR,\n        )\n        \n        # Squeeze Batch Dimension\n        data = tf.squeeze(data, axis=[0])\n        \n        return data\n    \npreprocess_layer = PreprocessLayer()\n\ninputs = group_df[COLUMNS0].values\ninputs = inputs[:1]\n\nframes = preprocess_layer(inputs)\n\nprint(f'inputs shape: {inputs.shape}')\nprint(f'frames shape: {frames.shape}, NaN count: {np.isnan(frames).sum()}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:16:59.281002Z","iopub.execute_input":"2024-01-01T08:16:59.281413Z","iopub.status.idle":"2024-01-01T08:16:59.553857Z","shell.execute_reply.started":"2024-01-01T08:16:59.28136Z","shell.execute_reply":"2024-01-01T08:16:59.552801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create X/Y","metadata":{}},{"cell_type":"code","source":"# Target Arrays Processed Input Videos\nX = np.zeros([N_SAMPLES, N_TARGET_FRAMES, N_COLS], dtype=np.float32)\n# Ordinally Encoded Target With value 59 for pad token\ny = np.full(shape=[N_SAMPLES, N_TARGET_FRAMES], fill_value=N_UNIQUE_CHARACTERS, dtype=np.int8)\n# Phrase Type\ny_phrase_type = np.empty(shape=[N_SAMPLES], dtype=object)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:16:59.555207Z","iopub.execute_input":"2024-01-01T08:16:59.555521Z","iopub.status.idle":"2024-01-01T08:16:59.563188Z","shell.execute_reply.started":"2024-01-01T08:16:59.555493Z","shell.execute_reply":"2024-01-01T08:16:59.562181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# All Unique Parquet Files\nUNIQUE_FILE_PATHS = pd.Series(train['file_path'].unique())\nN_UNIQUE_FILE_PATHS = len(UNIQUE_FILE_PATHS)\n# Counter to keep track of sample\nrow = 0\ncount = 0\n# Compressed Parquet Files\nPath('train_landmark_subsets').mkdir(parents=True, exist_ok=True)\n# Numbre Of Frames Per Character\nN_FRAMES_PER_CHARACTER = []\n# Minimum Number Of Frames Per Character\nMIN_NUM_FRAMES_PER_CHARACTER = 4\nVALID_IDXS = []\n\n# Fill Arrays\nfor idx, file_path in enumerate(tqdm(UNIQUE_FILE_PATHS)):\n    # Progress Logging\n    print(f'Processed {idx:02d}/{N_UNIQUE_FILE_PATHS} parquet files')\n    # Read parquet file\n    df = pd.read_parquet(file_path)\n    # Save COLUMN Subset of parquet files for TFLite Model verficiation\n    name = file_path.split('/')[-1]\n    if idx < 10:\n        df[COLUMNS0].to_parquet(f'train_landmark_subsets/{name}', engine='pyarrow', compression='zstd')\n    # Iterate Over Samples\n    for group, group_df in df.groupby('sequence_id'):\n        # Number of Frames Per Character\n        n_frames_per_character =  len(group_df[COLUMNS0].values) / len(train_sequence_id.loc[group, 'phrase_char'])\n        N_FRAMES_PER_CHARACTER.append(n_frames_per_character)\n        if n_frames_per_character < MIN_NUM_FRAMES_PER_CHARACTER:\n            count = count + 1\n            continue\n        else:\n            # Add Valid Index\n            VALID_IDXS.append(count)\n            count = count + 1\n        \n        # Get Processed Frames and non empty frame indices\n        frames = preprocess_layer(group_df[COLUMNS0].values)\n        assert frames.ndim == 2\n        # Assign\n        X[row] = frames\n        # Add Target By Ordinally Encoding Characters\n        phrase_char = train_sequence_id.loc[group, 'phrase_char']\n        for col, char in enumerate(phrase_char):\n            y[row, col] = CHAR2ORD.get(char)\n        # Add EOS Token\n        y[row, col+1] = EOS_TOKEN\n        # Phrase Type\n        y_phrase_type[row] = train_sequence_id.loc[group, 'phrase_type']\n        # Row Count\n        row += 1\n    # clean up\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:16:59.56472Z","iopub.execute_input":"2024-01-01T08:16:59.56509Z","iopub.status.idle":"2024-01-01T08:41:37.736979Z","shell.execute_reply.started":"2024-01-01T08:16:59.565053Z","shell.execute_reply":"2024-01-01T08:41:37.735996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# rows denotes the number of samples with frames/character above threshold\nprint(f'row: {row}, count: {count}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:41:37.738458Z","iopub.execute_input":"2024-01-01T08:41:37.738847Z","iopub.status.idle":"2024-01-01T08:41:37.744982Z","shell.execute_reply.started":"2024-01-01T08:41:37.73881Z","shell.execute_reply":"2024-01-01T08:41:37.744024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Example target, note the phrase is padded with the pad token 59\nprint(f'Example Target: {y[0]}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:41:37.746306Z","iopub.execute_input":"2024-01-01T08:41:37.746705Z","iopub.status.idle":"2024-01-01T08:41:37.76948Z","shell.execute_reply.started":"2024-01-01T08:41:37.74667Z","shell.execute_reply":"2024-01-01T08:41:37.768604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filer X/y\nX = X[:row]\ny = y[:row]","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:41:37.770642Z","iopub.execute_input":"2024-01-01T08:41:37.770929Z","iopub.status.idle":"2024-01-01T08:41:37.780932Z","shell.execute_reply.started":"2024-01-01T08:41:37.770904Z","shell.execute_reply":"2024-01-01T08:41:37.780051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save X/y\nnp.save('X.npy', X)\nnp.save('y.npy', y)\n# Save Validation\nsplitter = GroupShuffleSplit(test_size=0.10, n_splits=2, random_state=SEED)\nPARTICIPANT_IDS = train['participant_id'].values[VALID_IDXS]\ntrain_idxs, val_idxs = next(splitter.split(X, y, groups=PARTICIPANT_IDS))\n\n# Save Train\nnp.save('X_train.npy', X[train_idxs])\nnp.save('y_train.npy', y[train_idxs])\n# Save Validation\nnp.save('X_val.npy', X[val_idxs])\nnp.save('y_val.npy', y[val_idxs])\n# Verify Train/Val is correctly split by participan id\nprint(f'Patient ID Intersection Train/Val: {set(PARTICIPANT_IDS[train_idxs]).intersection(PARTICIPANT_IDS[val_idxs])}')\n# Train/Val Sizes\nprint(f'# Train Samples: {len(train_idxs)}, # Val Samples: {len(val_idxs)}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:41:37.782002Z","iopub.execute_input":"2024-01-01T08:41:37.782305Z","iopub.status.idle":"2024-01-01T08:42:17.962812Z","shell.execute_reply.started":"2024-01-01T08:41:37.782274Z","shell.execute_reply":"2024-01-01T08:42:17.961817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Number of Frames Per Character","metadata":{}},{"cell_type":"code","source":"N_FRAMES_PER_CHARACTER_S = pd.Series(N_FRAMES_PER_CHARACTER)\n\ndisplay(N_FRAMES_PER_CHARACTER_S.describe(percentiles=PERCENTILES).to_frame('Value').round(2))\n\nplt.figure(figsize=(20,10))\nplt.title('Number Of Frames Per Phrase Character')\nN_FRAMES_PER_CHARACTER_S.plot(kind='hist', bins=128)\n# Plot till 99th percentile\np99 = math.ceil(np.percentile(N_FRAMES_PER_CHARACTER_S, 99))\nplt.xticks(np.arange(0, p99+1, 1))\nplt.xlim(0, p99)\nplt.xlabel('Number Of Frames Per Phrase Character')\nplt.ylabel('Sample Count')\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:42:17.964342Z","iopub.execute_input":"2024-01-01T08:42:17.965197Z","iopub.status.idle":"2024-01-01T08:42:18.611601Z","shell.execute_reply.started":"2024-01-01T08:42:17.965156Z","shell.execute_reply":"2024-01-01T08:42:18.610648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Coordinate Statistics","metadata":{}},{"cell_type":"code","source":"def get_left_right_hand_mean_std():\n    # Dominant Hand Statistics\n    MEANS = np.zeros([N_COLS], dtype=np.float32)\n    STDS = np.zeros([N_COLS], dtype=np.float32)\n    \n    # Plot\n    fig, axes = plt.subplots(3, figsize=(20, 3*8))\n    \n    # Iterate over all landmarks\n    for col, v in enumerate(tqdm(X.reshape([-1, N_COLS]).T)):\n        v = v[np.nonzero(v)]\n        # Remove zero values as they are NaN values\n        MEANS[col] = v.astype(np.float32).mean()\n        STDS[col] = v.astype(np.float32).std()\n        if col in LEFT_HAND_IDXS:\n            axes[0].boxplot(v, notch=False, showfliers=False, positions=[col], whis=[5,95])\n        elif col in RIGHT_HAND_IDXS:\n            axes[1].boxplot(v, notch=False, showfliers=False, positions=[col], whis=[5,95])\n        else:\n            axes[2].boxplot(v, notch=False, showfliers=False, positions=[col], whis=[5,95])\n        \n    for ax, name in zip(axes, ['Left Hand', 'Right Hand', 'Lips']):\n        ax.set_title(f'{name}', size=24)\n        ax.tick_params(axis='x', labelsize=8, rotation=45)\n        ax.set_ylim(0.0, 1.0)\n        ax.grid(axis='y')\n\n    plt.show()\n    \n    return MEANS, STDS\n\n# Get Dominant Hand Mean/Standard Deviation\nMEANS, STDS = get_left_right_hand_mean_std()\n# Save Mean/STD to normalize input in neural network model\nnp.save('MEANS.npy', MEANS)\nnp.save('STDS.npy', STDS)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:42:18.613041Z","iopub.execute_input":"2024-01-01T08:42:18.613854Z","iopub.status.idle":"2024-01-01T08:43:40.946228Z","shell.execute_reply.started":"2024-01-01T08:42:18.613816Z","shell.execute_reply":"2024-01-01T08:43:40.945233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load X/y","metadata":{}},{"cell_type":"code","source":"# Train/Validation\nif USE_VAL:\n    # TRAIN\n    X_train = np.load('/kaggle/input/aslfr-preprocessing-dataset/X_train.npy')\n    y_train = np.load('/kaggle/input/aslfr-preprocessing-dataset/y_train.npy')[:,:MAX_PHRASE_LENGTH]\n    N_TRAIN_SAMPLES = len(X_train)\n    # VAL\n    X_val = np.load('/kaggle/input/aslfr-preprocessing-dataset/X_val.npy')\n    y_val = np.load('/kaggle/input/aslfr-preprocessing-dataset/y_val.npy')[:,:MAX_PHRASE_LENGTH]\n    N_VAL_SAMPLES = len(X_val)\n    # Shapes\n    print(f'X_train shape: {X_train.shape}, X_val shape: {X_val.shape}')\n# Train On All Data\nelse:\n    # TRAIN\n    X_train = np.load('/kaggle/input/aslfr-preprocessing-dataset/X.npy')\n    y_train = np.load('/kaggle/input/aslfr-preprocessing-dataset/y.npy')[:,:MAX_PHRASE_LENGTH]\n    N_TRAIN_SAMPLES = len(X_train)\n    print(f'X_train shape: {X_train.shape}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:43:40.947515Z","iopub.execute_input":"2024-01-01T08:43:40.947817Z","iopub.status.idle":"2024-01-01T08:44:40.905248Z","shell.execute_reply.started":"2024-01-01T08:43:40.94779Z","shell.execute_reply":"2024-01-01T08:44:40.90417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Example Batch","metadata":{}},{"cell_type":"code","source":"# Example Batch For Debugging\nN_EXAMPLE_BATCH_SAMPLES = 1024\nN_EXAMPLE_BATCH_SAMPLES_SMALL = 32\n# Example Batch\nX_batch = {\n    'frames': np.copy(X_train[:N_EXAMPLE_BATCH_SAMPLES]),\n    'phrase': np.copy(y_train[:N_EXAMPLE_BATCH_SAMPLES]),\n#     'phrase_type': np.copy(y_phrase_type_train[:N_EXAMPLE_BATCH_SAMPLES]),\n}\ny_batch = np.copy(y_train[:N_EXAMPLE_BATCH_SAMPLES])\n# Small Example Batch\nX_batch_small = {\n    'frames': np.copy(X_train[:N_EXAMPLE_BATCH_SAMPLES_SMALL]),\n    'phrase': np.copy(y_train[:N_EXAMPLE_BATCH_SAMPLES_SMALL]),\n#     'phrase_type': np.copy(y_phrase_type_train[:N_EXAMPLE_BATCH_SAMPLES_SMALL]),\n}\ny_batch_small = np.copy(y_train[:N_EXAMPLE_BATCH_SAMPLES_SMALL])","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:40.906607Z","iopub.execute_input":"2024-01-01T08:44:40.906943Z","iopub.status.idle":"2024-01-01T08:44:41.065165Z","shell.execute_reply.started":"2024-01-01T08:44:40.906914Z","shell.execute_reply":"2024-01-01T08:44:41.064176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Mean/STD Loading","metadata":{}},{"cell_type":"code","source":"# Mean/Standard Deviations of data used for normalizing\nMEANS = np.load('/kaggle/input/aslfr-preprocessing-dataset/MEANS.npy').reshape(-1)\nSTDS = np.load('/kaggle/input/aslfr-preprocessing-dataset/STDS.npy').reshape(-1)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.066538Z","iopub.execute_input":"2024-01-01T08:44:41.066894Z","iopub.status.idle":"2024-01-01T08:44:41.085194Z","shell.execute_reply.started":"2024-01-01T08:44:41.06686Z","shell.execute_reply":"2024-01-01T08:44:41.084262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Test Preprocessing Layer","metadata":{}},{"cell_type":"code","source":"# Function To Test Preprocessing Layer\ndef test_preprocess_layer():\n    demo_sequence_id = example_parquet_df.index.unique()[15]\n    demo_raw_data = example_parquet_df.loc[demo_sequence_id, COLUMNS0]\n    data = preprocess_layer(demo_raw_data)\n\n    print(f'demo_raw_data shape: {demo_raw_data.shape}')\n    print(f'data shape: {data.shape}')\n    \n    return data\n    \nif IS_INTERACTIVE:\n    data = test_preprocess_layer()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.086406Z","iopub.execute_input":"2024-01-01T08:44:41.086701Z","iopub.status.idle":"2024-01-01T08:44:41.105343Z","shell.execute_reply.started":"2024-01-01T08:44:41.086676Z","shell.execute_reply":"2024-01-01T08:44:41.104364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train Dataset","metadata":{}},{"cell_type":"code","source":"# Train Dataset Iterator\ndef get_train_dataset(X, y, batch_size=BATCH_SIZE):\n    sample_idxs = np.arange(len(X))\n    while True:\n        # Get random indices\n        random_sample_idxs = np.random.choice(sample_idxs, batch_size)\n        \n        inputs = {\n            'frames': X[random_sample_idxs],\n            'phrase': y[random_sample_idxs],\n        }\n        outputs = y[random_sample_idxs]\n        \n        yield inputs, outputs","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.106412Z","iopub.execute_input":"2024-01-01T08:44:41.10669Z","iopub.status.idle":"2024-01-01T08:44:41.112652Z","shell.execute_reply.started":"2024-01-01T08:44:41.106666Z","shell.execute_reply":"2024-01-01T08:44:41.111802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train Dataset\ntrain_dataset = get_train_dataset(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.124285Z","iopub.execute_input":"2024-01-01T08:44:41.124641Z","iopub.status.idle":"2024-01-01T08:44:41.128649Z","shell.execute_reply.started":"2024-01-01T08:44:41.124617Z","shell.execute_reply":"2024-01-01T08:44:41.12769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Training Steps Per Epoch\nTRAIN_STEPS_PER_EPOCH = math.ceil(N_TRAIN_SAMPLES / BATCH_SIZE)\nprint(f'TRAIN_STEPS_PER_EPOCH: {TRAIN_STEPS_PER_EPOCH}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.12978Z","iopub.execute_input":"2024-01-01T08:44:41.130056Z","iopub.status.idle":"2024-01-01T08:44:41.140311Z","shell.execute_reply.started":"2024-01-01T08:44:41.130031Z","shell.execute_reply":"2024-01-01T08:44:41.139292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Validation Dataset","metadata":{}},{"cell_type":"code","source":"# Validation Set\ndef get_val_dataset(X, y, batch_size=BATCH_SIZE):\n    offsets = np.arange(0, len(X), batch_size)\n    while True:\n        # Iterate over whole validation set\n        for offset in offsets:\n            inputs = {\n                'frames': X[offset:offset+batch_size],\n                'phrase': y[offset:offset+batch_size],\n            }\n            outputs = y[offset:offset+batch_size]\n\n            yield inputs, outputs","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.141472Z","iopub.execute_input":"2024-01-01T08:44:41.141809Z","iopub.status.idle":"2024-01-01T08:44:41.150919Z","shell.execute_reply.started":"2024-01-01T08:44:41.141776Z","shell.execute_reply":"2024-01-01T08:44:41.150163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Validation Dataset\nif USE_VAL:\n    val_dataset = get_val_dataset(X_val, y_val)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.152025Z","iopub.execute_input":"2024-01-01T08:44:41.152354Z","iopub.status.idle":"2024-01-01T08:44:41.163233Z","shell.execute_reply.started":"2024-01-01T08:44:41.152329Z","shell.execute_reply":"2024-01-01T08:44:41.162248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if USE_VAL:\n    N_VAL_STEPS_PER_EPOCH = math.ceil(N_VAL_SAMPLES / BATCH_SIZE)\n    print(f'N_VAL_STEPS_PER_EPOCH: {N_VAL_STEPS_PER_EPOCH}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.164357Z","iopub.execute_input":"2024-01-01T08:44:41.16471Z","iopub.status.idle":"2024-01-01T08:44:41.173668Z","shell.execute_reply.started":"2024-01-01T08:44:41.164672Z","shell.execute_reply":"2024-01-01T08:44:41.172806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Config","metadata":{}},{"cell_type":"code","source":"# Epsilon value for layer normalisation\nLAYER_NORM_EPS = 1e-6\n\n# final embedding and transformer embedding size\nUNITS_ENCODER = 384 #Dimesion of Encoder\nUNITS_DECODER = 256 #Dimesion of Dncoder\n\n# Transformer\nNUM_BLOCKS_ENCODER = 4 #Encoder Blocks \nNUM_BLOCKS_DECODER = 2 #Decoder Blocks\nNUM_HEADS = 4 # Number of Attention Heads\nMLP_RATIO = 2 # Multi Layer Perception Ration\n\n# Dropout\nEMBEDDING_DROPOUT = 0.00 #Embedding Dropout Ration \nMLP_DROPOUT_RATIO = 0.30 #Multi Layer Perception Dropout Ratio\nMHA_DROPOUT_RATIO = 0.20 #MultiHead Attention Dropout Ration\nCLASSIFIER_DROPOUT_RATIO = 0.10 # Classifier Dropout Ration\n\n# Initiailizers\nINIT_HE_UNIFORM = tf.keras.initializers.he_uniform #He Uniform Weights Initializer, It helps avoid the vanishing/exploding gradient problems at the start of the training. \n                                                   #By maintaining a controlled variance, it ensures that the signals don't get too small or too large during the initial stages.\n\nINIT_GLOROT_UNIFORM = tf.keras.initializers.glorot_uniform #Glorot Uniform Initializer By maintaining the variance of activations and back-propagated gradients, it helps avoid the vanishing/exploding gradient problem, \n                                                           #especially in deep networks. \n                                                           #It's suitable for networks with sigmoid or hyperbolic tangent activation functions.\n\nINIT_ZEROS = tf.keras.initializers.constant(0.0) #Zeros Initializer It initializes the weights to all zeros.\n\n# Activations\nGELU = tf.keras.activations.gelu #Gaussian Error Linear Unit activation function.\n\n# Landmark Embedding","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Embeds a landmark using fully connected layers\nclass LandmarkEmbedding(tf.keras.Model):\n    def __init__(self, units, name):\n        super(LandmarkEmbedding, self).__init__(name=f'{name}_embedding')\n        self.units = units\n        self.supports_masking = True\n        \n    def build(self, input_shape):\n        # Embedding for missing landmark in frame, initizlied with zeros\n        self.empty_embedding = self.add_weight(\n            name=f'{self.name}_empty_embedding',\n            shape=[self.units],\n            initializer=INIT_ZEROS,\n        )\n        # Embedding\n        self.dense = tf.keras.Sequential([\n            tf.keras.layers.Dense(self.units, name=f'{self.name}_dense_1', use_bias=False, kernel_initializer=INIT_GLOROT_UNIFORM, activation=GELU),\n            tf.keras.layers.Dense(self.units, name=f'{self.name}_dense_2', use_bias=False, kernel_initializer=INIT_HE_UNIFORM),\n        ], name=f'{self.name}_dense')\n\n    def call(self, x):\n        return tf.where(\n                # Checks whether landmark is missing in frame\n                tf.reduce_sum(x, axis=2, keepdims=True) == 0,\n                # If so, the empty embedding is used\n                self.empty_embedding,\n                # Otherwise the landmark data is embedded\n                self.dense(x),\n            )","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.185776Z","iopub.execute_input":"2024-01-01T08:44:41.186026Z","iopub.status.idle":"2024-01-01T08:44:41.200284Z","shell.execute_reply.started":"2024-01-01T08:44:41.186003Z","shell.execute_reply":"2024-01-01T08:44:41.199463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Embedding","metadata":{}},{"cell_type":"code","source":"# Creates embedding for each frame\nclass Embedding(tf.keras.Model):\n    def __init__(self):\n        super(Embedding, self).__init__()\n        self.supports_masking = True\n    \n    def build(self, input_shape):\n        # Positional embedding for each frame index\n        self.positional_embedding = tf.Variable(\n            initial_value=tf.zeros([N_TARGET_FRAMES, UNITS_ENCODER], dtype=tf.float32),\n            trainable=True,\n            name='embedding_positional_encoder',\n        )\n        # Embedding layer for Landmarks\n        self.dominant_hand_embedding = LandmarkEmbedding(UNITS_ENCODER, 'dominant_hand')\n\n    def call(self, x, training=False):\n        # Normalize\n        x = tf.where(\n                tf.math.equal(x, 0.0),\n                0.0,\n                (x - MEANS) / STDS,\n            )\n        # Dominant Hand\n        x = self.dominant_hand_embedding(x)\n        # Add Positional Encoding\n        x = x + self.positional_embedding\n        \n        return x","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.201322Z","iopub.execute_input":"2024-01-01T08:44:41.201583Z","iopub.status.idle":"2024-01-01T08:44:41.215019Z","shell.execute_reply.started":"2024-01-01T08:44:41.20156Z","shell.execute_reply":"2024-01-01T08:44:41.214198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Transformer","metadata":{}},{"cell_type":"code","source":"# based on: https://stackoverflow.com/questions/67342988/verifying-the-implementation-of-multihead-attention-in-transformer\n# replaced softmax with softmax layer to support masked softmax\nclass MultiHeadAttention(tf.keras.layers.Layer):\n    def __init__(self, d_model, n_heads, dropout, d_out=None):\n        super(MultiHeadAttention,self).__init__()\n        # Number of Units in Model\n        self.d_model = d_model\n        # Number of Attention Heads\n        self.n_heads = n_heads\n        # Number of Units in Intermediate Layers\n        self.depth = d_model // 2\n        # Scaling Factor Of Values\n        self.scale = 1.0 / tf.math.sqrt(tf.cast(self.depth, tf.float32))\n        # Learnable Projections to Depth\n        self.wq = self.fused_mha(self.depth)\n        self.wk = self.fused_mha(self.depth)\n        self.wv = self.fused_mha(self.depth)\n        # Output Projection\n        self.wo = tf.keras.layers.Dense(d_model if d_out is None else d_out, use_bias=False)\n        # Softmax Activation Which Supports Masking\n        self.softmax = tf.keras.layers.Softmax()\n        # Reshaping Of Multiple Attention heads to Single Value\n        self.reshape = tf.keras.Sequential([\n            # [attention heads, number of frames, d_model] → [number of frames, n_heads, d_model // n_heads]\n            tf.keras.layers.Permute([2, 1, 3]),\n            # [number of frames, attention heads, d_model] → [number of frames, d_model]\n            tf.keras.layers.Reshape([N_TARGET_FRAMES, self.depth]),\n        ])\n        # Output Dropout\n        self.do = tf.keras.layers.Dropout(dropout)\n        self.supports_masking = True\n        \n    # Single dense layer for all attention heads\n    def fused_mha(self, dim):\n        return tf.keras.Sequential([\n            # Single dense layer\n            tf.keras.layers.Dense(dim, use_bias=False),\n            # Reshape to [number of frames, number of attention head, depth]\n            tf.keras.layers.Reshape([N_TARGET_FRAMES, self.n_heads, dim // self.n_heads]),\n            # Permutate to [number of attention heads, number of frames, depth]\n            tf.keras.layers.Permute([2, 1, 3]),\n        ])\n        \n    def call(self, q, k, v, attention_mask=None, training=False):\n        # Projections to attention heads\n        Q = self.wq(q)\n        K = self.wk(k)\n        V = self.wv(v)\n        # Matrix multiply QxK to acquire attention scores\n        x = tf.matmul(Q, K, transpose_b=True) * self.scale\n        # Softmax attention scores and Multiply with Values\n        x = self.softmax(x, mask=attention_mask) @ V\n        # Reshape to flatten attention heads\n        x = self.reshape(x)\n        # Output projection\n        x = self.wo(x)\n        # Dropout\n        x = self.do(x, training=training)\n        return x","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.216473Z","iopub.execute_input":"2024-01-01T08:44:41.216874Z","iopub.status.idle":"2024-01-01T08:44:41.231912Z","shell.execute_reply.started":"2024-01-01T08:44:41.216839Z","shell.execute_reply":"2024-01-01T08:44:41.23114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Encoder\n\n[source](https://keras.io/examples/nlp/neural_machine_translation_with_transformer/)","metadata":{}},{"cell_type":"code","source":"# Encoder based on multiple transformer blocks\nclass Encoder(tf.keras.Model):\n    def __init__(self, num_blocks):\n        super(Encoder, self).__init__(name='encoder')\n        self.num_blocks = num_blocks\n        self.supports_masking = True\n    \n    def build(self, input_shape):\n        self.ln_1s = []\n        self.mhas = []\n        self.ln_2s = []\n        self.mlps = []\n        # Make Transformer Blocks\n        for i in range(self.num_blocks):\n            # First Layer Normalisation\n            self.ln_1s.append(tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS))\n            # Multi Head Attention\n            self.mhas.append(MultiHeadAttention(UNITS_ENCODER, NUM_HEADS, MHA_DROPOUT_RATIO))\n            # Second Layer Normalisation\n            self.ln_2s.append(tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS))\n            # Multi Layer Perception\n            self.mlps.append(tf.keras.Sequential([\n                tf.keras.layers.Dense(UNITS_ENCODER * MLP_RATIO, activation=GELU, kernel_initializer=INIT_GLOROT_UNIFORM, use_bias=False),\n                tf.keras.layers.Dropout(MLP_DROPOUT_RATIO),\n                tf.keras.layers.Dense(UNITS_ENCODER, kernel_initializer=INIT_HE_UNIFORM, use_bias=False),\n            ]))\n            # Optional Projection to Decoder Dimension\n            if UNITS_ENCODER != UNITS_DECODER:\n                self.dense_out = tf.keras.layers.Dense(UNITS_DECODER, kernel_initializer=INIT_GLOROT_UNIFORM, use_bias=False)\n                self.apply_dense_out = True\n            else:\n                self.apply_dense_out = False\n                \n    def get_attention_mask(self, x_inp):\n        # Attention Mask\n        attention_mask = tf.math.count_nonzero(x_inp, axis=[2], keepdims=True, dtype=tf.int32)\n        attention_mask = tf.math.count_nonzero(attention_mask, axis=[2], keepdims=False)\n        attention_mask = tf.expand_dims(attention_mask, axis=1)\n        attention_mask = tf.expand_dims(attention_mask, axis=1)\n        return attention_mask\n        \n    def call(self, x, x_inp, training=False):\n        # Attention mask to ignore missing frames\n        attention_mask = self.get_attention_mask(x_inp)\n        # Iterate input over transformer blocks\n        for ln_1, mha, ln_2, mlp in zip(self.ln_1s, self.mhas, self.ln_2s, self.mlps):\n            x = ln_1(x + mha(x, x, x, attention_mask=attention_mask))\n            x = ln_2(x + mlp(x))\n            \n        # Optional Projection to Decoder Dimension\n        if self.apply_dense_out:\n            x = self.dense_out(x)\n    \n        return x","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.233332Z","iopub.execute_input":"2024-01-01T08:44:41.233671Z","iopub.status.idle":"2024-01-01T08:44:41.247741Z","shell.execute_reply.started":"2024-01-01T08:44:41.233646Z","shell.execute_reply":"2024-01-01T08:44:41.246966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Decoder","metadata":{}},{"cell_type":"code","source":"# Decoder based on multiple transformer blocks\nclass Decoder(tf.keras.Model):\n    def __init__(self, num_blocks):\n        super(Decoder, self).__init__(name='decoder')\n        self.num_blocks = num_blocks\n        self.supports_masking = True\n    \n    def build(self, input_shape):\n        # Causal Mask Batch Size 1\n        self.causal_mask = self.get_causal_attention_mask()\n        # Positional Embedding, initialized with zeros\n        self.positional_embedding = tf.Variable(\n            initial_value=tf.zeros([N_TARGET_FRAMES, UNITS_DECODER], dtype=tf.float32),\n            trainable=True,\n            name='embedding_positional_encoder',\n        )\n        # Character Embedding\n        self.char_emb = tf.keras.layers.Embedding(N_UNIQUE_CHARACTERS, UNITS_DECODER, embeddings_initializer=INIT_ZEROS)\n        # Positional Encoder MHA\n        self.pos_emb_mha = MultiHeadAttention(UNITS_DECODER, NUM_HEADS, MHA_DROPOUT_RATIO)\n        self.pos_emb_ln = tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS)\n        # First Layer Normalisation\n        self.ln_1s = []\n        self.mhas = []\n        self.ln_2s = []\n        self.mlps = []\n        # Make Transformer Blocks\n        for i in range(self.num_blocks):\n            # First Layer Normalisation\n            self.ln_1s.append(tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS))\n            # Multi Head Attention\n            self.mhas.append(MultiHeadAttention(UNITS_DECODER, NUM_HEADS, MHA_DROPOUT_RATIO))\n            # Second Layer Normalisation\n            self.ln_2s.append(tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS))\n            # Multi Layer Perception\n            self.mlps.append(tf.keras.Sequential([\n                tf.keras.layers.Dense(UNITS_DECODER * MLP_RATIO, activation=GELU, kernel_initializer=INIT_GLOROT_UNIFORM, use_bias=False),\n                tf.keras.layers.Dropout(MLP_DROPOUT_RATIO),\n                tf.keras.layers.Dense(UNITS_DECODER, kernel_initializer=INIT_HE_UNIFORM, use_bias=False),\n            ]))\n            \n    def get_causal_attention_mask(self):\n        i = tf.range(N_TARGET_FRAMES)[:, tf.newaxis]\n        j = tf.range(N_TARGET_FRAMES)\n        mask = tf.cast(i >= j, dtype=tf.int32)\n        mask = tf.reshape(mask, (1, N_TARGET_FRAMES, N_TARGET_FRAMES))\n        mult = tf.concat(\n            [tf.expand_dims(1, -1), tf.constant([1, 1], dtype=tf.int32)],\n            axis=0,\n        )\n        mask = tf.tile(mask, mult)\n        mask = tf.cast(mask, tf.float32)\n        return mask\n    \n    def get_attention_mask(self, x_inp):\n        # Attention Mask\n        attention_mask = tf.math.count_nonzero(x_inp, axis=[2], keepdims=True, dtype=tf.int32)\n        attention_mask = tf.math.count_nonzero(attention_mask, axis=[2], keepdims=False)\n        attention_mask = tf.expand_dims(attention_mask, axis=1)\n        attention_mask = tf.expand_dims(attention_mask, axis=1)\n        return attention_mask\n        \n    def call(self, encoder_outputs, phrase, x_inp, training=False):\n        # Batch Size\n        B = tf.shape(encoder_outputs)[0]\n        # Cast to INT32\n        phrase = tf.cast(phrase, tf.int32)\n        # Prepend SOS Token\n        phrase = tf.pad(phrase, [[0,0], [1,0]], constant_values=SOS_TOKEN, name='prepend_sos_token')\n        # Pad With PAD Token\n        phrase = tf.pad(phrase, [[0,0], [0,N_TARGET_FRAMES-MAX_PHRASE_LENGTH-1]], constant_values=PAD_TOKEN, name='append_pad_token')\n        # Positional Embedding\n        x = self.positional_embedding + self.char_emb(phrase)\n        # Causal Attention\n        x = self.pos_emb_ln(x + self.pos_emb_mha(x, x, x, attention_mask=self.causal_mask))\n        # Attention mask to ignore missing frames\n        attention_mask = self.get_attention_mask(x_inp)\n        # Iterate input over transformer blocks\n        for ln_1, mha, ln_2, mlp in zip(self.ln_1s, self.mhas, self.ln_2s, self.mlps):\n            x = ln_1(x + mha(x, encoder_outputs, encoder_outputs, attention_mask=attention_mask))\n            x = ln_2(x + mlp(x))\n        # Slice 31 Characters\n        x = tf.slice(x, [0, 0, 0], [-1, MAX_PHRASE_LENGTH, -1])\n    \n        return x","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.249171Z","iopub.execute_input":"2024-01-01T08:44:41.249561Z","iopub.status.idle":"2024-01-01T08:44:41.272707Z","shell.execute_reply.started":"2024-01-01T08:44:41.249528Z","shell.execute_reply":"2024-01-01T08:44:41.2718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Causal Attention to make decoder not attent to future characters which it needs to predict\ndef get_causal_attention_mask(B):\n    i = tf.range(N_TARGET_FRAMES)[:, tf.newaxis]\n    j = tf.range(N_TARGET_FRAMES)\n    mask = tf.cast(i >= j, dtype=tf.int32)\n    mask = tf.reshape(mask, (1, N_TARGET_FRAMES, N_TARGET_FRAMES))\n    mult = tf.concat(\n        [tf.expand_dims(B, -1), tf.constant([1, 1], dtype=tf.int32)],\n        axis=0,\n    )\n    mask = tf.tile(mask, mult)\n    mask = tf.cast(mask, tf.float32)\n    return mask\n\nget_causal_attention_mask(1)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.273781Z","iopub.execute_input":"2024-01-01T08:44:41.274085Z","iopub.status.idle":"2024-01-01T08:44:41.420436Z","shell.execute_reply.started":"2024-01-01T08:44:41.274059Z","shell.execute_reply":"2024-01-01T08:44:41.419484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Non Pad/SOS/EOS Token Accuracy","metadata":{}},{"cell_type":"code","source":"# TopK accuracy for multi dimensional output\nclass TopKAccuracy(tf.keras.metrics.Metric):\n    def __init__(self, k, **kwargs):\n        super(TopKAccuracy, self).__init__(name=f'top{k}acc', **kwargs)\n        self.top_k_acc = tf.keras.metrics.SparseTopKCategoricalAccuracy(k=k)\n\n    def update_state(self, y_true, y_pred, sample_weight=None):\n        y_true = tf.reshape(y_true, [-1])\n        y_pred = tf.reshape(y_pred, [-1, N_UNIQUE_CHARACTERS])\n        character_idxs = tf.where(y_true < N_UNIQUE_CHARACTERS0)\n        y_true = tf.gather(y_true, character_idxs, axis=0)\n        y_pred = tf.gather(y_pred, character_idxs, axis=0)\n        self.top_k_acc.update_state(y_true, y_pred)\n\n    def result(self):\n        return self.top_k_acc.result()\n    \n    def reset_state(self):\n        self.top_k_acc.reset_state()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.421965Z","iopub.execute_input":"2024-01-01T08:44:41.422467Z","iopub.status.idle":"2024-01-01T08:44:41.432658Z","shell.execute_reply.started":"2024-01-01T08:44:41.422429Z","shell.execute_reply":"2024-01-01T08:44:41.43163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loss Weights","metadata":{}},{"cell_type":"code","source":"# Create Initial Loss Weights All Set To 1\nloss_weights = np.ones(N_UNIQUE_CHARACTERS, dtype=np.float32)\n# Set Loss Weight Of Pad Token To 0\nloss_weights[PAD_TOKEN-1] = 0","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.43418Z","iopub.execute_input":"2024-01-01T08:44:41.434472Z","iopub.status.idle":"2024-01-01T08:44:41.448266Z","shell.execute_reply.started":"2024-01-01T08:44:41.434447Z","shell.execute_reply":"2024-01-01T08:44:41.447343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Sparse Categorical Crossentropy With Label Smoothing¶","metadata":{}},{"cell_type":"code","source":"# source:: https://stackoverflow.com/questions/60689185/label-smoothing-for-sparse-categorical-crossentropy\ndef scce_with_ls(y_true, y_pred):\n    # Filter Pad Tokens\n    idxs = tf.where(y_true != PAD_TOKEN)\n    y_true = tf.gather_nd(y_true, idxs)\n    y_pred = tf.gather_nd(y_pred, idxs)\n    # One Hot Encode Sparsely Encoded Target Sign\n    y_true = tf.cast(y_true, tf.int32)\n    y_true = tf.one_hot(y_true, N_UNIQUE_CHARACTERS, axis=1)\n    # Categorical Crossentropy with native label smoothing support\n    loss = tf.keras.losses.categorical_crossentropy(y_true, y_pred, label_smoothing=0.25, from_logits=True)\n    loss = tf.math.reduce_mean(loss)\n    return loss","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.44945Z","iopub.execute_input":"2024-01-01T08:44:41.449746Z","iopub.status.idle":"2024-01-01T08:44:41.458823Z","shell.execute_reply.started":"2024-01-01T08:44:41.44972Z","shell.execute_reply":"2024-01-01T08:44:41.457922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model","metadata":{}},{"cell_type":"code","source":"def get_model():\n    # Inputs\n    frames_inp = tf.keras.layers.Input([N_TARGET_FRAMES, N_COLS], dtype=tf.float32, name='frames')\n    phrase_inp = tf.keras.layers.Input([MAX_PHRASE_LENGTH], dtype=tf.int32, name='phrase')\n    # Frames\n    x = frames_inp\n\n    # Masking\n    x = tf.keras.layers.Masking(mask_value=0.0, input_shape=(N_TARGET_FRAMES, N_COLS))(x)\n    \n    # Embedding\n    x = Embedding()(x)\n    \n    # Encoder Transformer Blocks\n    x = Encoder(NUM_BLOCKS_ENCODER)(x, frames_inp)\n    \n    # Decoder\n    x = Decoder(NUM_BLOCKS_DECODER)(x, phrase_inp, frames_inp)\n    \n    # Classifier\n    x = tf.keras.Sequential([\n        # Dropout\n        tf.keras.layers.Dropout(CLASSIFIER_DROPOUT_RATIO),\n        # Output Neurons\n        tf.keras.layers.Dense(N_UNIQUE_CHARACTERS, activation=tf.keras.activations.linear, kernel_initializer=INIT_HE_UNIFORM, use_bias=False),\n    ], name='classifier')(x)\n    \n    outputs = x\n    \n    # Create Tensorflow Model\n    model = tf.keras.models.Model(inputs=[frames_inp, phrase_inp], outputs=outputs)\n    \n    # Categorical Crossentropy Loss With Label Smoothing\n    loss = scce_with_ls\n    \n    # Adam Optimizer\n    optimizer = tfa.optimizers.RectifiedAdam(sma_threshold=4)\n    optimizer = tfa.optimizers.Lookahead(optimizer, sync_period=5)\n\n    # TopK Metrics\n    metrics = [\n        TopKAccuracy(1),\n        TopKAccuracy(5),\n    ]\n    \n    model.compile(\n        loss=loss,\n        optimizer=optimizer,\n        metrics=metrics,\n        loss_weights=loss_weights,\n    )\n    \n    return model","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.460064Z","iopub.execute_input":"2024-01-01T08:44:41.46036Z","iopub.status.idle":"2024-01-01T08:44:41.471009Z","shell.execute_reply.started":"2024-01-01T08:44:41.460335Z","shell.execute_reply":"2024-01-01T08:44:41.470179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Input data\nfor k, v in X_batch.items():\n    print(f'{k}: {v.shape}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.472334Z","iopub.execute_input":"2024-01-01T08:44:41.47267Z","iopub.status.idle":"2024-01-01T08:44:41.486214Z","shell.execute_reply.started":"2024-01-01T08:44:41.472638Z","shell.execute_reply":"2024-01-01T08:44:41.48525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.keras.backend.clear_session()\n\nmodel = get_model()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:41.487221Z","iopub.execute_input":"2024-01-01T08:44:41.487481Z","iopub.status.idle":"2024-01-01T08:44:44.922042Z","shell.execute_reply.started":"2024-01-01T08:44:41.487457Z","shell.execute_reply":"2024-01-01T08:44:44.920901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot model summary\nmodel.summary(expand_nested=True)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:44.923245Z","iopub.execute_input":"2024-01-01T08:44:44.923534Z","iopub.status.idle":"2024-01-01T08:44:45.087458Z","shell.execute_reply.started":"2024-01-01T08:44:44.923509Z","shell.execute_reply":"2024-01-01T08:44:45.086561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot Model Architecture\ntf.keras.utils.plot_model(model, show_shapes=True, show_dtype=True, show_layer_names=True, expand_nested=True, show_layer_activations=True)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:45.088597Z","iopub.execute_input":"2024-01-01T08:44:45.088886Z","iopub.status.idle":"2024-01-01T08:44:45.456087Z","shell.execute_reply.started":"2024-01-01T08:44:45.08886Z","shell.execute_reply":"2024-01-01T08:44:45.455185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Verify Training Flag","metadata":{}},{"cell_type":"code","source":"def verify_correct_training_flag():\n    # Verify static output for inference\n    pred = model(X_batch_small, training=False)\n    for _ in tqdm(range(10)):\n        assert tf.reduce_min(tf.cast(pred == model(X_batch_small, training=False), tf.int8)) == 1\n\n    # Verify at least 99% varying output due to dropout during training\n    for _ in tqdm(range(10)):\n        assert tf.reduce_mean(tf.cast(pred != model(X_batch_small, training=True), tf.float32)) > 0.99\n        \nverify_correct_training_flag()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:45.457417Z","iopub.execute_input":"2024-01-01T08:44:45.458108Z","iopub.status.idle":"2024-01-01T08:44:51.188653Z","shell.execute_reply.started":"2024-01-01T08:44:45.458071Z","shell.execute_reply":"2024-01-01T08:44:51.187712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Verify No NaN Predictions","metadata":{}},{"cell_type":"code","source":"# Verify No NaN predictions\ndef verify_no_nan_predictions():\n    y_pred = model.predict(\n        val_dataset if USE_VAL else train_dataset,\n        steps=N_VAL_STEPS_PER_EPOCH if USE_VAL else 100,\n        verbose=VERBOSE,\n    )\n\n    print(f'# NaN Values In Predictions: {np.isnan(y_pred).sum()}')\n    \n    plt.figure(figsize=(15,8))\n    plt.title(f'Logit Predictions Initialized Model')\n    pd.Series(y_pred.flatten()).plot(kind='hist', bins=128)\n    plt.xlabel('Logits')\n    plt.grid()\n    plt.show()\n    \nverify_no_nan_predictions()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:44:51.190147Z","iopub.execute_input":"2024-01-01T08:44:51.19078Z","iopub.status.idle":"2024-01-01T08:45:00.00295Z","shell.execute_reply.started":"2024-01-01T08:44:51.190744Z","shell.execute_reply":"2024-01-01T08:45:00.002007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Learning Rate Scheduler","metadata":{}},{"cell_type":"code","source":"def lrfn(current_step, num_warmup_steps, lr_max, num_cycles=0.50, num_training_steps=N_EPOCHS):\n    \n    if current_step < num_warmup_steps:\n        if WARMUP_METHOD == 'log':\n            return lr_max * 0.10 ** (num_warmup_steps - current_step)\n        else:\n            return lr_max * 2 ** -(num_warmup_steps - current_step)\n    else:\n        progress = float(current_step - num_warmup_steps) / float(max(1, num_training_steps - num_warmup_steps))\n\n        return max(0.0, 0.5 * (1.0 + math.cos(math.pi * float(num_cycles) * 2.0 * progress))) * lr_max","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:45:00.004358Z","iopub.execute_input":"2024-01-01T08:45:00.00474Z","iopub.status.idle":"2024-01-01T08:45:00.012326Z","shell.execute_reply.started":"2024-01-01T08:45:00.004704Z","shell.execute_reply":"2024-01-01T08:45:00.011293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_lr_schedule(lr_schedule, epochs):\n    fig = plt.figure(figsize=(20, 10))\n    plt.plot([None] + lr_schedule + [None])\n    # X Labels\n    x = np.arange(1, epochs + 1)\n    x_axis_labels = [i if epochs <= 40 or i % 5 == 0 or i == 1 else None for i in range(1, epochs + 1)]\n    plt.xlim([1, epochs])\n    plt.xticks(x, x_axis_labels) # set tick step to 1 and let x axis start at 1\n    \n    # Increase y-limit for better readability\n    plt.ylim([0, max(lr_schedule) * 1.1])\n    \n    # Title\n    schedule_info = f'start: {lr_schedule[0]:.1E}, max: {max(lr_schedule):.1E}, final: {lr_schedule[-1]:.1E}'\n    plt.title(f'Step Learning Rate Schedule, {schedule_info}', size=18, pad=12)\n    \n    # Plot Learning Rates\n    for x, val in enumerate(lr_schedule):\n        if epochs <= 40 or x % 5 == 0 or x is epochs - 1:\n            if x < len(lr_schedule) - 1:\n                if lr_schedule[x - 1] < val:\n                    ha = 'right'\n                else:\n                    ha = 'left'\n            elif x == 0:\n                ha = 'right'\n            else:\n                ha = 'left'\n            plt.plot(x + 1, val, 'o', color='black');\n            offset_y = (max(lr_schedule) - min(lr_schedule)) * 0.02\n            plt.annotate(f'{val:.1E}', xy=(x + 1, val + offset_y), size=12, ha=ha)\n    \n    plt.xlabel('Epoch', size=16, labelpad=5)\n    plt.ylabel('Learning Rate', size=16, labelpad=5)\n    plt.grid()\n    plt.show()\n\n# Learning rate for encoder\nLR_SCHEDULE = [lrfn(step, num_warmup_steps=N_WARMUP_EPOCHS, lr_max=LR_MAX, num_cycles=0.50) for step in range(N_EPOCHS)]\n# Plot Learning Rate Schedule\nplot_lr_schedule(LR_SCHEDULE, epochs=N_EPOCHS)\n# Learning Rate Callback\nlr_callback = tf.keras.callbacks.LearningRateScheduler(lambda step: LR_SCHEDULE[step], verbose=0)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:45:00.013822Z","iopub.execute_input":"2024-01-01T08:45:00.01423Z","iopub.status.idle":"2024-01-01T08:45:00.829024Z","shell.execute_reply.started":"2024-01-01T08:45:00.014195Z","shell.execute_reply":"2024-01-01T08:45:00.828059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Weight Decay Callback","metadata":{}},{"cell_type":"code","source":"# Custom callback to update weight decay with learning rate\nclass WeightDecayCallback(tf.keras.callbacks.Callback):\n    def __init__(self, wd_ratio=WD_RATIO):\n        self.step_counter = 0\n        self.wd_ratio = wd_ratio\n    \n    def on_epoch_begin(self, epoch, logs=None):\n        model.optimizer.weight_decay = model.optimizer.learning_rate * self.wd_ratio\n        print(f'learning rate: {model.optimizer.learning_rate.numpy():.2e}, weight decay: {model.optimizer.weight_decay.numpy():.2e}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:45:00.830334Z","iopub.execute_input":"2024-01-01T08:45:00.830703Z","iopub.status.idle":"2024-01-01T08:45:00.837494Z","shell.execute_reply.started":"2024-01-01T08:45:00.830668Z","shell.execute_reply":"2024-01-01T08:45:00.836514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Evaluate Initialized Model","metadata":{}},{"cell_type":"code","source":"# Evaluate Initialized Model On Validation Data\ny_pred = model.evaluate(\n    val_dataset if USE_VAL else train_dataset,\n    steps=N_VAL_STEPS_PER_EPOCH if USE_VAL else TRAIN_STEPS_PER_EPOCH,\n    verbose=VERBOSE,\n)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:45:00.838816Z","iopub.execute_input":"2024-01-01T08:45:00.839207Z","iopub.status.idle":"2024-01-01T08:45:08.059335Z","shell.execute_reply.started":"2024-01-01T08:45:00.839174Z","shell.execute_reply":"2024-01-01T08:45:08.058251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Baseline","metadata":{}},{"cell_type":"code","source":"# baseline accuracy when only pad token is predicted\nif USE_VAL:\n    baseline_accuracy = np.mean(y_val == PAD_TOKEN)\nelse:\n    baseline_accuracy = np.mean(y_train == PAD_TOKEN)\nprint(f'Baseline Accuracy: {baseline_accuracy:.4f}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:45:08.060883Z","iopub.execute_input":"2024-01-01T08:45:08.0612Z","iopub.status.idle":"2024-01-01T08:45:08.067459Z","shell.execute_reply.started":"2024-01-01T08:45:08.061172Z","shell.execute_reply":"2024-01-01T08:45:08.066518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train","metadata":{}},{"cell_type":"code","source":"gc.collect()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:45:08.068703Z","iopub.execute_input":"2024-01-01T08:45:08.069008Z","iopub.status.idle":"2024-01-01T08:45:08.386822Z","shell.execute_reply.started":"2024-01-01T08:45:08.068982Z","shell.execute_reply":"2024-01-01T08:45:08.385889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## log parameters during training","metadata":{}},{"cell_type":"code","source":"from tensorflow.keras.callbacks import Callback\n\nclass CustomNeptuneCallback(Callback):\n    def __init__(self, run, num_epochs, batch_size, learning_rate, weight_decay_ratio, use_val_set, num_warmup_epochs):\n        self.run = run\n        self.num_epochs = num_epochs\n        self.batch_size = batch_size\n        self.learning_rate = learning_rate\n        self.weight_decay_ratio = weight_decay_ratio\n        self.use_val_set = use_val_set\n        self.num_warmup_epochs= num_warmup_epochs\n        self.logged_params = False\n\n    def on_epoch_begin(self, epoch, logs=None):\n        # Log parameters only once at the beginning of training\n        if not self.logged_params:\n            self.run[\"parameters/num_epochs\"] = self.params['epochs']  # Total number of epochs\n            self.run[\"parameters/batch_size\"] = self.batch_size\n            self.run[\"parameters/maximum_learning_rate\"] = self.learning_rate\n            self.run[\"parameters/weight_decay_ratio\"] = self.weight_decay_ratio\n            self.run[\"parameters/use_val_set\"] = self.use_val_set\n            self.run[\"parameters/num_warmup_epochs\"] = self.num_warmup_epochs\n            self.logged_params = True\n\n    def on_epoch_end(self, epoch, logs=None):\n        logs = logs or {}\n        self.run[\"training/epoch\"].log(epoch)\n        for key, value in logs.items():\n            self.run[f\"training/{key}\"].log(value)\n","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:45:08.388247Z","iopub.execute_input":"2024-01-01T08:45:08.388604Z","iopub.status.idle":"2024-01-01T08:45:08.399511Z","shell.execute_reply.started":"2024-01-01T08:45:08.388575Z","shell.execute_reply":"2024-01-01T08:45:08.398522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a callback that saves the model's weights\ncp_callback = tf.keras.callbacks.ModelCheckpoint(filepath='/kaggle/working/Results:/',\n                                                 save_best_only = True,\n                                                 verbose=1)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:45:08.400896Z","iopub.execute_input":"2024-01-01T08:45:08.401212Z","iopub.status.idle":"2024-01-01T08:45:08.413055Z","shell.execute_reply.started":"2024-01-01T08:45:08.401183Z","shell.execute_reply":"2024-01-01T08:45:08.412156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train Model","metadata":{}},{"cell_type":"code","source":"if TRAIN_MODEL:\n    # Clear all models in GPU\n    tf.keras.backend.clear_session()\n\n    # Get new fresh model\n    model = get_model()\n\n    # Sanity Check\n    model.summary()\n\n    # Actual Training\n    history = model.fit(\n            x=train_dataset,\n            steps_per_epoch=TRAIN_STEPS_PER_EPOCH,\n            epochs=N_EPOCHS,\n            # Only used for validation data since training data is a generator\n            validation_data=val_dataset if USE_VAL else None,\n            validation_steps=N_VAL_STEPS_PER_EPOCH if USE_VAL else None,\n            callbacks=[\n                lr_callback, #Learning rate \n                WeightDecayCallback(), #Weight decay\n                cp_callback, # checkpoints\n                CustomNeptuneCallback( #Log parameters in your Neptune project\n                    run=run,\n                    num_epochs=N_EPOCHS,\n                    batch_size=BATCH_SIZE,\n                    learning_rate=LR_MAX,\n                    weight_decay_ratio=WD_RATIO,\n                    use_val_set=USE_VAL,\n                    num_warmup_epochs= N_WARMUP_EPOCHS\n                )\n            ],\n            verbose=VERBOSE, #2\n        )","metadata":{"execution":{"iopub.status.busy":"2024-01-01T08:45:08.415937Z","iopub.execute_input":"2024-01-01T08:45:08.416659Z","iopub.status.idle":"2024-01-01T12:11:23.672331Z","shell.execute_reply.started":"2024-01-01T08:45:08.416631Z","shell.execute_reply":"2024-01-01T12:11:23.671336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save Model Weights\nmodel.save_weights('ASL.h5')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:11:23.674263Z","iopub.execute_input":"2024-01-01T12:11:23.674569Z","iopub.status.idle":"2024-01-01T12:11:23.807382Z","shell.execute_reply.started":"2024-01-01T12:11:23.674541Z","shell.execute_reply":"2024-01-01T12:11:23.806329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Log the saved model\nrun[\"artifacts/model\"].upload('ASL.h5')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:11:23.808941Z","iopub.execute_input":"2024-01-01T12:11:23.809649Z","iopub.status.idle":"2024-01-01T12:11:23.81485Z","shell.execute_reply.started":"2024-01-01T12:11:23.80961Z","shell.execute_reply":"2024-01-01T12:11:23.813862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Verify Model is Loaded Correctly\nmodel.evaluate(\n    val_dataset if USE_VAL else train_dataset,\n    steps=N_VAL_STEPS_PER_EPOCH if USE_VAL else TRAIN_STEPS_PER_EPOCH,\n    batch_size=BATCH_SIZE,\n    verbose=VERBOSE,\n)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:11:23.81626Z","iopub.execute_input":"2024-01-01T12:11:23.816913Z","iopub.status.idle":"2024-01-01T12:11:29.000025Z","shell.execute_reply.started":"2024-01-01T12:11:23.816879Z","shell.execute_reply":"2024-01-01T12:11:28.999084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Levenshtein Distance","metadata":{}},{"cell_type":"code","source":"# Output Predictions to string\ndef outputs2phrase(outputs):\n    if outputs.ndim == 2:\n        outputs = np.argmax(outputs, axis=1)\n    \n    return ''.join([ORD2CHAR.get(s, '') for s in outputs])","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:11:29.001242Z","iopub.execute_input":"2024-01-01T12:11:29.001561Z","iopub.status.idle":"2024-01-01T12:11:29.008045Z","shell.execute_reply.started":"2024-01-01T12:11:29.001533Z","shell.execute_reply":"2024-01-01T12:11:29.007178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"@tf.function(jit_compile=True)\ndef predict_phrase(frames):\n    # Add Batch Dimension\n    frames = tf.expand_dims(frames, axis=0)\n    # Start Phrase\n    phrase = tf.fill([1,MAX_PHRASE_LENGTH], PAD_TOKEN)\n\n    for idx in tf.range(MAX_PHRASE_LENGTH):\n        # Cast phrase to int8\n        phrase = tf.cast(phrase, tf.int8)\n        # Predict Next Token\n        outputs = model({\n            'frames': frames,\n            'phrase': phrase,\n        })\n\n        # Add predicted token to input phrase\n        phrase = tf.cast(phrase, tf.int32)\n        phrase = tf.where(\n            tf.range(MAX_PHRASE_LENGTH) < idx + 1,\n            tf.argmax(outputs, axis=2, output_type=tf.int32),\n            phrase,\n        )\n\n    # Squeeze outputs\n    outputs = tf.squeeze(phrase, axis=0)\n    outputs = tf.one_hot(outputs, N_UNIQUE_CHARACTERS)\n\n    # Return a dictionary with the output tensor\n    return outputs","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:11:29.009185Z","iopub.execute_input":"2024-01-01T12:11:29.009491Z","iopub.status.idle":"2024-01-01T12:11:29.02663Z","shell.execute_reply.started":"2024-01-01T12:11:29.009466Z","shell.execute_reply":"2024-01-01T12:11:29.0258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Levenstein Distance Train","metadata":{}},{"cell_type":"code","source":"# Compute Levenstein Distances\ndef get_ld_train():\n    N = 100 if IS_INTERACTIVE else 1000\n    LD_TRAIN = []\n    for idx, (frames, phrase_true) in enumerate(zip(tqdm(X_train, total=N), y_train)):\n        # Predict Phrase and Convert to String\n        phrase_pred = predict_phrase(frames).numpy()\n        phrase_pred = outputs2phrase(phrase_pred)\n        # True Phrase Ordinal to String\n        phrase_true = outputs2phrase(phrase_true)\n        # Add Levenstein Distance\n        LD_TRAIN.append({\n            'phrase_true': phrase_true,\n            'phrase_true_len': len(phrase_true),\n            'phrase_pred': phrase_pred,\n            'levenshtein_distance': levenshtein(phrase_pred, phrase_true),\n        })\n        # Take subset in interactive mode\n        if idx == N:\n            break\n            \n    # Convert to DataFrame\n    LD_TRAIN_DF = pd.DataFrame(LD_TRAIN)\n    \n    return LD_TRAIN_DF","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:11:29.027766Z","iopub.execute_input":"2024-01-01T12:11:29.028138Z","iopub.status.idle":"2024-01-01T12:11:29.040676Z","shell.execute_reply.started":"2024-01-01T12:11:29.028089Z","shell.execute_reply":"2024-01-01T12:11:29.039905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"LD_TRAIN_DF = get_ld_train()\n\n# Display Errors\ndisplay(LD_TRAIN_DF.head(30))","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:11:29.041928Z","iopub.execute_input":"2024-01-01T12:11:29.04229Z","iopub.status.idle":"2024-01-01T12:11:36.687432Z","shell.execute_reply.started":"2024-01-01T12:11:29.042255Z","shell.execute_reply":"2024-01-01T12:11:36.686375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Value Counts\nLD_TRAIN_VC = dict([(i, 0) for i in range(LD_TRAIN_DF['levenshtein_distance'].max()+1)])\nfor ld in LD_TRAIN_DF['levenshtein_distance']:\n    LD_TRAIN_VC[ld] += 1\n\n# Evaluation Metric\nN = LD_TRAIN_DF['phrase_true_len'].sum()\nD = LD_TRAIN_DF['levenshtein_distance'].sum()\nnld = (N - D) / N\n\nLD_TRAIN_VC = dict([(i, 0) for i in range(LD_TRAIN_DF['levenshtein_distance'].max()+1)])\nfor ld in LD_TRAIN_DF['levenshtein_distance']:\n    LD_TRAIN_VC[ld] += 1\n\nplt.figure(figsize=(15,8))\npd.Series(LD_TRAIN_VC).plot(kind='bar', width=1)\nplt.title(f'Train Levenstein Distance Distribution | Mean: {LD_TRAIN_DF.levenshtein_distance.mean():.4f}, NLD: {nld:.3f}')\nplt.xlabel('Levenstein Distance')\nplt.ylabel('Sample Count')\nplt.xlim(-0.50, LD_TRAIN_DF.levenshtein_distance.max()+0.50)\nplt.grid(axis='y')\nplt.savefig('temp.png')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:11:36.688818Z","iopub.execute_input":"2024-01-01T12:11:36.689118Z","iopub.status.idle":"2024-01-01T12:11:37.334796Z","shell.execute_reply.started":"2024-01-01T12:11:36.689091Z","shell.execute_reply":"2024-01-01T12:11:37.333835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# BLEU and Rouge Evaluation Metrics","metadata":{}},{"cell_type":"code","source":"Phrase_true_list = []\nPhrase_pred_list = []\nfor idx, (frames, phrase_true) in enumerate(zip(tqdm(X_val), y_val)):\n        # Predict Phrase and Convert to String\n        phrase_pred = predict_phrase(frames).numpy()\n        phrase_pred = outputs2phrase(phrase_pred)\n        Phrase_pred_list.append(phrase_pred)\n        # True Phrase Ordinal to String\n        phrase_true = outputs2phrase(phrase_true)\n        Phrase_true_list.append(phrase_true)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for ref, pred in zip(Phrase_true_list, Phrase_pred_list):\n    # BLEU scores\n    total_bleu_1 += sentence_bleu(ref, pred, weights=(1, 0, 0, 0))","metadata":{"execution":{"iopub.status.busy":"2024-01-01T14:00:22.006255Z","iopub.execute_input":"2024-01-01T14:00:22.006665Z","iopub.status.idle":"2024-01-01T14:00:22.093563Z","shell.execute_reply.started":"2024-01-01T14:00:22.006632Z","shell.execute_reply":"2024-01-01T14:00:22.09253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Phrase_true_=[]\nfor Phrase in Phrase_true_list:\n    Phrase_true_.append([Phrase])","metadata":{"execution":{"iopub.status.busy":"2024-01-01T14:14:49.941113Z","iopub.execute_input":"2024-01-01T14:14:49.941557Z","iopub.status.idle":"2024-01-01T14:14:49.94656Z","shell.execute_reply.started":"2024-01-01T14:14:49.941521Z","shell.execute_reply":"2024-01-01T14:14:49.94554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Phrase_pred_=[]\nfor Phrase in Phrase_pred_list:\n    Phrase_pred_.append([Phrase])","metadata":{"execution":{"iopub.status.busy":"2024-01-01T14:14:53.893944Z","iopub.execute_input":"2024-01-01T14:14:53.894679Z","iopub.status.idle":"2024-01-01T14:14:53.899011Z","shell.execute_reply.started":"2024-01-01T14:14:53.894642Z","shell.execute_reply":"2024-01-01T14:14:53.898065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for ref, pred in zip(Phrase_true_, Phrase_pred_):\n    # BLEU scores\n    total_bleu_1 += sentence_bleu(ref, pred, weights=(1, 0, 0, 0))\n    total_bleu_2 += sentence_bleu(ref, pred, weights=(0.5, 0.5, 0, 0))\n    total_bleu_3 += sentence_bleu(ref, pred, weights=(0.33, 0.33, 0.33, 0))\n    total_bleu_4 += sentence_bleu(ref, pred, weights=(0.25, 0.25, 0.25, 0.25))\n    \n    # ROUGE-L score\n    rouge = Rouge()\n    rouge_scores = rouge.get_scores(pred, ref)\n    total_rouge_l += rouge_scores[0]['rouge-l']['f']\n\n\n# Calculate the average of each metric\navg_bleu_1 = (total_bleu_1 / len(Phrase_true_list))*0.01\navg_bleu_2 = (total_bleu_2 / len(Phrase_true_list))*0.01\navg_bleu_3 = (total_bleu_3 / len(Phrase_true_list))*0.01\navg_bleu_4 = (total_bleu_4 / len(Phrase_true_list))*0.01\navg_rouge_l = (total_rouge_l / len(Phrase_true_list))*0.1\n\n# Log the metrics\nprint(f\"BLEU-1: {avg_bleu_1:.4f}, BLEU-2: {avg_bleu_2:.4f}, BLEU-3: {avg_bleu_3:.4f}, BLEU-4: {avg_bleu_4:.4f}\")\nprint(f\"METEOR: {avg_rouge_l:.4f}\")","metadata":{"execution":{"iopub.status.busy":"2024-01-01T14:16:33.309098Z","iopub.execute_input":"2024-01-01T14:16:33.309521Z","iopub.status.idle":"2024-01-01T14:16:33.365066Z","shell.execute_reply.started":"2024-01-01T14:16:33.30948Z","shell.execute_reply":"2024-01-01T14:16:33.364004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Levenstein Distance Evaluation","metadata":{}},{"cell_type":"code","source":"# Compute Levenstein Distances\ndef get_ld_val():\n    N = 100 if IS_INTERACTIVE else 1000\n    LD_VAL = []\n    for idx, (frames, phrase_true) in enumerate(zip(tqdm(X_val, total=N), y_val)):\n        # Predict Phrase and Convert to String\n        phrase_pred = predict_phrase(frames).numpy()\n        phrase_pred = outputs2phrase(phrase_pred)\n        # True Phrase Ordinal to String\n        phrase_true = outputs2phrase(phrase_true)\n        # Add Levenstein Distance\n        LD_VAL.append({\n            'phrase_true': phrase_true,\n            'phrase_true_len': len(phrase_true),\n            'phrase_pred': phrase_pred,\n            'levenshtein_distance': levenshtein(phrase_pred, phrase_true),\n        })\n        # Take subset in interactive mode\n        if idx == N:\n            break\n            \n    # Convert to DataFrame\n    LD_VAL_DF = pd.DataFrame(LD_VAL)\n    \n    return LD_VAL_DF","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:28:27.042665Z","iopub.execute_input":"2024-01-01T12:28:27.043787Z","iopub.status.idle":"2024-01-01T12:28:27.0508Z","shell.execute_reply.started":"2024-01-01T12:28:27.043746Z","shell.execute_reply":"2024-01-01T12:28:27.049865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if USE_VAL:\n    LD_VAL_DF = get_ld_val()\n\n    # Display Errors\n    display(LD_VAL_DF.head(30))","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:28:27.382882Z","iopub.execute_input":"2024-01-01T12:28:27.383272Z","iopub.status.idle":"2024-01-01T12:28:30.657786Z","shell.execute_reply.started":"2024-01-01T12:28:27.383239Z","shell.execute_reply":"2024-01-01T12:28:30.656848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Value Counts\nif USE_VAL:\n    # Evaluation Metric\n    N = LD_VAL_DF['phrase_true_len'].sum()\n    D = LD_VAL_DF['levenshtein_distance'].sum()\n    nld = (N - D) / N\n    \n    LD_VAL_VC = dict([(i, 0) for i in range(LD_VAL_DF['levenshtein_distance'].max()+1)])\n    for ld in LD_VAL_DF['levenshtein_distance']:\n        LD_VAL_VC[ld] += 1\n\n    plt.figure(figsize=(15,8))\n    pd.Series(LD_VAL_VC).plot(kind='bar', width=1)\n    plt.title(f'Validation Levenstein Distance Distribution | Mean: {LD_VAL_DF.levenshtein_distance.mean():.4f}, NLD: {nld:.3f}')\n    plt.xlabel('Levenstein Distance')\n    plt.ylabel('Sample Count')\n    plt.xlim(-0.50, LD_VAL_DF.levenshtein_distance.max()+0.50)\n    plt.grid(axis='y')\n    plt.savefig('temp.png')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:28:30.65947Z","iopub.execute_input":"2024-01-01T12:28:30.659777Z","iopub.status.idle":"2024-01-01T12:28:31.923939Z","shell.execute_reply.started":"2024-01-01T12:28:30.65975Z","shell.execute_reply":"2024-01-01T12:28:31.923036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training History","metadata":{}},{"cell_type":"code","source":"def plot_history_metric(metric, f_best=np.argmax, ylim=None, yscale=None, yticks=None):\n    # Only plot when training\n    if not TRAIN_MODEL:\n        return\n    \n    plt.figure(figsize=(20, 10))\n    \n    values = history.history[metric]\n    N_EPOCHS = len(values)\n    val = 'val' in ''.join(history.history.keys())\n    # Epoch Ticks\n    if N_EPOCHS <= 20:\n        x = np.arange(1, N_EPOCHS + 1)\n    else:\n        x = [1, 5] + [10 + 5 * idx for idx in range((N_EPOCHS - 10) // 5 + 1)]\n\n    x_ticks = np.arange(1, N_EPOCHS+1)\n\n    # Validation\n    if val:\n        val_values = history.history[f'val_{metric}']\n        val_argmin = f_best(val_values)\n        plt.plot(x_ticks, val_values, label=f'val')\n\n    # summarize history for accuracy\n    plt.plot(x_ticks, values, label=f'train')\n    argmin = f_best(values)\n    plt.scatter(argmin + 1, values[argmin], color='red', s=75, marker='o', label=f'train_best')\n    if val:\n        plt.scatter(val_argmin + 1, val_values[val_argmin], color='purple', s=75, marker='o', label=f'val_best')\n\n    plt.title(f'Model {metric}', fontsize=24, pad=10)\n    plt.ylabel(metric, fontsize=20, labelpad=10)\n\n    if ylim:\n        plt.ylim(ylim)\n\n    if yscale is not None:\n        plt.yscale(yscale)\n        \n    if yticks is not None:\n        plt.yticks(yticks, fontsize=16)\n\n    plt.xlabel('epoch', fontsize=20, labelpad=10)        \n    plt.tick_params(axis='x', labelsize=8)\n    plt.xticks(x, fontsize=16) # set tick step to 1 and let x axis start at 1\n    plt.yticks(fontsize=16)\n    \n    plt.legend(prop={'size': 10})\n    plt.grid()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:33:16.305001Z","iopub.execute_input":"2024-01-01T12:33:16.305956Z","iopub.status.idle":"2024-01-01T12:33:16.319259Z","shell.execute_reply.started":"2024-01-01T12:33:16.305918Z","shell.execute_reply":"2024-01-01T12:33:16.318179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# training loss history\nplot_history_metric('loss', f_best=np.argmin)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:33:16.788197Z","iopub.execute_input":"2024-01-01T12:33:16.789098Z","iopub.status.idle":"2024-01-01T12:33:17.333263Z","shell.execute_reply.started":"2024-01-01T12:33:16.789062Z","shell.execute_reply":"2024-01-01T12:33:17.332297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Top1accuracy history during training\nplot_history_metric('top1acc', ylim=[0,1], yticks=np.arange(0.0, 1.1, 0.1))","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:33:17.69185Z","iopub.execute_input":"2024-01-01T12:33:17.692241Z","iopub.status.idle":"2024-01-01T12:33:18.244022Z","shell.execute_reply.started":"2024-01-01T12:33:17.692207Z","shell.execute_reply":"2024-01-01T12:33:18.243091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Top5accuracy history during training\nplot_history_metric('top5acc', ylim=[0,1], yticks=np.arange(0.0, 1.1, 0.1))","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:33:20.464114Z","iopub.execute_input":"2024-01-01T12:33:20.465006Z","iopub.status.idle":"2024-01-01T12:33:21.010747Z","shell.execute_reply.started":"2024-01-01T12:33:20.464965Z","shell.execute_reply":"2024-01-01T12:33:21.009677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Inference","metadata":{}},{"cell_type":"code","source":"# Model Layer Names\nfor l in model.layers:\n    print(l.name)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T12:34:24.324406Z","iopub.execute_input":"2024-01-01T12:34:24.325422Z","iopub.status.idle":"2024-01-01T12:34:24.332369Z","shell.execute_reply.started":"2024-01-01T12:34:24.325387Z","shell.execute_reply":"2024-01-01T12:34:24.331334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# TFLite model For model and inference saving\nclass TFLiteModel(tf.Module):\n    def __init__(self, model):\n        super(TFLiteModel, self).__init__()\n\n        # Load the feature generation and main models\n        self.preprocess_layer = preprocess_layer\n        self.model = model\n    \n    @tf.function(jit_compile=True)\n    def encoder(self, x, frames_inp):\n        x = self.model.get_layer('embedding')(x)\n        x = self.model.get_layer('encoder')(x, frames_inp)\n        \n        return x\n        \n    @tf.function(jit_compile=True)\n    def decoder(self, x, phrase_inp, frames_inp):\n        x = self.model.get_layer('decoder')(x, phrase_inp, frames_inp)\n        x = self.model.get_layer('classifier')(x)\n        \n        return x\n    \n    @tf.function(input_signature=[tf.TensorSpec(shape=[None, N_COLS0], dtype=tf.float32, name='inputs')])\n    def __call__(self, inputs):\n        # Number Of Input Frames\n        N_INPUT_FRAMES = tf.shape(inputs)[0]\n        # Preprocess Data\n        frames_inp = self.preprocess_layer(inputs)        \n        # Add Batch Dimension\n        frames_inp = tf.expand_dims(frames_inp, axis=0)\n        # Get Encoding\n        encoding = self.encoder(frames_inp, frames_inp)\n        # Make Prediction\n        phrase = tf.fill([1,MAX_PHRASE_LENGTH], PAD_TOKEN)\n        # Predict One Token At A Time\n        stop = False\n        for idx in tf.range(MAX_PHRASE_LENGTH):\n            # Cast phrase to int8\n            phrase = tf.cast(phrase, tf.int8)\n            # If EOS token is predicted, stop predicting\n            outputs = tf.cond(\n                stop,\n                lambda: tf.one_hot(tf.cast(phrase, tf.int32), N_UNIQUE_CHARACTERS),\n                lambda: self.decoder(encoding, phrase, frames_inp)\n            )\n            # Add predicted token to input phrase\n            phrase = tf.cast(phrase, tf.int32)\n            # Replcae PAD token with predicted token up to idx\n            phrase = tf.where(\n                tf.range(MAX_PHRASE_LENGTH) < idx + 1,\n                tf.argmax(outputs, axis=2, output_type=tf.int32),\n                phrase,\n            )\n            # Predicted Token\n            predicted_token = phrase[0,idx]\n            # If EOS (End Of Sentence) token is predicted stop\n            if not stop:\n                stop = predicted_token == EOS_TOKEN\n            \n        # Squeeze outputs\n        outputs = tf.squeeze(phrase, axis=0)\n        outputs = tf.one_hot(outputs, N_UNIQUE_CHARACTERS)\n            \n        # Return a dictionary with the output tensor\n        return {'outputs': outputs }\n\n# Define TF Lite Model\ntflite_keras_model = TFLiteModel(model)\n\n# Sanity Check\ndemo_sequence_id = example_parquet_df.index.unique()[900]\ndemo_raw_data = example_parquet_df.loc[demo_sequence_id, COLUMNS0].values\ndemo_phrase_true = train_sequence_id.loc[demo_sequence_id, 'phrase']\nprint(f'demo_raw_data shape: {demo_raw_data.shape}, dtype: {demo_raw_data.dtype}')\ndemo_output = tflite_keras_model(demo_raw_data)['outputs'].numpy()\nprint(f'demo_output shape: {demo_output.shape}, dtype: {demo_output.dtype}')\nprint(f'demo_outputs phrase decoded: {outputs2phrase(demo_output)}')\nprint(f'phrase true: {demo_phrase_true}')","metadata":{"execution":{"iopub.status.busy":"2024-01-01T13:03:14.625318Z","iopub.execute_input":"2024-01-01T13:03:14.625691Z","iopub.status.idle":"2024-01-01T13:03:17.433253Z","shell.execute_reply.started":"2024-01-01T13:03:14.625661Z","shell.execute_reply":"2024-01-01T13:03:17.432181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create Model Converter\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)\n# Convert Model\ntflite_model = keras_model_converter.convert()\n# Write Model\nwith open('/kaggle/working/model.tflite', 'wb') as f:\n    f.write(tflite_model)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T13:03:31.987972Z","iopub.execute_input":"2024-01-01T13:03:31.988884Z","iopub.status.idle":"2024-01-01T13:04:20.579802Z","shell.execute_reply.started":"2024-01-01T13:03:31.988847Z","shell.execute_reply":"2024-01-01T13:04:20.578766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add selected_columns json to only select specific columns from input frames\nwith open('inference_args.json', 'w') as f:\n     json.dump({ 'selected_columns': COLUMNS0.tolist() }, f)","metadata":{"execution":{"iopub.status.busy":"2024-01-01T13:04:27.536001Z","iopub.execute_input":"2024-01-01T13:04:27.536937Z","iopub.status.idle":"2024-01-01T13:04:27.542585Z","shell.execute_reply.started":"2024-01-01T13:04:27.5369Z","shell.execute_reply":"2024-01-01T13:04:27.541568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Zip Model\n!zip submission.zip /kaggle/working/model.tflite /kaggle/working/inference_args.json","metadata":{"execution":{"iopub.status.busy":"2024-01-01T13:04:28.776362Z","iopub.execute_input":"2024-01-01T13:04:28.776762Z","iopub.status.idle":"2024-01-01T13:04:31.453258Z","shell.execute_reply.started":"2024-01-01T13:04:28.776725Z","shell.execute_reply":"2024-01-01T13:04:31.451981Z"},"trusted":true},"execution_count":null,"outputs":[]}]}