{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\n\nfrom tqdm.notebook import tqdm\nfrom sklearn.model_selection import train_test_split, GroupShuffleSplit\n\n# TQDM Progress Bar With Pandas Apply Function\ntqdm.pandas()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import json\nimport os\nimport matplotlib as mpl\nimport glob\nimport math","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code reads a JSON file containing a character-to-ordinal encoding mapping, creates an ordinal-to-character mapping, and displays the character-to-ordinal encoding mapping using a Pandas DataFrame. Here's an explanation of the code:\n\n1. Reading JSON file: The code uses the `open()` function to open the JSON file `character_to_prediction_index.json` located at `/kaggle/input/asl-fingerspelling/`. The file contains the character-to-ordinal encoding mapping. The `json.load()` function is used to load the contents of the JSON file into a Python dictionary called `CHAR2ORD`.\n\n2. Creating ordinal-to-character mapping: The code creates a new dictionary called `ORD2CHAR` using a dictionary comprehension. It swaps the keys and values of `CHAR2ORD`, so the ordinal values become keys, and the corresponding characters become values.\n\n3. Displaying the character-to-ordinal encoding mapping: The code creates a Pandas Series from `CHAR2ORD` using `pd.Series(CHAR2ORD)`. It then converts the Series to a DataFrame using the `to_frame()` method and assigns the column name as 'Ordinal Encoding'. Finally, the `display()` function is used to display the DataFrame.","metadata":{}},{"cell_type":"code","source":"# Read Character to Ordinal Encoding Mapping\nwith open('/kaggle/input/asl-fingerspelling/character_to_prediction_index.json') as json_file:\n    CHAR2ORD = json.load(json_file)\n    \n# Ordinal to Character Mapping\nORD2CHAR = {j:i for i,j in CHAR2ORD.items()}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code sets several configuration variables for model training and data preprocessing. Here's an explanation of each variable:\n\n1. `IS_INTERACTIVE`: It checks if the notebook is running in interactive mode or by committing. It uses the `KAGGLE_KERNEL_RUN_TYPE` environment variable to determine this.\n\n2. `VERBOSE`: It sets the verbosity level for training. If the notebook is running in interactive mode (`IS_INTERACTIVE` is `True`), the verbosity is set to 1; otherwise, it is set to 2.\n\n3. `SEED`: It sets the global random seed to a fixed value (42 in this case) to ensure reproducibility.\n\n4. `N_TARGET_FRAMES`: It specifies the number of frames to resize the recording to. The value is set to 128.\n\n5. `DEBUG`: It is a global debug flag that determines whether to use a subset of the train data for debugging purposes. By default, it is set to `False`.\n\n6. `N_UNIQUE_CHARACTERS0` and `N_UNIQUE_CHARACTERS`: They define the number of unique characters to predict in the model. `N_UNIQUE_CHARACTERS0` is set to the length of the `CHAR2ORD` dictionary, and `N_UNIQUE_CHARACTERS` is calculated by adding 1 for the PAD token, 1 for the SOS token, and 1 for the EOS token.\n\n7. `PAD_TOKEN`, `SOS_TOKEN`, and `EOS_TOKEN`: They define the token values for padding, start of sentence, and end of sentence, respectively. These values are based on the length of the `CHAR2ORD` dictionary.\n\n8. `USE_VAL`: It determines whether to use 10% of the data for validation during training. By default, it is set to `True`.\n\n9. `BATCH_SIZE`: It sets the batch size for training. The value is set to 96.\n\n10. `N_EPOCHS`: It determines the number of epochs to train for. If the notebook is running in interactive mode (`IS_INTERACTIVE` is `True`), the value is set to 2; otherwise, it is set to 55.\n\n11. `N_WARMUP_EPOCHS`: It specifies the number of warm-up epochs in the learning rate scheduler. The value is set to 0.\n\n12. `LR_MAX`: It sets the maximum learning rate for the model. The value is set to 1e-3.\n\n13. `WD_RATIO`: It determines the weight decay ratio as a ratio of the learning rate. The value is set to 0.05.\n\n14. `MAX_PHRASE_LENGTH`: It defines the maximum length of the phrase plus the EOS token. The value is set to 31 + 1.\n\n15. `TRAIN_MODEL`: It specifies whether to train the model. By default, it is set to `True`.\n\n16. `LOAD_WEIGHTS`: It determines whether to load pretrained weights for the model. By default, it is set to `False`.","metadata":{}},{"cell_type":"code","source":"IS_INTERACTIVE = os.environ['KAGGLE_KERNEL_RUN_TYPE'] == 'Interactive'\nVERBOSE = 1 if IS_INTERACTIVE else 2\nSEED = 42\nN_TARGET_FRAMES = 128\nDEBUG = False\nN_UNIQUE_CHARACTERS0 = len(CHAR2ORD)\nN_UNIQUE_CHARACTERS = len(CHAR2ORD) + 1 + 1 + 1\nPAD_TOKEN = len(CHAR2ORD) # Padding\nSOS_TOKEN = len(CHAR2ORD) + 1 # Start Of Sentence\nEOS_TOKEN = len(CHAR2ORD) + 2 # End Of Sentence\nUSE_VAL = False\nBATCH_SIZE = 96\nN_EPOCHS = 2 if IS_INTERACTIVE else 55\nN_WARMUP_EPOCHS = 0\nLR_MAX = 1e-3\nWD_RATIO = 0.05\nMAX_PHRASE_LENGTH = 31 + 1\nTRAIN_MODEL = True\nLOAD_WEIGHTS = False","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code snippet updates the global settings of Matplotlib to customize the appearance of plots. Here's an explanation of each line:\n\n1. `mpl.rcParams.update(mpl.rcParamsDefault)`: This line resets the global Matplotlib settings to their default values before applying customizations.\n\n2. `mpl.rcParams['xtick.labelsize'] = 16`: It sets the font size of x-axis tick labels to 16.\n\n3. `mpl.rcParams['ytick.labelsize'] = 16`: It sets the font size of y-axis tick labels to 16.\n\n4. `mpl.rcParams['axes.labelsize'] = 18`: It sets the font size of axis labels to 18.\n\n5. `mpl.rcParams['axes.titlesize'] = 24`: It sets the font size of plot titles to 24.\n\nThese settings affect the entire Matplotlib environment and will be applied to all subsequent plots generated in the code.","metadata":{}},{"cell_type":"code","source":"mpl.rcParams.update(mpl.rcParamsDefault)\nmpl.rcParams['xtick.labelsize'] = 16\nmpl.rcParams['ytick.labelsize'] = 16\nmpl.rcParams['axes.labelsize'] = 18\nmpl.rcParams['axes.titlesize'] = 24","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code reads a CSV file containing training data and performs some operations on the DataFrame. Here's an explanation of each part of the code:\n\n1. Reading the train DataFrame:\n   - If `DEBUG` is `True`, it reads the first 5000 rows from the CSV file located at '/kaggle/input/asl-fingerspelling/train.csv' using `pd.read_csv()`.\n   - If `DEBUG` is `False`, it reads the entire CSV file.\n\n2. Setting train DataFrame indexed by 'sequence_id':\n   The code sets the index of the `train` DataFrame to the 'sequence_id' column using the `set_index()` function and assigns the result to `train_sequence_id`.\n\n3. Getting the number of train samples:\n   The code calculates the number of samples in the train DataFrame by getting the length of the DataFrame and assigns it to the variable `N_SAMPLES`. It then prints the value of `N_SAMPLES`.\n\n4. Displaying information about the train DataFrame:\n   The code uses `train.info()` to display information about the train DataFrame, such as the number of rows, column names, and data types. It uses `train.head()` to display the first few rows of the DataFrame.\n\n5. Defining a function to get the complete file path:\n   The code defines a function `get_file_path()` that takes a path as input and returns the complete file path by appending it to '/kaggle/input/asl-fingerspelling/'.\n\n6. Adding a 'file_path' column to the train DataFrame:\n   The code applies the `get_file_path()` function to the 'path' column of the `train` DataFrame using the `apply()` function and assigns the result to the 'file_path' column.\n\n7. Getting the unique inference pickle files:\n   The code uses the `glob.glob()` function to get a list of file paths that match the pattern '/kaggle/input/aslfr-preprocessing-dataset/train_landmark_subsets/*' and assigns it to the `INFERENCE_FILE_PATHS` variable. It prints the number of inference pickle files found.","metadata":{}},{"cell_type":"code","source":"# Read Train DataFrame\nif DEBUG:\n    train = pd.read_csv('/kaggle/input/asl-fingerspelling/train.csv').head(5000)\nelse:\n    train = pd.read_csv('/kaggle/input/asl-fingerspelling/train.csv')\n    \n# Set Train Indexed By sqeuence_id\ntrain_sequence_id = train.set_index('sequence_id')\n\n# Number Of Train Samples\nN_SAMPLES = len(train)\n\ndisplay(train.info())\ndisplay(train.head())\n\n# Get complete file path to file\ndef get_file_path(path):\n    return f'/kaggle/input/asl-fingerspelling/{path}'\ntrain['file_path'] = train['path'].apply(get_file_path)\n\n# Unique Parquet Files\nINFERENCE_FILE_PATHS = pd.Series(\n        glob.glob('/kaggle/input/aslfr-preprocessing-dataset/train_landmark_subsets/*')\n    )\nprint(f'Found {len(INFERENCE_FILE_PATHS)} Inference Pickle Files')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code snippet handles the loading and preprocessing of training and validation data for the model. Here's an explanation of each part of the code:\n\n1. Using validation data (`USE_VAL` is `True`):\n   - It loads the training data and labels from the numpy files: `X_train.npy` and `y_train.npy`. `X_train` contains the training data, and `y_train` contains the corresponding labels.\n   - It retrieves the number of training samples from `X_train` and assigns it to the variable `N_TRAIN_SAMPLES`.\n   - It loads the validation data and labels from the numpy files: `X_val.npy` and `y_val.npy`. `X_val` contains the validation data, and `y_val` contains the corresponding labels.\n   - It retrieves the number of validation samples from `X_val` and assigns it to the variable `N_VAL_SAMPLES`.\n   - It prints the shapes of `X_train` and `X_val`.\n\n2. Using all the data for training (`USE_VAL` is `False`):\n   - It loads the entire data and labels from the numpy files: `X.npy` and `y.npy`. `X_train` contains the training data, and `y_train` contains the corresponding labels.\n   - It retrieves the number of training samples from `X_train` and assigns it to the variable `N_TRAIN_SAMPLES`.\n   - It prints the shape of `X_train`.\n\n3. Loading mean and standard deviations:\n   - It loads the mean and standard deviations used for normalizing the data from the numpy files: `MEANS.npy` and `STDS.npy`. `MEANS` contains the mean values, and `STDS` contains the standard deviations.\n   - The mean and standard deviations are reshaped to a one-dimensional array.\n\n4. Creating an example batch for debugging:\n   - It selects a batch of examples for debugging purposes. The batch size is determined by the variable `N_EXAMPLE_BATCH_SAMPLES`.\n   - The batch data is stored in the `X_batch` dictionary, with 'frames' representing the input frames and 'phrase' representing the corresponding phrases.\n   - The batch labels are stored in the `y_batch` array.\n\n5. Reading the first parquet file:\n   - It reads the first parquet file from the `INFERENCE_FILE_PATHS` series, which contains the file paths of inference pickle files.\n   - It assigns the resulting DataFrame to `example_parquet_df`.","metadata":{}},{"cell_type":"code","source":"# Train/Validation\nif USE_VAL:\n    # TRAIN\n    X_train = np.load('/kaggle/input/aslfr-preprocessing-dataset/X_train.npy')\n    y_train = np.load('/kaggle/input/aslfr-preprocessing-dataset/y_train.npy')[:,:MAX_PHRASE_LENGTH]\n    N_TRAIN_SAMPLES = len(X_train)\n    # VAL\n    X_val = np.load('/kaggle/input/aslfr-preprocessing-dataset/X_val.npy')\n    y_val = np.load('/kaggle/input/aslfr-preprocessing-dataset/y_val.npy')[:,:MAX_PHRASE_LENGTH]\n    N_VAL_SAMPLES = len(X_val)\n    # Shapes\n    print(f'X_train shape: {X_train.shape}, X_val shape: {X_val.shape}')\n# Train On All Data\nelse:\n    # TRAIN\n    X_train = np.load('/kaggle/input/aslfr-preprocessing-dataset/X.npy')\n    y_train = np.load('/kaggle/input/aslfr-preprocessing-dataset/y.npy')[:,:MAX_PHRASE_LENGTH]\n    N_TRAIN_SAMPLES = len(X_train)\n    print(f'X_train shape: {X_train.shape}')\n    \n# Mean/Standard Deviations of data used for normalizing\nMEANS = np.load('/kaggle/input/aslfr-preprocessing-dataset/MEANS.npy').reshape(-1)\nSTDS = np.load('/kaggle/input/aslfr-preprocessing-dataset/STDS.npy').reshape(-1)\n    \n# Example Batch For Debugging\nN_EXAMPLE_BATCH_SAMPLES = 1024\n\nX_batch = {\n    'frames': np.copy(X_train[:N_EXAMPLE_BATCH_SAMPLES]),\n    'phrase': np.copy(y_train[:N_EXAMPLE_BATCH_SAMPLES]),\n}\ny_batch = np.copy(y_train[:N_EXAMPLE_BATCH_SAMPLES])\n\n# Read First Parquet File\nexample_parquet_df = pd.read_parquet(INFERENCE_FILE_PATHS[0])\n\n# Each parquet file contains 1000 recordings\nprint(f'# Unique Recording: {example_parquet_df.index.nunique()}')\n# Display DataFrame layout\ndisplay(example_parquet_df.head())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code defines a function called `get_idxs` that extracts indices from a DataFrame based on certain conditions. Here's an explanation of the function:\n\n1. Function signature:\n   - `get_idxs(df, words_pos, words_neg=[], ret_names=True, idxs_pos=None)`\n\n   - `df`: The DataFrame from which indices need to be extracted.\n   - `words_pos`: A list of positive words. The function will look for column names that contain these words.\n   - `words_neg`: An optional list of negative words. The function will exclude column names that contain any of these words.\n   - `ret_names`: A boolean indicating whether to return both the column indices and names (`True`) or only the column indices (`False`).\n   - `idxs_pos`: An optional list of positive indices. The function will consider only columns with these indices. If `None`, all column indices will be considered.\n\n2. Iterating over column names:\n   The function iterates over the column names of the DataFrame `example_parquet_df`.\n\n3. Excluding non-landmark columns:\n   The function skips the iteration if the column name is 'frame' to exclude non-landmark columns.\n\n4. Extracting indices based on conditions:\n   - The function extracts the column index from the column name by splitting it and taking the last element as an integer (`col_idx = int(col.split('_')[-1])`).\n   - It checks if the column name contains all positive words (`w in col`).\n   - It checks if the column index is in the list of positive indices (`idxs_pos is None or col_idx in idxs_pos`).\n   - It checks if none of the negative words are present in the column name (`all([w not in col for w in words_neg])`).\n   - If the above conditions are met, it appends the column index to the `idxs` list and the column name to the `names` list.\n\n5. Converting to numpy arrays and returning the results:\n   - The function converts the `idxs` and `names` lists to numpy arrays (`np.array(idxs)` and `np.array(names)`).\n   - If `ret_names` is `True`, it returns both the column indices and names as a tuple (`return idxs, names`).\n   - If `ret_names` is `False`, it returns only the column indices (`return idxs`).","metadata":{}},{"cell_type":"code","source":"# Get indices in original dataframe\ndef get_idxs(df, words_pos, words_neg=[], ret_names=True, idxs_pos=None):\n    idxs = []\n    names = []\n    for w in words_pos:\n        for col_idx, col in enumerate(example_parquet_df.columns):\n            # Exclude Non Landmark Columns\n            if col in ['frame']:\n                continue\n                \n            col_idx = int(col.split('_')[-1])\n            # Check if column name contains all words\n            if (w in col) and (idxs_pos is None or col_idx in idxs_pos) and all([w not in col for w in words_neg]):\n                idxs.append(col_idx)\n                names.append(col)\n    # Convert to Numpy arrays\n    idxs = np.array(idxs)\n    names = np.array(names)\n    # Returns either both column indices and names\n    if ret_names:\n        return idxs, names\n    # Or only columns indices\n    else:\n        return idxs","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code snippet defines and initializes several variables related to landmark indices and names. Here's an explanation of each part of the code:\n\n1. `LIPS_LANDMARK_IDXS`:\n   - It is a numpy array containing a list of indices representing the landmark positions for the lips.\n   - The indices are used to select specific columns related to the lips in the DataFrame.\n\n2. `LEFT_HAND_IDXS0` and `LEFT_HAND_NAMES0`:\n   - These variables are initialized by calling the `get_idxs()` function to extract column indices and names related to the left hand.\n   - The positive word used is 'left_hand', and the negative word is 'z' to exclude columns related to the z-axis.\n   - The `idxs_pos` argument is not specified, so all column indices are considered.\n\n3. `RIGHT_HAND_IDXS0` and `RIGHT_HAND_NAMES0`:\n   - These variables are initialized similar to the left hand variables, but for the right hand.\n   - The positive word used is 'right_hand', and the negative word is 'z' to exclude columns related to the z-axis.\n\n4. `LIPS_IDXS0` and `LIPS_NAMES0`:\n   - These variables are initialized similar to the hand variables, but for the lips.\n   - The positive word used is 'face', and the negative word is 'z'.\n   - The `idxs_pos` argument is set to `LIPS_LANDMARK_IDXS`, which contains specific indices for the lips landmarks.\n\n5. `COLUMNS0`:\n   - It is a numpy array that concatenates the column names related to the left hand, right hand, and lips.\n   - It represents all the selected columns used for modeling.\n\n6. `N_COLS0`:\n   - It stores the number of columns in `COLUMNS0` using the `len()` function.\n   - It represents the total number of columns used for modeling.\n\n7. `N_DIMS0`:\n   - It is set to 2 and represents the number of dimensions used for modeling (only the X and Y axes).","metadata":{}},{"cell_type":"code","source":"# Lips Landmark Face Ids\nLIPS_LANDMARK_IDXS = np.array([\n        61, 185, 40, 39, 37, 0, 267, 269, 270, 409,\n        291, 146, 91, 181, 84, 17, 314, 405, 321, 375,\n        78, 191, 80, 81, 82, 13, 312, 311, 310, 415,\n        95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n    ])\n\n# Landmark Indices for Left/Right hand without z axis in raw data\nLEFT_HAND_IDXS0, LEFT_HAND_NAMES0 = get_idxs(example_parquet_df, ['left_hand'], ['z'])\nRIGHT_HAND_IDXS0, RIGHT_HAND_NAMES0 = get_idxs(example_parquet_df, ['right_hand'], ['z'])\nLIPS_IDXS0, LIPS_NAMES0 = get_idxs(example_parquet_df, ['face'], ['z'], idxs_pos=LIPS_LANDMARK_IDXS)\nCOLUMNS0 = np.concatenate((LEFT_HAND_NAMES0, RIGHT_HAND_NAMES0, LIPS_NAMES0))\nN_COLS0 = len(COLUMNS0)\n# Only X/Y axes are used\nN_DIMS0 = 2\n\nprint(f'N_COLS0: {N_COLS0}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code snippet extracts the indices of selected landmark columns from a subset of the DataFrame. Here's an explanation of each part of the code:\n\n1. `LEFT_HAND_IDXS`, `RIGHT_HAND_IDXS`, and `LIPS_IDXS`:\n   - These variables use the `np.argwhere()` function to find the indices where the column names from `COLUMNS0` match the respective hand or lips names.\n   - The `np.isin()` function is used to check if the column names in `COLUMNS0` are present in the corresponding hand or lips names.\n   - `np.argwhere()` returns an array of indices for the matching column names, and `squeeze()` is used to remove any unnecessary dimensions.\n   - These variables represent the indices of the selected columns related to the left hand, right hand, and lips.\n\n2. `HAND_IDXS`:\n   - This variable concatenates the indices of the selected columns related to the left hand and right hand using `np.concatenate()`.\n   - It represents the combined indices of the hand columns.\n\n3. `N_COLS`:\n   - It is assigned the value of `N_COLS0`, which represents the total number of selected columns.\n\n4. `N_DIMS`:\n   - It is set to 2, representing the number of dimensions used for modeling (only the X and Y axes).","metadata":{}},{"cell_type":"code","source":"# Landmark Indices in subset of dataframe with only COLUMNS selected\nLEFT_HAND_IDXS = np.argwhere(np.isin(COLUMNS0, LEFT_HAND_NAMES0)).squeeze()\nRIGHT_HAND_IDXS = np.argwhere(np.isin(COLUMNS0, RIGHT_HAND_NAMES0)).squeeze()\nLIPS_IDXS = np.argwhere(np.isin(COLUMNS0, LIPS_NAMES0)).squeeze()\nHAND_IDXS = np.concatenate((LEFT_HAND_IDXS, RIGHT_HAND_IDXS), axis=0)\nN_COLS = N_COLS0\n# Only X/Y axes are used\nN_DIMS = 2\n\nprint(f'N_COLS: {N_COLS}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code snippet extracts the indices and names of the processed data columns based on the axes (X and Y) for the dominant hand (left hand). Here's an explanation of each part of the code:\n\n1. `HAND_X_IDXS`:\n   - It uses a list comprehension to iterate over the indices and names of the columns related to the left hand (`LEFT_HAND_NAMES0`).\n   - For each column, it checks if 'x' is present in the name, and if so, it appends the index to the list.\n   - The resulting list is converted to a numpy array using `np.array()` and assigned to `HAND_X_IDXS`.\n   - `squeeze()` is then used to remove any unnecessary dimensions.\n\n2. `HAND_Y_IDXS`:\n   - It is similar to `HAND_X_IDXS`, but it checks if 'y' is present in the name instead.\n   - It extracts the indices of the columns related to the Y axis of the left hand.\n\n3. `HAND_X_NAMES` and `HAND_Y_NAMES`:\n   - These variables store the names of the columns in `LEFT_HAND_NAMES0` corresponding to the X and Y axes, respectively.\n   - The names are extracted using the indices in `HAND_X_IDXS` and `HAND_Y_IDXS`.","metadata":{}},{"cell_type":"code","source":"# Indices in processed data by axes with only dominant hand\nHAND_X_IDXS = np.array(\n        [idx for idx, name in enumerate(LEFT_HAND_NAMES0) if 'x' in name]\n    ).squeeze()\nHAND_Y_IDXS = np.array(\n        [idx for idx, name in enumerate(LEFT_HAND_NAMES0) if 'y' in name]\n    ).squeeze()\n# Names in processed data by axes\nHAND_X_NAMES = LEFT_HAND_NAMES0[HAND_X_IDXS]\nHAND_Y_NAMES = LEFT_HAND_NAMES0[HAND_Y_IDXS]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code snippet defines a TensorFlow layer called `PreprocessLayer` that performs data preprocessing in TFLite. Here's an explanation of each part of the code:\n\n1. `PreprocessLayer` class:\n   - It is a custom layer that subclasses `tf.keras.layers.Layer`.\n   - The `__init__` method initializes the layer and defines a constant `normalisation_correction` as a TensorFlow constant.\n   - The `call` method is decorated with `@tf.function` and specifies the input signature using `input_signature`. It takes `data0` as input, which is a tensor of shape `[None, N_COLS0]`.\n   - The `call` method performs various preprocessing steps on the input data and returns the processed data.\n\n2. `normalisation_correction`:\n   - It is a TensorFlow constant that is initialized based on the `LEFT_HAND_NAMES0` list.\n   - It contains values to correct the normalization of the X coordinates of the left hand (original right hand) and the Y coordinates of the right hand (original left hand).\n\n3. `call` method:\n   - It takes the input data `data0` and a `resize` flag as input.\n   - It first replaces any NaN values in the input data with 0 using `tf.where(tf.math.is_nan(data0), 0.0, data0)`.\n   - It applies a hacky operation by adding a batch dimension of size 1 using `data[None]`.\n   - It performs empty hand frame filtering by calculating the absolute values of the hands, summing them along the third axis, and creating a mask that checks if the sum is not equal to 0. The data is then filtered based on this mask.\n   - If the number of frames is less than `N_TARGET_FRAMES`, it pads the data with zeros to match the desired number of frames.\n   - It downsamples the data by resizing it to have `N_TARGET_FRAMES` frames using `tf.image.resize` with the `BILINEAR` method.\n   - It removes the batch dimension using `tf.squeeze`.\n   - The processed data is then returned.\n\n4. `preprocess_layer`:\n   - It instantiates an object of the `PreprocessLayer` class.\n\nThe code snippet defines a custom TensorFlow layer that performs data preprocessing in TFLite, taking care of handling NaN values, empty hand frame filtering, padding, downsampling, and removing unnecessary dimensions. The `preprocess_layer` object can be used as part of a TensorFlow model for further processing of the data.","metadata":{}},{"cell_type":"code","source":"class PreprocessLayer(tf.keras.layers.Layer):\n    def __init__(self):\n        super(PreprocessLayer, self).__init__()\n        self.normalisation_correction = tf.constant(\n                    # Add 0.50 to x coordinates of left hand (original right hand) and substract 0.50 of right hand (original left hand)\n                     [0.50 if 'x' in name else 0.00 for name in LEFT_HAND_NAMES0],\n                dtype=tf.float32,\n            )\n    \n    @tf.function(\n        input_signature=(tf.TensorSpec(shape=[None,N_COLS0], dtype=tf.float32),),\n    )\n    def call(self, data0, resize=True):\n        # Fill NaN Values With 0\n        data = tf.where(tf.math.is_nan(data0), 0.0, data0)\n        \n        # Hacky\n        data = data[None]\n        \n        # Empty Hand Frame Filtering\n        hands = tf.slice(data, [0,0,0], [-1, -1, 84])\n        hands = tf.abs(hands)\n        mask = tf.reduce_sum(hands, axis=2)\n        mask = tf.not_equal(mask, 0)\n        data = data[mask][None]\n        \n        # Pad Zeros\n        N_FRAMES = len(data[0])\n        if N_FRAMES < N_TARGET_FRAMES:\n            data = tf.concat((\n                data,\n                tf.zeros([1,N_TARGET_FRAMES-N_FRAMES,N_COLS], dtype=tf.float32)\n            ), axis=1)\n        # Downsample\n        data = tf.image.resize(\n            data,\n            [1, N_TARGET_FRAMES],\n            method=tf.image.ResizeMethod.BILINEAR,\n        )\n        \n        # Squeeze Batch Dimension\n        data = tf.squeeze(data, axis=[0])\n        \n        return data\n    \npreprocess_layer = PreprocessLayer()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code defines a function called `test_preprocess_layer()` to test the `PreprocessLayer` layer. Here's an explanation of the function:\n\n1. `demo_sequence_id`:\n   - It selects a demonstration sequence ID from the unique indices of the `example_parquet_df` DataFrame.\n\n2. `demo_raw_data`:\n   - It retrieves the raw data for the selected demonstration sequence ID from the `example_parquet_df` DataFrame using `loc`.\n   - The columns used for the demonstration are based on the `COLUMNS0` array.\n\n3. `data`:\n   - It applies the `preprocess_layer` to the `demo_raw_data` to preprocess the data.\n\n4. Printing the shape of the data:\n   - It prints the shape of the `demo_raw_data` and the processed `data` to compare their shapes.\n\n5. Returning the processed data:\n   - The processed `data` is returned from the function.\n\n6. Testing the preprocess layer:\n   - If the notebook is running in interactive mode (`IS_INTERACTIVE` is `True`), the `test_preprocess_layer()` function is called, and the processed data is stored in the `data` variable.\n\nThe purpose of the `test_preprocess_layer()` function is to demonstrate the usage of the `preprocess_layer` by preprocessing a sample raw data sequence and comparing the shapes of the raw data and the processed data. If the notebook is running in interactive mode, the function is executed, and the processed data is stored in the `data` variable for further examination.","metadata":{}},{"cell_type":"code","source":"# Function To Test Preprocessing Layer\ndef test_preprocess_layer():\n    demo_sequence_id = example_parquet_df.index.unique()[15]\n    demo_raw_data = example_parquet_df.loc[demo_sequence_id, COLUMNS0]\n    data = preprocess_layer(demo_raw_data)\n\n    print(f'demo_raw_data shape: {demo_raw_data.shape}')\n    print(f'data shape: {data.shape}')\n    \n    return data\n    \nif IS_INTERACTIVE:\n    data = test_preprocess_layer()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code defines a function called `get_train_dataset()` that generates a train dataset iterator. Here's an explanation of the function:\n\n1. `X` and `y`:\n   - These are the input features (`X`) and corresponding target labels (`y`) for the training dataset.\n\n2. `batch_size`:\n   - It specifies the number of samples in each batch. The default value is `BATCH_SIZE`.\n\n3. `sample_idxs`:\n   - It is an array containing the indices of the samples in the training dataset.\n\n4. Infinite loop:\n   - The function runs in an infinite loop to continuously generate batches of data.\n\n5. Random sampling:\n   - In each iteration, random indices are generated using `np.random.choice()` from the `sample_idxs` array.\n   - The number of random indices generated is equal to the `batch_size`.\n\n6. Noise generation:\n   - It generates random noise values using `np.random.uniform()` in the range [-1, 1].\n   - The noise is divided by 20 to scale it down.\n\n7. Creating inputs and outputs dictionaries:\n   - The `inputs` dictionary contains the features and labels for the input to the model.\n   - The 'frames' key in the `inputs` dictionary is assigned the values from `X` based on the random indices.\n   - The 'phrase' key in the `inputs` dictionary is assigned the values from `y` based on the random indices.\n   - The `outputs` dictionary is assigned the values from `y` based on the random indices.\n\n8. Yielding inputs and outputs:\n   - The function uses the `yield` keyword to return the `inputs` and `outputs` dictionaries as a generator.\n\nThe `get_train_dataset()` function generates a train dataset iterator that can be used in the training loop to provide batches of input and output data for the model. The generator continuously samples random batches of data from the training dataset and returns them as dictionaries.","metadata":{}},{"cell_type":"code","source":"# Train Dataset Iterator\ndef get_train_dataset(X, y, batch_size=BATCH_SIZE):\n    sample_idxs = np.arange(len(X))\n    while True:\n        # Get random indices\n        random_sample_idxs = np.random.choice(sample_idxs, batch_size)\n        \n        noise = np.random.uniform(-1,1, size=X[random_sample_idxs].shape) / 20\n        inputs = {\n            'frames': X[random_sample_idxs], \n            'phrase': y[random_sample_idxs],\n        }\n        outputs = y[random_sample_idxs]\n        \n        yield inputs, outputs","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train Dataset\ntrain_dataset = get_train_dataset(X_train, y_train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Training Steps Per Epoch\nTRAIN_STEPS_PER_EPOCH = math.ceil(N_TRAIN_SAMPLES / BATCH_SIZE)\nprint(f'TRAIN_STEPS_PER_EPOCH: {TRAIN_STEPS_PER_EPOCH}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code defines a function called `get_val_dataset()` that generates a validation dataset iterator. Here's an explanation of the function:\n\n1. `X` and `y`:\n   - These are the input features (`X`) and corresponding target labels (`y`) for the validation dataset.\n\n2. `batch_size`:\n   - It specifies the number of samples in each batch. The default value is `BATCH_SIZE`.\n\n3. `offsets`:\n   - It is an array containing the offsets to define the start indices of each batch in the validation dataset.\n   - The offsets are created using `np.arange()` with a step size of `batch_size`.\n\n4. Infinite loop:\n   - The function runs in an infinite loop to continuously generate batches of data from the validation dataset.\n\n5. Iterating over the validation set:\n   - In each iteration, the function iterates over the whole validation set using the `offsets` array.\n   - For each offset, it creates the `inputs` dictionary and assigns the corresponding batch of features and labels from `X` and `y`.\n   - The 'frames' key in the `inputs` dictionary is assigned the batch of features from `X`.\n   - The 'phrase' key in the `inputs` dictionary is assigned the batch of labels from `y`.\n   - The `outputs` dictionary is assigned the corresponding batch of labels from `y`.\n\n6. Yielding inputs and outputs:\n   - The function uses the `yield` keyword to return the `inputs` and `outputs` dictionaries as a generator.\n\nThe `get_val_dataset()` function generates a validation dataset iterator that can be used in the validation loop to provide batches of input and output data for evaluating the model's performance. The generator iterates over the validation dataset and returns batches of data as dictionaries.","metadata":{}},{"cell_type":"code","source":"def get_val_dataset(X, y, batch_size=BATCH_SIZE):\n    offsets = np.arange(0, len(X), batch_size)\n    while True:\n        # Iterate over whole validation set\n        for offset in offsets:\n            inputs = {\n                'frames': X[offset:offset+batch_size],\n                'phrase': y[offset:offset+batch_size],\n            }\n            outputs = y[offset:offset+batch_size]\n\n            yield inputs, outputs","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if USE_VAL:\n    val_dataset = get_val_dataset(X_val, y_val)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if USE_VAL:\n    N_VAL_STEPS_PER_EPOCH = math.ceil(N_VAL_SAMPLES / BATCH_SIZE)\n    print(f'N_VAL_STEPS_PER_EPOCH: {N_VAL_STEPS_PER_EPOCH}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code snippet defines various constants and parameters related to the model architecture and training. Here's an explanation of each part of the code:\n\n1. `LAYER_NORM_EPS`:\n   - It is a constant representing the epsilon value used in layer normalization. The value is set to 1e-6.\n\n2. `UNITS_ENCODER` and `UNITS_DECODER`:\n   - These constants represent the final embedding size for the encoder and decoder, respectively.\n   - The values are set to 256.\n\n3. `NUM_BLOCKS_ENCODER` and `NUM_BLOCKS_DECODER`:\n   - These constants represent the number of blocks in the encoder and decoder layers of the transformer model, respectively.\n   - The values are set to 2.\n\n4. `MLP_RATIO`:\n   - It is a constant representing the ratio of the hidden units in the feed-forward layer of the transformer model.\n   - The value is set to 4.\n\n5. Dropout rates:\n   - `EMBEDDING_DROPOUT` represents the dropout rate for the embedding layer of the transformer model. The value is set to 0.00 (no dropout).\n   - `MLP_DROPOUT_RATIO` represents the dropout ratio for the feed-forward layer in the transformer model. The value is set to 0.30 (30% dropout).\n   - `CLASSIFIER_DROPOUT_RATIO` represents the dropout ratio for the classifier layer in the transformer model. The value is set to 0.00 (no dropout).\n\n6. Initializers:\n   - `INIT_HE_UNIFORM` represents the He uniform initializer, which is commonly used for initializing weights in neural networks.\n   - `INIT_GLOROT_UNIFORM` represents the Glorot uniform initializer, which is another commonly used initializer.\n   - `INIT_ZEROS` represents the constant initializer with a value of 0.0.\n\n7. Activation function:\n   - `GELU` represents the Gaussian Error Linear Unit (GELU) activation function, which is a popular choice in transformer models.","metadata":{}},{"cell_type":"code","source":"# Epsilon value for layer normalisation\nLAYER_NORM_EPS = 1e-6\n\n# final embedding and transformer embedding size\nUNITS_ENCODER = 256\nUNITS_DECODER = 256\n\n# Transformer\nNUM_BLOCKS_ENCODER = 2\nNUM_BLOCKS_DECODER = 2\nMLP_RATIO = 4\n\n# Dropout\nEMBEDDING_DROPOUT = 0.00\nMLP_DROPOUT_RATIO = 0.30\nCLASSIFIER_DROPOUT_RATIO = 0.00\n\n# Initiailizers\nINIT_HE_UNIFORM = tf.keras.initializers.he_uniform\nINIT_GLOROT_UNIFORM = tf.keras.initializers.glorot_uniform\nINIT_ZEROS = tf.keras.initializers.constant(0.0)\n# Activations\nGELU = tf.keras.activations.gelu","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code defines a `LandmarkEmbedding` model class that embeds a landmark using fully connected layers. Here's an explanation of each part of the code:\n\n1. `LandmarkEmbedding` class:\n   - It is a subclass of `tf.keras.Model`.\n   - The `units` and `name` parameters are passed to the superclass constructor to set the name of the model.\n\n2. `build` method:\n   - It is overridden from the parent class and is called during model construction.\n   - The method initializes the empty embedding weights and the fully connected layers.\n   - The `empty_embedding` is a trainable weight initialized with zeros, representing the embedding for missing landmarks in a frame.\n   - The `dense` layer is a sequential model consisting of two dense layers.\n   - The dense layers have `units` number of units and are initialized using the Glorot uniform initializer for the first layer and the He uniform initializer for the second layer.\n   - The GELU activation function is used in the first dense layer.\n\n3. `call` method:\n   - It is overridden from the parent class and defines the forward pass of the model.\n   - The method takes an input tensor `x`, which represents the landmark data.\n   - It uses `tf.where()` to conditionally apply the empty embedding or the dense layers based on whether the landmark is missing in the frame.\n   - If the landmark is missing (sum of the landmark data along the last axis is 0), the `empty_embedding` is used.\n   - If the landmark is present, the dense layers are applied to embed the landmark data.\n\nThe `LandmarkEmbedding` model can be used to embed landmarks, considering the case where some landmarks might be missing in certain frames. It handles missing landmarks by using a separate embedding for them. The model's behavior is defined by the number of units in the embedding layers and can be customized by adjusting the initialization and activation functions.","metadata":{}},{"cell_type":"code","source":"# Embeds a landmark using fully connected layers\nclass LandmarkEmbedding(tf.keras.Model):\n    def __init__(self, units, name):\n        super(LandmarkEmbedding, self).__init__(name=f'{name}_embedding')\n        self.units = units\n        \n    def build(self, input_shape):\n        # Embedding for missing landmark in frame, initizlied with zeros\n        self.empty_embedding = self.add_weight(\n            name=f'{self.name}_empty_embedding',\n            shape=[self.units],\n            initializer=INIT_ZEROS,\n        )\n        # Embedding\n        self.dense = tf.keras.Sequential([\n            tf.keras.layers.Dense(self.units, name=f'{self.name}_dense_1', use_bias=False, kernel_initializer=INIT_GLOROT_UNIFORM, activation=GELU),\n            tf.keras.layers.Dense(self.units, name=f'{self.name}_dense_2', use_bias=False, kernel_initializer=INIT_HE_UNIFORM),\n        ], name=f'{self.name}_dense')\n\n    def call(self, x):\n        return tf.where(\n                # Checks whether landmark is missing in frame\n                tf.reduce_sum(x, axis=2, keepdims=True) == 0,\n                # If so, the empty embedding is used\n                self.empty_embedding,\n                # Otherwise the landmark data is embedded\n                self.dense(x),\n            )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code defines an `Embedding` model class that creates an embedding for each frame. Here's an explanation of each part of the code:\n\n1. `Embedding` class:\n   - It is a subclass of `tf.keras.Model`.\n   - The constructor does not require any parameters.\n\n2. `build` method:\n   - It is overridden from the parent class and is called during model construction.\n   - The method initializes the positional embedding and the `dominant_hand_embedding` (a `LandmarkEmbedding` instance).\n\n3. `positional_embedding`:\n   - It is a trainable variable representing the positional embedding for each frame index.\n   - The initial value is a tensor of zeros with a shape of `[N_TARGET_FRAMES, UNITS_ENCODER]`.\n   - The variable is trainable, allowing the model to learn the positional encoding.\n\n4. `dominant_hand_embedding`:\n   - It is an instance of the `LandmarkEmbedding` class, which creates an embedding for the dominant hand landmarks.\n   - The embedding size is set to `UNITS_ENCODER` (256).\n\n5. `call` method:\n   - It is overridden from the parent class and defines the forward pass of the model.\n   - The method takes an input tensor `x`, which represents the input data.\n   - It applies normalization to the input data by subtracting the means and dividing by the standard deviations.\n   - The `LandmarkEmbedding` instance is used to embed the dominant hand landmarks.\n   - The positional embedding is added to the embedded landmarks.\n   - The resulting tensor is returned as the output of the model.\n\nThe `Embedding` model takes input data and performs normalization, embedding of landmarks, and addition of positional encoding. It can be used as a component of a larger model for further processing and analysis of the data.","metadata":{}},{"cell_type":"code","source":"class Embedding(tf.keras.Model):\n    def __init__(self):\n        super(Embedding, self).__init__()\n\n    def build(self, input_shape):\n        # Positional embedding for each frame index\n        self.positional_embedding = tf.Variable(\n            initial_value=tf.zeros([N_TARGET_FRAMES, UNITS_ENCODER], dtype=tf.float32),\n            trainable=True,\n            name='embedding_positional_encoder',\n        )\n        # Embedding layer for Landmarks\n        self.dominant_hand_embedding = LandmarkEmbedding(UNITS_ENCODER, 'dominant_hand')\n\n    def call(self, x, training=False):\n        # Normalize\n        x = tf.where(\n                tf.math.equal(x, 0.0),\n                0.0,\n                (x - MEANS) / STDS,\n            )\n        # Dominant Hand\n        x = self.dominant_hand_embedding(x)\n        # Add Positional Encoding\n        x = x + self.positional_embedding\n        \n        return x","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code defines a scaled dot-product attention mechanism in the `scaled_dot_product` function and a multi-head attention layer in the `MultiHeadAttention` class. Here's an explanation of each part of the code:\n\n1. `scaled_dot_product` function:\n   - It takes query (`q`), key (`k`), and value (`v`) tensors as input.\n   - It performs the scaled dot-product attention operation by calculating the dot product between `q` and the transpose of `k`.\n   - The result is divided by the square root of the depth of `q` to scale the dot product.\n   - The `softmax` function is then applied to the scaled dot product, using the `attention_mask` to mask certain elements if provided.\n   - The result is multiplied with `v` to obtain the weighted sum of values based on the attention scores.\n   - The final output `z` has the same shape as `q`, `k`, and `v`.\n\n2. `MultiHeadAttention` class:\n   - It is a subclass of `tf.keras.layers.Layer`.\n   - The constructor takes `d_model` (dimension of the model) and `num_of_heads` as parameters.\n   - It initializes the number of heads, the depth of each head, and the trainable weight matrices (`wq`, `wk`, `wv`, and `wo`) for each head.\n   - It also initializes a `softmax` layer for the attention softmax calculation.\n\n3. `call` method:\n   - It takes query (`q`), key (`k`), and value (`v`) tensors as input.\n   - It applies the multi-head attention mechanism by computing attention scores for each head and concatenating the results.\n   - For each head, it applies separate linear transformations (`wq`, `wk`, and `wv`) to the query, key, and value tensors.\n   - It then calls the `scaled_dot_product` function to compute the attention output for each head.\n   - The attention outputs are concatenated along the last axis.\n   - Finally, the concatenated outputs are transformed using the `wo` linear layer to obtain the multi-head attention output.\n\nThe `MultiHeadAttention` layer can be used in a transformer model to perform multi-head attention operations. It takes query, key, and value tensors as input and produces the multi-head attention output. The attention mechanism uses the `scaled_dot_product` function for computing attention scores.","metadata":{}},{"cell_type":"code","source":"# based on: https://stackoverflow.com/questions/67342988/verifying-the-implementation-of-multihead-attention-in-transformer\n# replaced softmax with softmax layer to support masked softmax\ndef scaled_dot_product(q,k,v, softmax, attention_mask):\n    #calculates Q . K(transpose)\n    qkt = tf.matmul(q,k,transpose_b=True)\n    #caculates scaling factor\n    dk = tf.math.sqrt(tf.cast(q.shape[-1],dtype=tf.float32))\n    scaled_qkt = qkt/dk\n    softmax = softmax(scaled_qkt, mask=attention_mask)\n    z = tf.matmul(softmax,v)\n    #shape: (m,Tx,depth), same shape as q,k,v\n    return z\n\nclass MultiHeadAttention(tf.keras.layers.Layer):\n    def __init__(self,d_model,num_of_heads):\n        super(MultiHeadAttention,self).__init__()\n        self.d_model = d_model\n        self.num_of_heads = num_of_heads\n        self.depth = d_model//num_of_heads\n        self.wq = [tf.keras.layers.Dense(self.depth) for i in range(num_of_heads)]\n        self.wk = [tf.keras.layers.Dense(self.depth) for i in range(num_of_heads)]\n        self.wv = [tf.keras.layers.Dense(self.depth) for i in range(num_of_heads)]\n        self.wo = tf.keras.layers.Dense(d_model)\n        self.softmax = tf.keras.layers.Softmax()\n        \n    def call(self, q, k, v, attention_mask=None):\n        \n        multi_attn = []\n        for i in range(self.num_of_heads):\n            Q = self.wq[i](q)\n            K = self.wk[i](k)\n            V = self.wv[i](v)\n            multi_attn.append(scaled_dot_product(Q,K,V, self.softmax, attention_mask))\n            \n        multi_head = tf.concat(multi_attn, axis=-1)\n        multi_head_attention = self.wo(multi_head)\n        return multi_head_attention","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code defines an `Encoder` model class based on multiple transformer blocks. Here's an explanation of each part of the code:\n\n1. `Encoder` class:\n   - It is a subclass of `tf.keras.Model`.\n   - The constructor takes the number of transformer blocks (`num_blocks`) as a parameter.\n   - It initializes the number of blocks.\n\n2. `build` method:\n   - It is overridden from the parent class and is called during model construction.\n   - The method initializes the layer normalization (`ln_1s` and `ln_2s`), multi-head attention (`mhas`), and multi-layer perception (`mlps`) layers for each transformer block.\n\n3. Transformer blocks:\n   - For each transformer block, the following layers are initialized:\n     - `ln_1`: The first layer normalization layer.\n     - `mha`: The multi-head attention layer with `UNITS_ENCODER` dimension and 8 heads.\n     - `ln_2`: The second layer normalization layer.\n     - `mlp`: The multi-layer perception layer, consisting of two dense layers with activation, dropout, and weight initialization.\n\n4. `call` method:\n   - It is overridden from the parent class and defines the forward pass of the model.\n   - The method takes input tensors `x` and `x_inp`, where `x` represents the input data and `x_inp` represents the attention mask to ignore missing frames.\n   - The attention mask is calculated based on the sum of `x_inp` along the last axis, where a non-zero value indicates the presence of a landmark.\n   - The attention mask is expanded and repeated to match the shape of the input data.\n   - The input data is iterated over the transformer blocks.\n   - For each block, the input is passed through a layer normalization, multi-head attention, and another layer normalization.\n   - The output of the multi-head attention is added to the input (residual connection) and passed through the first layer normalization.\n   - The output is then passed through the multi-layer perception and added to the previous output (residual connection).\n   - The final output of the encoder is returned.\n\nThe `Encoder` model consists of multiple transformer blocks and can be used to process the input data by applying self-attention and feed-forward operations. It incorporates residual connections and layer normalization for each block, facilitating the flow of information through the blocks.","metadata":{}},{"cell_type":"code","source":"# Encoder based on multiple transformer blocks\nclass Encoder(tf.keras.Model):\n    def __init__(self, num_blocks):\n        super(Encoder, self).__init__(name='encoder')\n        self.num_blocks = num_blocks\n    \n    def build(self, input_shape):\n        self.ln_1s = []\n        self.mhas = []\n        self.ln_2s = []\n        self.mlps = []\n        # Make Transformer Blocks\n        for i in range(self.num_blocks):\n            # First Layer Normalisation\n            self.ln_1s.append(tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS))\n            # Multi Head Attention\n            self.mhas.append(MultiHeadAttention(UNITS_ENCODER, 8))\n            # Second Layer Normalisation\n            self.ln_2s.append(tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS))\n            # Multi Layer Perception\n            self.mlps.append(tf.keras.Sequential([\n                tf.keras.layers.Dense(UNITS_ENCODER * MLP_RATIO, activation=GELU, kernel_initializer=INIT_GLOROT_UNIFORM),\n                tf.keras.layers.Dropout(MLP_DROPOUT_RATIO),\n                tf.keras.layers.Dense(UNITS_ENCODER, kernel_initializer=INIT_HE_UNIFORM),\n            ]))\n        \n    def call(self, x, x_inp):\n        # Attention mask to ignore missing frames\n        attention_mask = tf.where(tf.math.reduce_sum(x_inp, axis=[2]) == 0.0, 0.0, 1.0)\n        attention_mask = tf.expand_dims(attention_mask, axis=1)\n        attention_mask = tf.repeat(attention_mask, repeats=N_TARGET_FRAMES, axis=1)\n        # Iterate input over transformer blocks\n        for ln_1, mha, ln_2, mlp in zip(self.ln_1s, self.mhas, self.ln_2s, self.mlps):\n            x = ln_1(x + mha(x, x, x, attention_mask=attention_mask))\n            x = ln_2(x + mlp(x))\n    \n        return x","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code defines a `Decoder` model class based on multiple transformer blocks. Here's an explanation of each part of the code:\n\n1. `Decoder` class:\n   - It is a subclass of `tf.keras.Model`.\n   - The constructor takes the number of transformer blocks (`num_blocks`) as a parameter.\n   - It initializes the number of blocks.\n\n2. `build` method:\n   - It is overridden from the parent class and is called during model construction.\n   - The method initializes the positional embedding (`positional_embedding`), character embedding (`char_emb`), positional encoder multi-head attention (`pos_emb_mha`), and layer normalization (`pos_emb_ln`) layers.\n   - It also initializes the layer normalization (`ln_1s`), multi-head attention (`mhas`), and multi-layer perception (`mlps`) layers for each transformer block.\n\n3. Causal attention mask:\n   - The `get_causal_attention_mask` method creates a causal attention mask that masks future positions in the sequence to enforce causality.\n\n4. `call` method:\n   - It is overridden from the parent class and defines the forward pass of the model.\n   - The method takes encoder outputs (`encoder_outputs`) and a phrase tensor (`phrase`) as input.\n   - The phrase tensor represents the target phrase for decoding.\n   - The input phrase is preprocessed by casting it to `int32`, prepending the start of sentence (SOS) token, and padding it with the pad (PAD) token.\n   - The causal attention mask is created using the `get_causal_attention_mask` method.\n   - The positional embedding and character embedding are added together and passed through the positional encoder multi-head attention, followed by layer normalization.\n   - The input is then iterated over the transformer blocks.\n   - For each block, the input is passed through a layer normalization, multi-head attention (with encoder outputs as the key and value), and another layer normalization.\n   - The output is then passed through the multi-layer perception and added to the previous output (residual connection).\n   - Finally, the output is sliced to retain only the desired number of characters in the output sequence.\n\nThe `Decoder` model takes encoder outputs and a target phrase as input and generates the decoder output by applying self-attention and feed-forward operations. It incorporates residual connections, layer normalization, and positional encoding for each block, facilitating the flow of information and capturing the positional information in the sequence.","metadata":{}},{"cell_type":"code","source":"# Decoder based on multiple transformer blocks\nclass Decoder(tf.keras.Model):\n    def __init__(self, num_blocks):\n        super(Decoder, self).__init__(name='decoder')\n        self.num_blocks = num_blocks\n    \n    def build(self, input_shape):\n        # Positional Embedding, initialized with zeros\n        self.positional_embedding = tf.Variable(\n            initial_value=tf.zeros([N_TARGET_FRAMES, UNITS_DECODER], dtype=tf.float32),\n            trainable=True,\n            name='embedding_positional_encoder',\n        )\n        # Character Embedding\n        self.char_emb = tf.keras.layers.Embedding(N_UNIQUE_CHARACTERS, UNITS_DECODER, embeddings_initializer=INIT_ZEROS)\n        # Positional Encoder MHA\n        self.pos_emb_mha = MultiHeadAttention(UNITS_DECODER, 8)\n        self.pos_emb_ln = tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS)\n        # First Layer Normalisation\n        self.ln_1s = []\n        self.mhas = []\n        self.ln_2s = []\n        self.mlps = []\n        # Make Transformer Blocks\n        for i in range(self.num_blocks):\n            # First Layer Normalisation\n            self.ln_1s.append(tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS))\n            # Multi Head Attention\n            self.mhas.append(MultiHeadAttention(UNITS_DECODER, 8))\n            # Second Layer Normalisation\n            self.ln_2s.append(tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS))\n            # Multi Layer Perception\n            self.mlps.append(tf.keras.Sequential([\n                tf.keras.layers.Dense(UNITS_DECODER * MLP_RATIO, activation=GELU, kernel_initializer=INIT_GLOROT_UNIFORM),\n                tf.keras.layers.Dropout(MLP_DROPOUT_RATIO),\n                tf.keras.layers.Dense(UNITS_DECODER, kernel_initializer=INIT_HE_UNIFORM),\n            ]))\n            \n    def get_causal_attention_mask(self, B):\n        i = tf.range(N_TARGET_FRAMES)[:, tf.newaxis]\n        j = tf.range(N_TARGET_FRAMES)\n        mask = tf.cast(i >= j, dtype=tf.int32)\n        mask = tf.reshape(mask, (1, N_TARGET_FRAMES, N_TARGET_FRAMES))\n        mult = tf.concat(\n            [tf.expand_dims(B, -1), tf.constant([1, 1], dtype=tf.int32)],\n            axis=0,\n        )\n        mask = tf.tile(mask, mult)\n        mask = tf.cast(mask, tf.float32)\n        return mask\n        \n    def call(self, encoder_outputs, phrase):\n        # Batch Size\n        B = tf.shape(encoder_outputs)[0]\n        # Cast to INT32\n        phrase = tf.cast(phrase, tf.int32)\n        # Prepend SOS Token\n        phrase = tf.pad(phrase, [[0,0], [1,0]], constant_values=SOS_TOKEN, name='prepend_sos_token')\n        # Pad With PAD Token\n        phrase = tf.pad(phrase, [[0,0], [0,N_TARGET_FRAMES-MAX_PHRASE_LENGTH-1]], constant_values=PAD_TOKEN, name='append_pad_token')\n        # Causal Mask\n        causal_mask = self.get_causal_attention_mask(B)\n        # Positional Embedding\n        x = self.positional_embedding + self.char_emb(phrase)\n        # Causal Attention\n        x = self.pos_emb_ln(x + self.pos_emb_mha(x, x, x, attention_mask=causal_mask))\n        # Iterate input over transformer blocks\n        for ln_1, mha, ln_2, mlp in zip(self.ln_1s, self.mhas, self.ln_2s, self.mlps):\n            x = ln_1(x + mha(x, encoder_outputs, encoder_outputs, attention_mask=causal_mask))\n            x = ln_2(x + mlp(x))\n        # Slice 31 Characters\n        x = tf.slice(x, [0, 0, 0], [-1, MAX_PHRASE_LENGTH, -1])\n    \n        return x","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code defines the `get_causal_attention_mask` function, which generates a causal attention mask to prevent the decoder from attending to future characters that it needs to predict. Here's an explanation of the function:\n\n1. `get_causal_attention_mask` function:\n   - It takes the batch size (`B`) as a parameter.\n   - It generates a mask that has a shape of `(B, N_TARGET_FRAMES, N_TARGET_FRAMES)` to represent the attention weights between each target frame and all previous frames.\n   - The mask is created using two range tensors, `i` and `j`, which represent the indices of the target frames.\n   - The mask is computed by comparing `i` with `j` element-wise, where a value of 1 indicates that the target frame is before or at the same position as the frame being attended to.\n   - The mask is reshaped to have a shape of `(1, N_TARGET_FRAMES, N_TARGET_FRAMES)` to match the batch size.\n   - The mask is tiled along the batch dimension and broadcasted to have a shape of `(B, N_TARGET_FRAMES, N_TARGET_FRAMES)`.\n   - Finally, the mask is cast to `float32` and returned.\n\nWhen calling the `get_causal_attention_mask` function with a batch size of 1, it generates a causal attention mask with the shape `(1, N_TARGET_FRAMES, N_TARGET_FRAMES)`. The mask ensures that each target frame attends to only previous frames or the same frame, preventing the decoder from attending to future characters during the decoding process.","metadata":{}},{"cell_type":"code","source":"# Causal Attention to make decoder not attent to future characters which it needs to predict\ndef get_causal_attention_mask(B):\n    i = tf.range(N_TARGET_FRAMES)[:, tf.newaxis]\n    j = tf.range(N_TARGET_FRAMES)\n    mask = tf.cast(i >= j, dtype=tf.int32)\n    mask = tf.reshape(mask, (1, N_TARGET_FRAMES, N_TARGET_FRAMES))\n    mult = tf.concat(\n        [tf.expand_dims(B, -1), tf.constant([1, 1], dtype=tf.int32)],\n        axis=0,\n    )\n    mask = tf.tile(mask, mult)\n    mask = tf.cast(mask, tf.float32)\n    return mask\n\nget_causal_attention_mask(1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code defines a `TopKAccuracy` class, which is a custom metric subclassed from `tf.keras.metrics.Metric`. This metric calculates the top-K accuracy for multi-dimensional output. Here's an explanation of each part of the code:\n\n1. `TopKAccuracy` class:\n   - It inherits from `tf.keras.metrics.Metric` and represents the custom metric for top-K accuracy.\n   - The constructor takes `k` as a parameter to specify the value of K for top-K accuracy.\n   - It initializes the `top_k_acc` metric, which is the `SparseTopKCategoricalAccuracy` metric from TensorFlow.\n\n2. `update_state` method:\n   - It is overridden from the parent class and is called to update the metric state.\n   - The method takes `y_true` and `y_pred` as input, representing the true labels and predicted probabilities, respectively.\n   - The `y_true` and `y_pred` tensors are reshaped to have the shape `[-1]` and `[-1, N_UNIQUE_CHARACTERS]`, respectively, to handle multi-dimensional outputs.\n   - The character indices are extracted from `y_true` to filter out padding tokens.\n   - The filtered `y_true` and `y_pred` tensors are passed to the `update_state` method of `top_k_acc` metric to update its state.\n\n3. `result` method:\n   - It is overridden from the parent class and returns the current result of the metric, which is the top-K accuracy.\n\n4. `reset_state` method:\n   - It is overridden from the parent class and resets the state of the metric.\n\nThe `TopKAccuracy` metric calculates the top-K accuracy for multi-dimensional output by utilizing the `SparseTopKCategoricalAccuracy` metric from TensorFlow. It provides the `update_state` method to update the metric state, the `result` method to get the current result, and the `reset_state` method to reset the metric state when needed.","metadata":{}},{"cell_type":"code","source":"# TopK accuracy for multi dimensional output\nclass TopKAccuracy(tf.keras.metrics.Metric):\n    def __init__(self, k, **kwargs):\n        super(TopKAccuracy, self).__init__(name=f'top{k}acc', **kwargs)\n        self.top_k_acc = tf.keras.metrics.SparseTopKCategoricalAccuracy(k=k)\n\n    def update_state(self, y_true, y_pred, sample_weight=None):\n        y_true = tf.reshape(y_true, [-1])\n        y_pred = tf.reshape(y_pred, [-1, N_UNIQUE_CHARACTERS])\n        character_idxs = tf.where(y_true < N_UNIQUE_CHARACTERS0)\n        y_true = tf.gather(y_true, character_idxs, axis=0)\n        y_pred = tf.gather(y_pred, character_idxs, axis=0)\n        self.top_k_acc.update_state(y_true, y_pred)\n\n    def result(self):\n        return self.top_k_acc.result()\n    \n    def reset_state(self):\n        self.top_k_acc.reset_state()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code creates an array `loss_weights` with all elements initially set to 1. It also sets the loss weight of the pad token to 0. Here's a breakdown of the code:\n\n1. `loss_weights` array:\n   - It is initialized as an array of shape `(N_UNIQUE_CHARACTERS,)`.\n   - All elements of the array are initially set to 1 using `np.ones`.\n   - The `dtype` is set to `np.float32` to ensure the loss weights are floating-point numbers.\n\n2. Setting the loss weight of the pad token:\n   - The pad token index is represented by `PAD_TOKEN`.\n   - The code assigns a value of 0 to the element at the index `PAD_TOKEN` in the `loss_weights` array.\n\nThe `loss_weights` array is useful for assigning different weights to different tokens during training or calculating losses. In this case, all tokens except the pad token have a weight of 1, while the pad token has a weight of 0, indicating that it should not contribute to the loss calculation.","metadata":{}},{"cell_type":"code","source":"# Create Initial Loss Weights All Set To 1\nloss_weights = np.ones(N_UNIQUE_CHARACTERS, dtype=np.float32)\n# Set Loss Weight Of Pad Token To 0\nloss_weights[PAD_TOKEN] = 0","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below code defines a custom loss function `scce_with_ls` (sparse categorical cross-entropy with label smoothing) for training a model. It utilizes the `tf.keras.losses.categorical_crossentropy` function with native label smoothing support. Here's an explanation of the code:\n\n1. `scce_with_ls` function:\n   - It takes `y_true` and `y_pred` as input, representing the true labels and predicted probabilities, respectively.\n   - `y_true` is cast to `int32` and then one-hot encoded using `tf.one_hot`. The one-hot encoding is applied along the last dimension (axis=2) to match the shape of `y_pred`.\n   - The `categorical_crossentropy` function from `tf.keras.losses` is used to compute the loss between the one-hot encoded `y_true` and `y_pred`.\n   - The `label_smoothing` parameter is set to 0.25, indicating the amount of smoothing to be applied to the one-hot encoded labels. This helps prevent overfitting and encourage generalization by reducing the confidence of the true label and spreading it across other labels.\n\nThe `scce_with_ls` loss function provides a way to apply label smoothing during training, which can be useful for improving the performance and generalization of the model.","metadata":{}},{"cell_type":"code","source":"# source:: https://stackoverflow.com/questions/60689185/label-smoothing-for-sparse-categorical-crossentropy\ndef scce_with_ls(y_true, y_pred):\n    # One Hot Encode Sparsely Encoded Target Sign\n    y_true = tf.cast(y_true, tf.int32)\n    y_true = tf.one_hot(y_true, N_UNIQUE_CHARACTERS, axis=2)\n    # Categorical Crossentropy with native label smoothing support\n    return tf.keras.losses.categorical_crossentropy(y_true, y_pred, label_smoothing=0.25)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `get_model` function constructs and compiles a TensorFlow model for the ASL Fingerspelling task. Here's an overview of the function:\n\n1. Input layers:\n   - Two input layers are defined: `frames_inp` for the frames data and `phrase_inp` for the phrase sequence.\n\n2. Frames data processing:\n   - The `frames_inp` is passed through the `Embedding` layer to perform data normalization and embedding.\n\n3. Encoder Transformer Blocks:\n   - The processed frames data is passed through the `Encoder` layer, which consists of multiple transformer blocks. This layer performs encoding and contextual representation of the frames.\n\n4. Decoder:\n   - The encoded frames data and the `phrase_inp` are passed through the `Decoder` layer, which also consists of multiple transformer blocks. This layer performs decoding and generates the predicted phrase sequence.\n\n5. Classifier:\n   - The output of the decoder is passed through a classifier, which consists of a dropout layer and a dense layer with softmax activation. This layer generates the final output probabilities for each unique character.\n\n6. Model compilation:\n   - The TensorFlow model is created using the input and output layers.\n   - The loss function is set to `scce_with_ls`, which is the sparse categorical cross-entropy with label smoothing.\n   - The Adam optimizer with weight decay and gradient clipping (clipnorm) is used.\n   - The top-k accuracy metrics are defined using the `TopKAccuracy` class.\n   - The loss weights are set using the `loss_weights` array defined earlier.\n\n7. Model return:\n   - The compiled model is returned.\n\nOverall, this function provides a convenient way to create and compile the model architecture for the ASL Fingerspelling task, using the specified layers, loss function, optimizer, and metrics.","metadata":{}},{"cell_type":"code","source":"def get_model():\n    # Inputs\n    frames_inp = tf.keras.layers.Input([N_TARGET_FRAMES, N_COLS], dtype=tf.float32, name='frames')\n    phrase_inp = tf.keras.layers.Input([MAX_PHRASE_LENGTH], dtype=tf.int32, name='phrase')\n    # Frames\n    x = frames_inp\n\n    # Embedding\n    x = Embedding()(x, frames_inp)\n    \n    # Encoder Transformer Blocks\n    x = Encoder(NUM_BLOCKS_ENCODER)(x, frames_inp)\n    \n    # Decoder\n    x = Decoder(NUM_BLOCKS_DECODER)(x, phrase_inp)\n    \n    # Classifier\n    x = tf.keras.Sequential([\n        # Dropout\n        tf.keras.layers.Dropout(CLASSIFIER_DROPOUT_RATIO),\n        # Output Neurons\n        tf.keras.layers.Dense(N_UNIQUE_CHARACTERS, activation=tf.keras.activations.softmax, kernel_initializer=INIT_HE_UNIFORM),\n    ], name='classifier')(x)\n    \n    outputs = x\n    \n    # Create Tensorflow Model\n    model = tf.keras.models.Model(inputs=[frames_inp, phrase_inp], outputs=outputs)\n    \n    # Simple Categorical Crossentropy Loss\n#     loss = tf.keras.losses.SparseCategoricalCrossentropy()\n    # Categorical Crossentropy Loss With Label Smoothing\n    loss = scce_with_ls\n    \n    # Adam Optimizer with weight decay\n    optimizer = tf.keras.optimizers.Adam(clipnorm=5.0)\n    \n    # TopK Metrics\n    metrics = [\n        TopKAccuracy(1),\n        TopKAccuracy(5),\n    ]\n    \n    model.compile(\n        loss=loss,\n        optimizer=optimizer,\n        metrics=metrics,\n        loss_weights=loss_weights,\n    )\n    \n    return model","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Input data\nfor k, v in X_batch.items():\n    print(f'{k}: {v.shape}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.keras.backend.clear_session()\n\nmodel = get_model()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot model summary\nmodel.summary(expand_nested=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot Model Architecture\ntf.keras.utils.plot_model(model, show_shapes=True, show_dtype=True, show_layer_names=True, expand_nested=True, show_layer_activations=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Dummy Prediction\ny_pred = model.predict(\n    val_dataset if USE_VAL else train_dataset,\n    steps=N_VAL_STEPS_PER_EPOCH if USE_VAL else TRAIN_STEPS_PER_EPOCH,\n    verbose=VERBOSE,\n)\n\nprint(f'# NaN Values In Predictions: {np.isnan(y_pred).sum()}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `lrfn` function implements a learning rate schedule based on a cosine annealing with warmup method. Here's how it works:\n\n1. Arguments:\n   - `current_step`: The current training step (iteration).\n   - `num_warmup_steps`: The number of warmup steps at the beginning of training.\n   - `lr_max`: The maximum learning rate.\n   - `num_cycles`: The number of cycles for the cosine annealing schedule. Default is 0.50, indicating half a cycle.\n   - `num_training_steps`: The total number of training steps (iterations).\n\n2. Warmup phase:\n   - If the `current_step` is less than `num_warmup_steps`, the learning rate is adjusted for the warmup phase.\n   - The warmup phase gradually increases the learning rate from 0 to `lr_max`.\n   - There are two warmup methods available: logarithmic (`log`) and linear (`linear`).\n     - Logarithmic warmup: The learning rate is calculated as `lr_max * 0.10 ** (num_warmup_steps - current_step)`.\n     - Linear warmup: The learning rate is calculated as `lr_max * 2 ** -(num_warmup_steps - current_step)`.\n\n3. Cosine annealing phase:\n   - If the `current_step` is greater than or equal to `num_warmup_steps`, the learning rate is adjusted using the cosine annealing schedule.\n   - The progress is calculated as the ratio of completed steps to the total number of steps in the training schedule.\n   - The learning rate is calculated using the cosine annealing formula:\n     - `0.5 * (1.0 + math.cos(math.pi * num_cycles * 2.0 * progress)) * lr_max`\n   - The learning rate gradually decreases from `lr_max` to 0 over the specified number of cycles.\n\n4. Return:\n   - The calculated learning rate is returned.\n\nThis learning rate function provides a smooth transition from warmup to the cosine annealing schedule, allowing the model to explore a wider range of learning rates initially and then gradually decrease the learning rate for better convergence.","metadata":{}},{"cell_type":"code","source":"def lrfn(current_step, num_warmup_steps, lr_max, num_cycles=0.50, num_training_steps=N_EPOCHS):\n    \n    if current_step < num_warmup_steps:\n        if WARMUP_METHOD == 'log':\n            return lr_max * 0.10 ** (num_warmup_steps - current_step)\n        else:\n            return lr_max * 2 ** -(num_warmup_steps - current_step)\n    else:\n        progress = float(current_step - num_warmup_steps) / float(max(1, num_training_steps - num_warmup_steps))\n\n        return max(0.0, 0.5 * (1.0 + math.cos(math.pi * float(num_cycles) * 2.0 * progress))) * lr_max","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `plot_lr_schedule` function is used to plot the learning rate schedule over the epochs. Here's how it works:\n\n1. Arguments:\n   - `lr_schedule`: A list of learning rates calculated for each epoch.\n   - `epochs`: The total number of epochs.\n\n2. Plotting:\n   - The function creates a figure with a size of 20x10 inches.\n   - It plots the learning rates as a line plot with `None` values at the beginning and end for better visualization.\n   - The x-axis represents the epochs, and the y-axis represents the learning rates.\n   - The x-axis labels are set based on the number of epochs, with specific intervals for better readability.\n   - The y-axis limit is increased by 10% to provide better readability.\n   - The title of the plot includes information about the start, maximum, and final learning rates.\n   - Each learning rate value is plotted as a black circle marker on the line plot.\n   - The learning rate values are annotated with their corresponding values above the markers.\n   - The x-axis is labeled as \"Epoch,\" and the y-axis is labeled as \"Learning Rate.\"\n   - Grid lines are added to the plot for better visualization.\n\n3. Display:\n   - The plot is displayed using `plt.show()`.\n\nThe `plot_lr_schedule` function can be used to visualize the learning rate schedule and understand how the learning rate changes over the epochs.","metadata":{}},{"cell_type":"code","source":"def plot_lr_schedule(lr_schedule, epochs):\n    fig = plt.figure(figsize=(20, 10))\n    plt.plot([None] + lr_schedule + [None])\n    # X Labels\n    x = np.arange(1, epochs + 1)\n    x_axis_labels = [i if epochs <= 40 or i % 5 == 0 or i == 1 else None for i in range(1, epochs + 1)]\n    plt.xlim([1, epochs])\n    plt.xticks(x, x_axis_labels) # set tick step to 1 and let x axis start at 1\n    \n    # Increase y-limit for better readability\n    plt.ylim([0, max(lr_schedule) * 1.1])\n    \n    # Title\n    schedule_info = f'start: {lr_schedule[0]:.1E}, max: {max(lr_schedule):.1E}, final: {lr_schedule[-1]:.1E}'\n    plt.title(f'Step Learning Rate Schedule, {schedule_info}', size=18, pad=12)\n    \n    # Plot Learning Rates\n    for x, val in enumerate(lr_schedule):\n        if epochs <= 40 or x % 5 == 0 or x is epochs - 1:\n            if x < len(lr_schedule) - 1:\n                if lr_schedule[x - 1] < val:\n                    ha = 'right'\n                else:\n                    ha = 'left'\n            elif x == 0:\n                ha = 'right'\n            else:\n                ha = 'left'\n            plt.plot(x + 1, val, 'o', color='black');\n            offset_y = (max(lr_schedule) - min(lr_schedule)) * 0.02\n            plt.annotate(f'{val:.1E}', xy=(x + 1, val + offset_y), size=12, ha=ha)\n    \n    plt.xlabel('Epoch', size=16, labelpad=5)\n    plt.ylabel('Learning Rate', size=16, labelpad=5)\n    plt.grid()\n    plt.show()\n\n# Learning rate for encoder\nLR_SCHEDULE = [lrfn(step, num_warmup_steps=N_WARMUP_EPOCHS, lr_max=LR_MAX, num_cycles=0.50) for step in range(N_EPOCHS)]\n# Plot Learning Rate Schedule\nplot_lr_schedule(LR_SCHEDULE, epochs=N_EPOCHS)\n# Learning Rate Callback\nlr_callback = tf.keras.callbacks.LearningRateScheduler(lambda step: LR_SCHEDULE[step], verbose=0)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Evaluate Initialized Model On Validation Data\ny_pred = model.evaluate(\n    val_dataset if USE_VAL else train_dataset,\n    steps=N_VAL_STEPS_PER_EPOCH if USE_VAL else TRAIN_STEPS_PER_EPOCH,\n    verbose=VERBOSE,\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# baseline accuracy when only pad token is predicted\nif USE_VAL:\n    baseline_accuracy = np.mean(y_val == PAD_TOKEN)\nelse:\n    baseline_accuracy = np.mean(y_train == PAD_TOKEN)\nprint(f'Baseline Accuracy: {baseline_accuracy:.4f}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Custom callback to update weight decay with learning rate\nclass WeightDecayCallback(tf.keras.callbacks.Callback):\n    def __init__(self, wd_ratio=WD_RATIO):\n        self.step_counter = 0\n        self.wd_ratio = wd_ratio\n    \n    def on_epoch_begin(self, epoch, logs=None):\n        model.optimizer.weight_decay = model.optimizer.learning_rate * self.wd_ratio\n        print(f'learning rate: {model.optimizer.learning_rate.numpy():.2e}, weight decay: {model.optimizer.weight_decay.numpy():.2e}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The code is for training the model if `TRAIN_MODEL` is `True`. Here's a breakdown of the training process:\n\n1. `tf.keras.backend.clear_session()`: Clears all models in the GPU to start with a fresh model.\n\n2. `model = get_model()`: Creates a new instance of the model using the `get_model()` function.\n\n3. `model.summary()`: Prints a summary of the model architecture.\n\n4. `history = model.fit(...)`: Trains the model using the `fit()` method. The parameters used are as follows:\n   - `x=train_dataset`: Specifies the training dataset.\n   - `steps_per_epoch=TRAIN_STEPS_PER_EPOCH`: Specifies the number of steps per epoch.\n   - `epochs=N_EPOCHS`: Specifies the number of epochs for training.\n   - `validation_data=val_dataset if USE_VAL else None`: Specifies the validation dataset if `USE_VAL` is `True`, otherwise set to `None`.\n   - `validation_steps=N_VAL_STEPS_PER_EPOCH if USE_VAL else None`: Specifies the number of validation steps per epoch if `USE_VAL` is `True`, otherwise set to `None`.\n   - `callbacks=[lr_callback, WeightDecayCallback(), ealystop_callback]`: Specifies the callbacks to be used during training, including the learning rate callback, weight decay callback, and early stopping callback.\n   - `verbose=VERBOSE`: Specifies the verbosity mode for training.\n\nThe `model.fit()` function returns a `history` object that contains information about the training process, such as the loss and metrics values at each epoch.","metadata":{}},{"cell_type":"code","source":"ealystop_callback = tf.keras.callbacks.EarlyStopping(monitor='loss', patience=3)\n\nif TRAIN_MODEL:\n    # Clear all models in GPU\n    tf.keras.backend.clear_session()\n\n    # Get new fresh model\n    model = get_model()\n\n    # Sanity Check\n    model.summary()\n\n    # Actual Training\n    history = model.fit(\n            x=train_dataset,\n            steps_per_epoch=TRAIN_STEPS_PER_EPOCH,\n            epochs=N_EPOCHS,\n            # Only used for validation data since training data is a generator\n            validation_data=val_dataset if USE_VAL else None,\n            validation_steps=N_VAL_STEPS_PER_EPOCH if USE_VAL else None,\n            callbacks=[\n                lr_callback,\n                WeightDecayCallback(),\n                ealystop_callback\n            ],\n            verbose = VERBOSE,\n        )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load Weights\nif LOAD_WEIGHTS:\n    model.load_weights('/kaggle/input/aslfr-training-python37/model.h5')\n    print(f'Successfully Loaded Pretrained Weights')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save Model Weights\nmodel.save_weights('model.h5')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Verify Model is Loaded Correctly\nmodel.evaluate(\n    val_dataset if USE_VAL else train_dataset,\n    steps=N_VAL_STEPS_PER_EPOCH if USE_VAL else TRAIN_STEPS_PER_EPOCH,\n    batch_size=BATCH_SIZE,\n    verbose=VERBOSE,\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Output Predictions to string\ndef outputs2phrase(outputs):\n    if outputs.ndim == 2:\n        outputs = np.argmax(outputs, axis=1)\n    \n    return ''.join([ORD2CHAR.get(s, '') for s in outputs])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The code defines a TensorFlow function `predict_phrase()` for generating a phrase prediction based on input frames. Here's a breakdown of the function:\n\n1. `@tf.function()`: Decorator that converts the Python function into a TensorFlow graph function for improved performance.\n\n2. `frames = tf.expand_dims(frames, axis=0)`: Adds a batch dimension to the input frames.\n\n3. `phrase = tf.fill([1,MAX_PHRASE_LENGTH], PAD_TOKEN)`: Initializes the phrase tensor with PAD_TOKEN.\n\n4. The function then iterates over the `MAX_PHRASE_LENGTH` to predict each token in the phrase:\n\n   a. `phrase = tf.cast(phrase, tf.int8)`: Casts the phrase tensor to `int8`.\n   \n   b. `outputs = model({'frames': frames, 'phrase': phrase})`: Calls the trained model with the input frames and current phrase to predict the next token.\n   \n   c. `phrase = tf.cast(phrase, tf.int32)`: Casts the phrase tensor back to `int32`.\n   \n   d. `phrase = tf.where(...)`: Updates the phrase tensor by replacing the tokens at positions less than `idx + 1` with the predicted tokens from the model outputs.\n   \n5. `outputs = tf.squeeze(phrase, axis=0)`: Removes the batch dimension from the output phrase tensor.\n","metadata":{}},{"cell_type":"code","source":"@tf.function()\ndef predict_phrase(frames):\n    # Add Batch Dimension\n    frames = tf.expand_dims(frames, axis=0)\n    # Start Phrase\n    phrase = tf.fill([1,MAX_PHRASE_LENGTH], PAD_TOKEN)\n\n    for idx in tf.range(MAX_PHRASE_LENGTH):\n        # Cast phrase to int8\n        phrase = tf.cast(phrase, tf.int8)\n        # Predict Next Token\n        outputs = model({\n            'frames': frames,\n            'phrase': phrase,\n        })\n\n        # Add predicted token to input phrase\n        phrase = tf.cast(phrase, tf.int32)\n        phrase = tf.where(\n            tf.range(MAX_PHRASE_LENGTH) < idx + 1,\n            tf.argmax(outputs, axis=2, output_type=tf.int32),\n            phrase,\n        )\n\n    # Squeeze outputs\n    outputs = tf.squeeze(phrase, axis=0)\n    outputs = tf.one_hot(outputs, N_UNIQUE_CHARACTERS)\n\n    # Return a dictionary with the output tensor\n    return outputs","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The function `get_ld_train()` computes the Levenshtein distances between the predicted phrases and the true phrases for the training data. Here's a breakdown of the function:\n\n1. `N = 100 if IS_INTERACTIVE else 1000`: Sets the number of samples to compute the Levenshtein distances. In interactive mode, it is set to 100, otherwise 1000.\n\n2. `LD_TRAIN = []`: Initializes an empty list to store the Levenshtein distances.\n\n3. The function iterates over the training data and performs the following steps for each sample:\n\n   a. `phrase_pred = predict_phrase(frames).numpy()`: Calls the `predict_phrase()` function to predict the phrase based on the input frames and converts the output to a NumPy array.\n   \n   b. `phrase_pred = outputs2phrase(phrase_pred)`: Converts the predicted phrase from the output tensor to a string representation.\n   \n   c. `phrase_true = outputs2phrase(phrase_true)`: Converts the true phrase from the output tensor to a string representation.\n   \n   d. `LD_TRAIN.append(...)`: Appends a dictionary containing the true phrase, predicted phrase, and the Levenshtein distance between them to the `LD_TRAIN` list.\n   \n   e. The loop breaks if the number of iterations reaches the specified limit (`N`).\n   \n4. `LD_TRAIN_DF = pd.DataFrame(LD_TRAIN)`: Converts the `LD_TRAIN` list to a pandas DataFrame for easier analysis and manipulation.\n\n5. The function returns the DataFrame `LD_TRAIN_DF` containing the Levenshtein distances for the training data.\n\nNote: The code uses the `tqdm` function to display a progress bar during the iteration. If `IS_INTERACTIVE` is set to `True`, it limits the number of iterations to `N` for faster execution.","metadata":{}},{"cell_type":"code","source":"# Compute Levenstein Distances\ndef get_ld_train():\n    N = 100 if IS_INTERACTIVE else 1000\n    LD_TRAIN = []\n    for idx, (frames, phrase_true) in enumerate(zip(tqdm(X_train, total=N), y_train)):\n        # Predict Phrase and Convert to String\n        phrase_pred = predict_phrase(frames).numpy()\n        phrase_pred = outputs2phrase(phrase_pred)\n        # True Phrase Ordinal to String\n        phrase_true = outputs2phrase(phrase_true)\n        # Add Levenstein Distance\n        LD_TRAIN.append({\n            'phrase_true': phrase_true,\n            'phrase_pred': phrase_pred,\n            'levenshtein_distance': levenshtein(phrase_pred, phrase_true),\n        })\n        # Take subset in interactive mode\n        if idx == N:\n            break\n            \n    # Convert to DataFrame\n    LD_TRAIN_DF = pd.DataFrame(LD_TRAIN)\n    \n    return LD_TRAIN_DF","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"LD_TRAIN_DF = get_ld_train()\n\n# Display Errors\ndisplay(LD_TRAIN_DF.head(30))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Value Counts\nLD_TRAIN_VC = dict([(i, 0) for i in range(LD_TRAIN_DF['levenshtein_distance'].max()+1)])\nfor ld in LD_TRAIN_DF['levenshtein_distance']:\n    LD_TRAIN_VC[ld] += 1\n\nplt.figure(figsize=(15,8))\npd.Series(LD_TRAIN_VC).plot(kind='bar', width=1)\nplt.title(f'Train Levenstein Distance Distribution | Mean: {LD_TRAIN_DF.levenshtein_distance.mean():.4f}')\nplt.xlabel('Levenstein Distance')\nplt.ylabel('Sample Count')\nplt.xlim(-0.50, LD_TRAIN_DF.levenshtein_distance.max()+0.50)\nplt.grid(axis='y')\nplt.savefig('temp.png')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The function `get_ld_val()` computes the Levenshtein distances between the predicted phrases and the true phrases for the validation data. Here's a breakdown of the function:\n\n1. `N = 100 if IS_INTERACTIVE else 1000`: Sets the number of samples to compute the Levenshtein distances. In interactive mode, it is set to 100, otherwise 1000.\n\n2. `LD_VAL = []`: Initializes an empty list to store the Levenshtein distances.\n\n3. The function iterates over the validation data and performs the following steps for each sample:\n\n   a. `phrase_pred = predict_phrase(frames).numpy()`: Calls the `predict_phrase()` function to predict the phrase based on the input frames and converts the output to a NumPy array.\n   \n   b. `phrase_pred = outputs2phrase(phrase_pred)`: Converts the predicted phrase from the output tensor to a string representation.\n   \n   c. `phrase_true = outputs2phrase(phrase_true)`: Converts the true phrase from the output tensor to a string representation.\n   \n   d. `LD_VAL.append(...)`: Appends a dictionary containing the true phrase, predicted phrase, and the Levenshtein distance between them to the `LD_VAL` list.\n   \n   e. The loop breaks if the number of iterations reaches the specified limit (`N`).\n   \n4. `LD_VAL_DF = pd.DataFrame(LD_VAL)`: Converts the `LD_VAL` list to a pandas DataFrame for easier analysis and manipulation.\n\n5. The function returns the DataFrame `LD_VAL_DF` containing the Levenshtein distances for the validation data.\n\nNote: The code uses the `tqdm` function to display a progress bar during the iteration. If `IS_INTERACTIVE` is set to `True`, it limits the number of iterations to `N` for faster execution.","metadata":{}},{"cell_type":"code","source":"# Compute Levenstein Distances\ndef get_ld_val():\n    N = 100 if IS_INTERACTIVE else 1000\n    LD_VAL = []\n    for idx, (frames, phrase_true) in enumerate(zip(tqdm(X_val, total=N), y_val)):\n        # Predict Phrase and Convert to String\n        phrase_pred = predict_phrase(frames).numpy()\n        phrase_pred = outputs2phrase(phrase_pred)\n        # True Phrase Ordinal to String\n        phrase_true = outputs2phrase(phrase_true)\n        # Add Levenstein Distance\n        LD_VAL.append({\n            'phrase_true': phrase_true,\n            'phrase_pred': phrase_pred,\n            'levenshtein_distance': levenshtein(phrase_pred, phrase_true),\n        })\n        # Take subset in interactive mode\n        if idx == N:\n            break\n            \n    # Convert to DataFrame\n    LD_VAL_DF = pd.DataFrame(LD_VAL)\n    \n    return LD_VAL_DF","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if USE_VAL:\n    LD_VAL_DF = get_ld_val()\n\n    # Display Errors\n    display(LD_VAL_DF.head(30))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Value Counts\nif USE_VAL:\n    LD_VAL_VC = dict([(i, 0) for i in range(LD_VAL_DF['levenshtein_distance'].max()+1)])\n    for ld in LD_VAL_DF['levenshtein_distance']:\n        LD_VAL_VC[ld] += 1\n\n    plt.figure(figsize=(15,8))\n    pd.Series(LD_VAL_VC).plot(kind='bar', width=1)\n    plt.title(f'Validation Levenstein Distance Distribution | Mean: {LD_VAL_DF.levenshtein_distance.mean():.4f}')\n    plt.xlabel('Levenstein Distance')\n    plt.ylabel('Sample Count')\n    plt.xlim(0-0.50, LD_VAL_DF.levenshtein_distance.max()+0.50)\n    plt.grid(axis='y')\n    plt.savefig('temp.png')\n    plt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `plot_history_metric()` function is used to plot the training history of a specific metric during the training process. Here's a breakdown of the function:\n\n1. Check if the model is being trained (`TRAIN_MODEL`). If not, the function returns without plotting.\n\n2. Create a new figure with a specified figure size.\n\n3. Retrieve the values of the specified metric from the `history` object.\n\n4. Determine the number of epochs (`N_EPOCHS`) based on the length of the metric values.\n\n5. Check if the validation metric values are available (`val`).\n\n6. If validation metric values are available, retrieve them from the `history` object and determine the index of the best value using the `f_best` function (default is `np.argmax`). Plot the validation metric values.\n\n7. Plot the training metric values.\n\n8. Determine the index of the best training metric value using the `f_best` function (default is `np.argmax`).\n\n9. If validation metric values are available, plot a red marker at the best training metric value and a purple marker at the best validation metric value.\n\n10. Set the title, x-axis label, and y-axis label for the plot.\n\n11. If a y-axis limit (`ylim`) is specified, set the y-axis limit.\n\n12. If a y-axis scale (`yscale`) is specified, set the y-axis scale.\n\n13. If custom y-ticks (`yticks`) are specified, set the y-axis ticks.\n\n14. Set the x-ticks based on the number of epochs.\n\n15. Set the tick labels for the x-axis and y-axis.\n\n16. Add a legend to the plot.\n\n17. Display the grid and show the plot.\n\nThe function allows you to visualize the training history of a specific metric, including the best training and validation metric values.","metadata":{}},{"cell_type":"code","source":"def plot_history_metric(metric, f_best=np.argmax, ylim=None, yscale=None, yticks=None):\n    # Only plot when training\n    if not TRAIN_MODEL:\n        return\n    \n    plt.figure(figsize=(20, 10))\n    \n    values = history.history[metric]\n    N_EPOCHS = len(values)\n    val = 'val' in ''.join(history.history.keys())\n    # Epoch Ticks\n    if N_EPOCHS <= 20:\n        x = np.arange(1, N_EPOCHS + 1)\n    else:\n        x = [1, 5] + [10 + 5 * idx for idx in range((N_EPOCHS - 10) // 5 + 1)]\n\n    x_ticks = np.arange(1, N_EPOCHS+1)\n\n    # Validation\n    if val:\n        val_values = history.history[f'val_{metric}']\n        val_argmin = f_best(val_values)\n        plt.plot(x_ticks, val_values, label=f'val')\n\n    # summarize history for accuracy\n    plt.plot(x_ticks, values, label=f'train')\n    argmin = f_best(values)\n    plt.scatter(argmin + 1, values[argmin], color='red', s=75, marker='o', label=f'train_best')\n    if val:\n        plt.scatter(val_argmin + 1, val_values[val_argmin], color='purple', s=75, marker='o', label=f'val_best')\n\n    plt.title(f'Model {metric}', fontsize=24, pad=10)\n    plt.ylabel(metric, fontsize=20, labelpad=10)\n\n    if ylim:\n        plt.ylim(ylim)\n\n    if yscale is not None:\n        plt.yscale(yscale)\n        \n    if yticks is not None:\n        plt.yticks(yticks, fontsize=16)\n\n    plt.xlabel('epoch', fontsize=20, labelpad=10)        \n    plt.tick_params(axis='x', labelsize=8)\n    plt.xticks(x, fontsize=16) # set tick step to 1 and let x axis start at 1\n    plt.yticks(fontsize=16)\n    \n    plt.legend(prop={'size': 10})\n    plt.grid()\n    plt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_history_metric('loss', f_best=np.argmin)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_history_metric('top1acc', ylim=[0,1], yticks=np.arange(0.0, 1.1, 0.1))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_history_metric('top5acc', ylim=[0,1], yticks=np.arange(0.0, 1.1, 0.1))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Model Layer Names\nfor l in model.layers:\n    print(l.name)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `TFLiteModel` class is a TensorFlow module that serves as a wrapper for the TF Lite model. Here's an overview of the class:\n\n- The `TFLiteModel` class inherits from `tf.Module` and contains the necessary methods to define the TF Lite model.\n- The constructor takes the `model` as an input, which is the trained model.\n- The class includes two `tf.function` decorated methods, `encoder` and `decoder`, that perform the encoding and decoding steps of the model, respectively.\n- The `__call__` method is decorated with `tf.function` and `input_signature` to specify the input signature for the TF Lite model.\n- The `__call__` method preprocesses the input frames using the `preprocess_layer`, passes them through the encoder, and then generates predictions using the decoder.\n- The predictions are generated one token at a time until the end-of-sentence (EOS) token is predicted.\n- The final outputs are converted to one-hot encoded tensors and returned as a dictionary with the key 'outputs'.\n\nAdditionally, the code includes a section to instantiate the `TFLiteModel` with the trained model and perform a sanity check on the TF Lite model using a sample input. The input frames are passed through the TF Lite model, and the predicted phrase is compared with the true phrase to verify the functionality of the TF Lite model.","metadata":{}},{"cell_type":"code","source":"# TFLite model for submission\nclass TFLiteModel(tf.Module):\n    def __init__(self, model):\n        super(TFLiteModel, self).__init__()\n\n        # Load the feature generation and main models\n        self.preprocess_layer = preprocess_layer\n        self.model = model\n    \n    @tf.function(jit_compile=True)\n    def encoder(self, x, frames_inp):\n        x = self.model.get_layer('embedding')(x, frames_inp)\n        x = self.model.get_layer('encoder')(x, frames_inp)\n        \n        return x\n        \n    @tf.function(jit_compile=True)\n    def decoder(self, x, phrase_inp):\n        x = self.model.get_layer('decoder')(x, phrase_inp)\n        x = self.model.get_layer('classifier')(x)\n        \n        return x\n    \n    @tf.function(input_signature=[tf.TensorSpec(shape=[None, N_COLS0], dtype=tf.float32, name='inputs')])\n    def __call__(self, inputs):\n        # Number Of Input Frames\n        N_INPUT_FRAMES = tf.shape(inputs)[0]\n        # Preprocess Data\n        frames_inp = self.preprocess_layer(inputs)        \n        # Add Batch Dimension\n        frames_inp = tf.expand_dims(frames_inp, axis=0)\n        # Get Encoding\n        encoding = self.encoder(frames_inp, frames_inp)\n        # Make Prediction\n        phrase = tf.fill([1,MAX_PHRASE_LENGTH], PAD_TOKEN)\n        # Predict One Token At A Time\n        stop = False\n        for idx in tf.range(MAX_PHRASE_LENGTH):\n            # Cast phrase to int8\n            phrase = tf.cast(phrase, tf.int8)\n            # If EOS token is predicted, stop predicting\n            outputs = tf.cond(\n                stop,\n                lambda: tf.one_hot(tf.cast(phrase, tf.int32), N_UNIQUE_CHARACTERS),\n                lambda: self.decoder(encoding, phrase)\n            )\n            # Add predicted token to input phrase\n            phrase = tf.cast(phrase, tf.int32)\n            # Replcae PAD token with predicted token up to idx\n            phrase = tf.where(\n                tf.range(MAX_PHRASE_LENGTH) < idx + 1,\n                tf.argmax(outputs, axis=2, output_type=tf.int32),\n                phrase,\n            )\n            # Predicted Token\n            predicted_token = phrase[0,idx]\n            # If EOS (End Of Sentence) token is predicted stop\n            if not stop:\n                stop = predicted_token == EOS_TOKEN\n            \n        # Squeeze outputs\n        outputs = tf.squeeze(phrase, axis=0)\n        outputs = tf.one_hot(outputs, N_UNIQUE_CHARACTERS)\n            \n        # Return a dictionary with the output tensor\n        return {'outputs': outputs }\n\n# Define TF Lite Model\ntflite_keras_model = TFLiteModel(model)\n\n# Sanity Check\n# demo_sequence_id = 1816796431\ndemo_sequence_id = example_parquet_df.index.unique()[0]\ndemo_raw_data = example_parquet_df.loc[demo_sequence_id, COLUMNS0].values\ndemo_phrase_true = train_sequence_id.loc[demo_sequence_id, 'phrase']\nprint(f'demo_raw_data shape: {demo_raw_data.shape}, dtype: {demo_raw_data.dtype}')\ndemo_output = tflite_keras_model(demo_raw_data)['outputs'].numpy()\nprint(f'demo_output shape: {demo_output.shape}, dtype: {demo_output.dtype}')\nprint(f'demo_outputs phrase decoded: {outputs2phrase(demo_output)}')\nprint(f'phrase true: {demo_phrase_true}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create Model Converter\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)\n# Convert Model\ntflite_model = keras_model_converter.convert()\n# Write Model\nwith open('/kaggle/working/model.tflite', 'wb') as f:\n    f.write(tflite_model)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add selected_columns json to only select specific columns from input frames\nwith open('inference_args.json', 'w') as f:\n     json.dump({ 'selected_columns': COLUMNS0.tolist() }, f)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Zip Model\n!zip submission.zip /kaggle/working/model.tflite /kaggle/working/inference_args.json","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!rm -f inference_args.json model.h5 model.png model.tflite temp.png","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}