{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Hello Fellow Kagglers,\n\nThis competition has been challenging to say the least, especially the submission.\n\nAfter a few weeks I finally got a working training + inference pipeline which is shared through this notebook.\n\nThe model consists of a transformer embedding + encoder + decoder.\n\nInference is performed by starting with an Start of Sentence (SOS) token and predicting one character at a time using the previous prediction.\n\nFeel free to ask for clarafications or comment.\n\nNotebook will be updated periodically.\n\n[Preprocessing Notebook](https://www.kaggle.com/code/m4nugnzl/aslfr-eda-preprocessing-dataset-for-beginners)\n\n**V6**\n\nThis competition has an inference limit of 5 hours which requires careful allocation of computational resources in the model. Most changes are based on the assymetrical number of encoder/decoder calls during inference.\n\nInference requires the encoder to encode the input frames and subsequently use that encoding to predict the 1st character by inputting the encoding and Start of Sentence token. Next, the encoding, SOS token and 1st predicted token are used to predict the 2nd character. Inference thus requires 1 call to the encoder and multiple calls to the encoder. On average a phrase is 18 characters long, requiring 18+1(SOS token) calls to the decoder. To stay within the 5 hour inference limit the encoder can be computationally heavy, however the decoder should be light.\n\nSome inspiration is taken from the [1st place solution - training](https://www.kaggle.com/code/hoyso48/1st-place-solution-training) from the last [Google - Isolated Sign Language Recognition\n](https://www.kaggle.com/competitions/asl-signs) competition.\n\n* Increased training epochs 30 -> 100\n* Using all data for training, no validation set\n* Increased number of decoder blocks 2 -> 3\n* Increased encoder dimensions 256 -> 384\n* Halved attention dimension to decrease computational intensity of Multi Head Attention\n* Added 20% dropout to multi head attention output\n* Batch size 128 -> 64\n* Classification layer linear activation for logits in loss function\n\n**V7**\n\nA small update, since several other notebooks got released with a LB score as high as 0.689!\n\nThis will most likely be the last version of this notebook, it seems like the way to get to 0.70 is CTC loss.\n\nThe 5 hour inference limit could is limitation with this encoder/decoder architecture.\n\nPerformance could be improved by using FlashAttention, however this is not yet implemented in Tensorflow.\n\nAnother observation is the huge speedup from jit_compile, which allows for inferencing up to ~30 samples per second. The TFLite model obtains about ~3 samples/s and it seems the `jit_compile` flag does not impact the TFLite model speed.\n\nModifications in this version\n\n* More efficient transformer architecture based on 1st place solution - training by HOYSO48\n* NUM_BLOCKS_ENCODER = 3 → 4\n* Correct attention mask in decoder: causal → non empty frames\n\n**Helpful Tutorials**\n\n[English-to-Spanish translation with a sequence-to-sequence Transformer](https://keras.io/examples/nlp/neural_machine_translation_with_transformer/)\n\n[Lecture 12.1 Self-attention](https://www.youtube.com/watch?v=KmAISyVvE1Y&list=LL&index=3)\n\n\n> This notebook is a fully explained version of [ASLFR - Transformer Training + Inference](https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference/notebook)","metadata":{"papermill":{"duration":0.023581,"end_time":"2023-07-02T11:32:59.503233","exception":false,"start_time":"2023-07-02T11:32:59.479652","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport matplotlib as mpl\nimport seaborn as sn\nimport tensorflow as tf\nimport tensorflow_addons as tfa\n\nfrom tqdm.notebook import tqdm\nfrom sklearn.model_selection import train_test_split, GroupShuffleSplit\nfrom leven import levenshtein\n\nimport glob\nimport sys\nimport os\nimport math\nimport gc\nimport sys\nimport sklearn\nimport time\nimport json\n\n# TQDM Progress Bar With Pandas Apply Function\ntqdm.pandas()\n\nprint(f'Tensorflow Version {tf.__version__}')\nprint(f'Python Version: {sys.version}')","metadata":{"papermill":{"duration":9.205219,"end_time":"2023-07-02T11:33:08.731341","exception":false,"start_time":"2023-07-02T11:32:59.526122","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:08:50.478956Z","iopub.execute_input":"2023-07-31T07:08:50.479315Z","iopub.status.idle":"2023-07-31T07:08:50.488077Z","shell.execute_reply.started":"2023-07-31T07:08:50.479287Z","shell.execute_reply":"2023-07-31T07:08:50.486932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Character 2 Ordinal Encoding","metadata":{"papermill":{"duration":0.022966,"end_time":"2023-07-02T11:33:08.779005","exception":false,"start_time":"2023-07-02T11:33:08.756039","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Read Character to Ordinal Encoding Mapping\nwith open('/kaggle/input/asl-fingerspelling/character_to_prediction_index.json') as json_file:\n    CHAR2ORD = json.load(json_file)\n    \n# Ordinal to Character Mapping\nORD2CHAR = {j:i for i,j in CHAR2ORD.items()}\n\n# convert dictionary to pandas dataframe\nCHAR2ORD_DF = pd.DataFrame(CHAR2ORD.values(),index=CHAR2ORD.keys(),columns=['Ordinal Encoding'])\nCHAR2ORD_DF.head()","metadata":{"papermill":{"duration":0.057156,"end_time":"2023-07-02T11:33:08.859134","exception":false,"start_time":"2023-07-02T11:33:08.801978","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:08:51.224233Z","iopub.execute_input":"2023-07-31T07:08:51.224603Z","iopub.status.idle":"2023-07-31T07:08:51.240618Z","shell.execute_reply.started":"2023-07-31T07:08:51.224573Z","shell.execute_reply":"2023-07-31T07:08:51.239337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Hyperparameters\n\nWe define the hyperparameters that will be used later:\n\n- `IS_INTERACTIVE`: This variable indicates whether the notebook is running in interactive mode, which means it is being executed in an environment where changes and experiments can be performed interactively. In this mode, you can run code cells individually, modify the code, and see the results immediately. This is useful when developing and debugging code as it allows you to make quick changes and observe their effects in real-time. **If you want to execute the complete code (5 hours), run it without interactive mode (save version or submit)**.\n\n- `SEED`: This is a global random seed used to reproduce the results. By setting a seed, it ensures that the results are consistent across different runs.\n\n- `N_TARGET_FRAMES`: This variable indicates the number of frames to which the recordings will be resized. This is done to **standardize the length of input sequences**.\n\n- `DEBUG`: This variable indicates whether it is running in debug mode. If `True`, it will run in debug mode, which may involve running a subset of the data for faster execution and easier debugging.\n\n- `N_UNIQUE_CHARACTERS0`: This variable represents the number of **unique characters that the model should predict**. It does not include the padding token, the start of sentence token, and the end of sentence token.\n\n- `N_UNIQUE_CHARACTERS`: This variable is similar to `N_UNIQUE_CHARACTERS0`, but includes the padding token, the start of sentence token, and the end of sentence token. Therefore, it represents the total number of classes the model should predict.\n\n- `PAD_TOKEN`, `START_TOKEN`, and `END_TOKEN`: These variables represent the indices of the padding, start of sentence, and end of sentence tokens in the vocabulary. They are used during data processing and target text sequence generation.\n\n- `USE_VAL`: This variable controls whether to use 10% of the data for validation during training.\n\n- `BATCH_SIZE`: This variable determines the batch size for training.\n\n- `N_EPOCHS`: This variable sets the number of epochs for training the model. If `IS_INTERACTIVE` is True, it is set to 2 for quick interactive training; otherwise, it is set to 100 for normal training.\n\n- `N_WARMUP_EPOCHS`: This variable defines the number of warm-up epochs in the learning rate scheduler.\n\n- `LR_MAX`: This variable sets the maximum learning rate for the training process.\n\n- `WD_RATIO`: This variable represents the weight decay ratio based on the learning rate.\n\n- `MAX_PHRASE_LENGTH`: This variable defines the maximum length of a phrase, including the end of sentence token.\n\n- `TRAIN_MODEL`: This variable controls whether to train the model.\n\n- `LOAD_WEIGHTS`: This variable controls whether to load pre-trained weights for the model.\n\n- `WARMUP_METHOD`: This variable specifies the warm-up method for the learning rate scheduler, either 'log' or 'exp'. If 'log', the learning rate gradually increases in a logarithmic manner during the initial warm-up epochs. This means that it increases slowly at first and then accelerates as the warm-up epochs progress. It is useful when a smoother and more gradual warm-up is desired. If 'exp', the learning rate increases exponentially during the initial warm-up epochs. This implies a faster increase at the beginning and a higher learning rate in the initial epochs. It is useful when a faster and more aggressive warm-up is desired.\n","metadata":{"papermill":{"duration":0.022903,"end_time":"2023-07-02T11:33:08.905507","exception":false,"start_time":"2023-07-02T11:33:08.882604","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"As we currently have it, the notebook runs in interactive mode, with verbosity 1, without a training subset or a validation set (we will use everything for training). We will use a configuration that allows for fast execution, with a number of epochs set to 2.","metadata":{}},{"cell_type":"code","source":"# If Notebook Is Run By Committing or In Interactive Mode For Development\nIS_INTERACTIVE = os.environ['KAGGLE_KERNEL_RUN_TYPE'] == 'Interactive'\n#IS_INTERACTIVE = False # to run full code\n# Verbose Setting during training\nVERBOSE = 1 if IS_INTERACTIVE else 2\n# Global Random Seed\nSEED = 42\n# Number of Frames to resize recording to\nN_TARGET_FRAMES = 128\n# Global debug flag, takes subset of train\nDEBUG = False\n# Number of Unique Characters To Predict + Pad Token + SOS Token + EOS Token\nN_UNIQUE_CHARACTERS0 = len(CHAR2ORD)\nN_UNIQUE_CHARACTERS = len(CHAR2ORD) + 1 + 1 + 1\nPAD_TOKEN = len(CHAR2ORD) # Padding # 59\nSTART_TOKEN = len(CHAR2ORD) + 1 # Start Of Sentence # 60\nEND_TOKEN = len(CHAR2ORD) + 2 # End Of Sentence # 61\n# Whether to use 10% of data for validation\nUSE_VAL = False\n# Batch Size\nBATCH_SIZE = 64\n# Number of Epochs to Train for\nN_EPOCHS = 2 if IS_INTERACTIVE else 100\n# Number of Warmup Epochs in Learning Rate Scheduler\nN_WARMUP_EPOCHS = 10\n# Maximum Learning Rate\nLR_MAX = 1e-3\n# Weight Decay Ratio as Ratio of Learning Rate\nWD_RATIO = 0.05\n# Length of Phrase + EOS Token\nMAX_PHRASE_LENGTH = 31 + 1\n# Whether to Train The model\nTRAIN_MODEL = True\n# Whether to Load Pretrained Weights\nLOAD_WEIGHTS = False\n# Learning Rate Warmup Method [log,exp]\nWARMUP_METHOD = 'exp'","metadata":{"papermill":{"duration":0.034637,"end_time":"2023-07-02T11:33:08.963209","exception":false,"start_time":"2023-07-02T11:33:08.928572","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:08:51.242666Z","iopub.execute_input":"2023-07-31T07:08:51.243707Z","iopub.status.idle":"2023-07-31T07:08:51.252508Z","shell.execute_reply.started":"2023-07-31T07:08:51.24367Z","shell.execute_reply":"2023-07-31T07:08:51.251499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read an Prepare the Data\n\n# Read Train\n\nSince `DEBUG=False`, we load the entire training dataset. We add a new column file_path, containing the complete `file_path` for the parquet files.","metadata":{"papermill":{"duration":0.023106,"end_time":"2023-07-02T11:33:09.114637","exception":false,"start_time":"2023-07-02T11:33:09.091531","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Read Train DataFrame\nif DEBUG:\n    train = pd.read_csv('/kaggle/input/asl-fingerspelling/train.csv').head(5000)\nelse:\n    train = pd.read_csv('/kaggle/input/asl-fingerspelling/train.csv')\n\n# this will be used to construct TFLite model\ntrain_sequence_id = train.set_index('sequence_id')\n\n# Get complete file path to file\ndef get_file_path(path):\n    return f'/kaggle/input/asl-fingerspelling/{path}'\n\ntrain['file_path'] = train['path'].apply(get_file_path)\n\ntrain.head()","metadata":{"papermill":{"duration":0.236261,"end_time":"2023-07-02T11:33:09.374134","exception":false,"start_time":"2023-07-02T11:33:09.137873","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:08:52.022675Z","iopub.execute_input":"2023-07-31T07:08:52.0231Z","iopub.status.idle":"2023-07-31T07:08:52.206677Z","shell.execute_reply.started":"2023-07-31T07:08:52.023072Z","shell.execute_reply":"2023-07-31T07:08:52.205655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Example File Paths\n\nLoading the 10 parquet elements from `train_landmark_subsets`","metadata":{"papermill":{"duration":0.023756,"end_time":"2023-07-02T11:33:09.552626","exception":false,"start_time":"2023-07-02T11:33:09.52887","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Unique Parquet Files\nINFERENCE_FILE_PATHS = pd.Series(\n        glob.glob('/kaggle/input//aslfr-eda-preprocessing-dataset-for-beginners/train_landmark_subsets/*')\n    )\n\nprint(f'Found {len(INFERENCE_FILE_PATHS)} Inference Pickle Files')","metadata":{"papermill":{"duration":0.038118,"end_time":"2023-07-02T11:33:09.614891","exception":false,"start_time":"2023-07-02T11:33:09.576773","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:08:52.209477Z","iopub.execute_input":"2023-07-31T07:08:52.210268Z","iopub.status.idle":"2023-07-31T07:08:52.217785Z","shell.execute_reply.started":"2023-07-31T07:08:52.210228Z","shell.execute_reply":"2023-07-31T07:08:52.216406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Example Parquet\n\nLet's remember the format of a Parquet file that we have saved before preprocessing:","metadata":{}},{"cell_type":"code","source":"# Read First Parquet File\n# example_parquet_df = pd.read_parquet(train['file_path'][0])\nexample_parquet_df = pd.read_parquet(INFERENCE_FILE_PATHS[0])\n\n# Each parquet file contains 1000 recordings\nprint(f'Number of Unique Recording: {example_parquet_df.index.nunique()}')\n# Display DataFrame layout\ndisplay(example_parquet_df.head())","metadata":{"papermill":{"duration":1.512276,"end_time":"2023-07-02T11:33:49.075196","exception":false,"start_time":"2023-07-02T11:33:47.56292","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:08:52.811468Z","iopub.execute_input":"2023-07-31T07:08:52.812398Z","iopub.status.idle":"2023-07-31T07:08:54.030296Z","shell.execute_reply.started":"2023-07-31T07:08:52.812336Z","shell.execute_reply":"2023-07-31T07:08:54.029178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load X/y\n\nSimilarly, we load the training and validation files. In our case, `USE_VAL=False`, so we load the complete `X` and `y`. The original shape of the `y` files is `[61955, 128]`, due to the padding we have applied. Since we have chosen the phrase limit to be `31+1` spaces, with the last space as the end-of-sentence token, we restrict the size from 128 to 32.\n","metadata":{"papermill":{"duration":0.023663,"end_time":"2023-07-02T11:33:09.662914","exception":false,"start_time":"2023-07-02T11:33:09.639251","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Train/Validation\nif USE_VAL:\n    # TRAIN\n    X_train = np.load('/kaggle/input/aslfr-eda-preprocessing-dataset-for-beginners/X_train.npy')\n    y_train = np.load('/kaggle/input/aslfr-eda-preprocessing-dataset-for-beginners/y_train.npy')[:,:MAX_PHRASE_LENGTH]\n    N_TRAIN_SAMPLES = len(X_train)\n    # VAL\n    X_val = np.load('/kaggle/input/aslfr-eda-preprocessing-dataset-for-beginners/X_val.npy')\n    y_val = np.load('/kaggle/input/aslfr-eda-preprocessing-dataset-for-beginners/y_val.npy')[:,:MAX_PHRASE_LENGTH]\n    N_VAL_SAMPLES = len(X_val)\n    # Shapes\n    print(f'X_train shape: {X_train.shape}, X_val shape: {X_val.shape}')\n    print(f'y_train shape: {X_train.shape}, y_val shape: {X_val.shape}')\n    \n# Train On All Data\nelse:\n    # TRAIN\n    X_train = np.load('/kaggle/input/aslfr-eda-preprocessing-dataset-for-beginners/X.npy')\n    y_train = np.load('/kaggle/input/aslfr-eda-preprocessing-dataset-for-beginners/y.npy')[:,:MAX_PHRASE_LENGTH]\n    N_TRAIN_SAMPLES = len(X_train)\n    print(f'X_train shape: {X_train.shape}')\n    print(f'y_train shape: {y_train.shape}')","metadata":{"papermill":{"duration":37.639487,"end_time":"2023-07-02T11:33:47.326164","exception":false,"start_time":"2023-07-02T11:33:09.686677","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:08:54.033453Z","iopub.execute_input":"2023-07-31T07:08:54.033868Z","iopub.status.idle":"2023-07-31T07:08:58.065814Z","shell.execute_reply.started":"2023-07-31T07:08:54.033837Z","shell.execute_reply":"2023-07-31T07:08:58.064798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Example Batch\n\nNow, we create batches of examples to debug or test the code. We create two sets of examples: one with a larger sample size (`N_EXAMPLE_BATCH_SAMPLES`) and another with a smaller sample size (`N_EXAMPLE_BATCH_SAMPLES_SMALL`).\n\nIn each set of examples, we create a dictionary `X_batch` containing the input features, and an array `y_batch` containing the output labels.\n\nIn the `X_batch` dictionary, we copy the first `N_EXAMPLE_BATCH_SAMPLES` training samples from `X_train`. The `frames` represent the data of the frames (rows in this case, where each row is a set of `[1, 128, 164]`), and `phrase` represents the phrases associated with the frames.\n\nIn the `y_batch` array, we copy the first `N_EXAMPLE_BATCH_SAMPLES` training labels from `y_train`, which correspond to the phrases associated with the frames.\n\nWe perform the same procedure for the smaller set, using the variables `X_batch_small` and `y_batch_small`.","metadata":{"papermill":{"duration":0.02373,"end_time":"2023-07-02T11:33:47.374154","exception":false,"start_time":"2023-07-02T11:33:47.350424","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Example Batch For Debugging\nN_EXAMPLE_BATCH_SAMPLES = 1024\nN_EXAMPLE_BATCH_SAMPLES_SMALL = 32\n# Example Batch\nX_batch = {\n    'frames': np.copy(X_train[:N_EXAMPLE_BATCH_SAMPLES]),\n    'phrase': np.copy(y_train[:N_EXAMPLE_BATCH_SAMPLES]),\n#     'phrase_type': np.copy(y_phrase_type_train[:N_EXAMPLE_BATCH_SAMPLES]),\n}\ny_batch = np.copy(y_train[:N_EXAMPLE_BATCH_SAMPLES])\n# Small Example Batch\nX_batch_small = {\n    'frames': np.copy(X_train[:N_EXAMPLE_BATCH_SAMPLES_SMALL]),\n    'phrase': np.copy(y_train[:N_EXAMPLE_BATCH_SAMPLES_SMALL]),\n#     'phrase_type': np.copy(y_phrase_type_train[:N_EXAMPLE_BATCH_SAMPLES_SMALL]),\n}\ny_batch_small = np.copy(y_train[:N_EXAMPLE_BATCH_SAMPLES_SMALL])","metadata":{"papermill":{"duration":0.093503,"end_time":"2023-07-02T11:33:47.491607","exception":false,"start_time":"2023-07-02T11:33:47.398104","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:08:58.067442Z","iopub.execute_input":"2023-07-31T07:08:58.067839Z","iopub.status.idle":"2023-07-31T07:08:58.204711Z","shell.execute_reply.started":"2023-07-31T07:08:58.067804Z","shell.execute_reply":"2023-07-31T07:08:58.203608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"First element of frames:\", X_batch_small['frames'][0])\nprint(f'Shape of every frames element: {X_batch_small[\"frames\"][0].shape} \\n\\n')\nprint(\"First element of phrase:\", X_batch_small['phrase'][0])\nprint(f'Shape of every phrase element: {X_batch_small[\"phrase\"][0].shape}')","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:08:58.207474Z","iopub.execute_input":"2023-07-31T07:08:58.207869Z","iopub.status.idle":"2023-07-31T07:08:58.216523Z","shell.execute_reply.started":"2023-07-31T07:08:58.207828Z","shell.execute_reply":"2023-07-31T07:08:58.215295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Landmark Indices\n\nSimilarly to the preprocessing notebook, we define the `get_idxs` function to obtain the indices of the left hand, right hand, and lips.","metadata":{"papermill":{"duration":0.039462,"end_time":"2023-07-02T11:33:49.150383","exception":false,"start_time":"2023-07-02T11:33:49.110921","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Get indices in original dataframe\ndef get_idxs(df, words_pos, words_neg=[], ret_names=True, idxs_pos=None):\n    idxs = []\n    names = []\n    for w in words_pos:\n        for col_idx, col in enumerate(example_parquet_df.columns):\n            # Exclude Non Landmark Columns\n            if col in ['frame']:\n                continue\n                \n            col_idx = int(col.split('_')[-1])\n            # Check if column name contains all words\n            if (w in col) and (idxs_pos is None or col_idx in idxs_pos) and all([w not in col for w in words_neg]):\n                idxs.append(col_idx)\n                names.append(col)\n    # Convert to Numpy arrays\n    idxs = np.array(idxs)\n    names = np.array(names)\n    # Returns either both column indices and names\n    if ret_names:\n        return idxs, names\n    # Or only columns indices\n    else:\n        return idxs","metadata":{"papermill":{"duration":0.051023,"end_time":"2023-07-02T11:33:49.235748","exception":false,"start_time":"2023-07-02T11:33:49.184725","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:08:58.218377Z","iopub.execute_input":"2023-07-31T07:08:58.218789Z","iopub.status.idle":"2023-07-31T07:08:58.229754Z","shell.execute_reply.started":"2023-07-31T07:08:58.218754Z","shell.execute_reply":"2023-07-31T07:08:58.228889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lips Landmark Face Ids\nLIPS_LANDMARK_IDXS = np.array([\n        61, 185, 40, 39, 37, 0, 267, 269, 270, 409,\n        291, 146, 91, 181, 84, 17, 314, 405, 321, 375,\n        78, 191, 80, 81, 82, 13, 312, 311, 310, 415,\n        95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n    ])\n\n# Landmark Indices for Left/Right hand without z axis in raw data\nLEFT_HAND_IDXS0, LEFT_HAND_NAMES0 = get_idxs(example_parquet_df, ['left_hand'], ['z'])\nRIGHT_HAND_IDXS0, RIGHT_HAND_NAMES0 = get_idxs(example_parquet_df, ['right_hand'], ['z'])\nLIPS_IDXS0, LIPS_NAMES0 = get_idxs(example_parquet_df, ['face'], ['z'], idxs_pos=LIPS_LANDMARK_IDXS)\nCOLUMNS0 = np.concatenate((LEFT_HAND_NAMES0, RIGHT_HAND_NAMES0, LIPS_NAMES0))\nN_COLS0 = len(COLUMNS0)\n# Only X/Y axes are used\nN_DIMS0 = 2\n\nprint(f'N_COLS0: {N_COLS0}')","metadata":{"papermill":{"duration":0.053701,"end_time":"2023-07-02T11:33:49.324986","exception":false,"start_time":"2023-07-02T11:33:49.271285","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:08:58.231465Z","iopub.execute_input":"2023-07-31T07:08:58.231995Z","iopub.status.idle":"2023-07-31T07:08:58.243425Z","shell.execute_reply.started":"2023-07-31T07:08:58.231962Z","shell.execute_reply":"2023-07-31T07:08:58.24234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Landmark Indices in subset of dataframe with only COLUMNS selected\nLEFT_HAND_IDXS = np.argwhere(np.isin(COLUMNS0, LEFT_HAND_NAMES0)).squeeze()\nRIGHT_HAND_IDXS = np.argwhere(np.isin(COLUMNS0, RIGHT_HAND_NAMES0)).squeeze()\nLIPS_IDXS = np.argwhere(np.isin(COLUMNS0, LIPS_NAMES0)).squeeze()\nHAND_IDXS = np.concatenate((LEFT_HAND_IDXS, RIGHT_HAND_IDXS), axis=0)\nN_COLS = N_COLS0\n# Only X/Y axes are used\nN_DIMS = 2\n\nprint(f'N_COLS: {N_COLS}')","metadata":{"papermill":{"duration":0.050026,"end_time":"2023-07-02T11:33:49.408684","exception":false,"start_time":"2023-07-02T11:33:49.358658","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:08:58.244943Z","iopub.execute_input":"2023-07-31T07:08:58.245364Z","iopub.status.idle":"2023-07-31T07:08:58.257534Z","shell.execute_reply.started":"2023-07-31T07:08:58.245324Z","shell.execute_reply":"2023-07-31T07:08:58.256296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Indices in processed data by axes with only dominant hand\nHAND_X_IDXS = np.array(\n        [idx for idx, name in enumerate(LEFT_HAND_NAMES0) if 'x' in name]\n    ).squeeze()\nHAND_Y_IDXS = np.array(\n        [idx for idx, name in enumerate(LEFT_HAND_NAMES0) if 'y' in name]\n    ).squeeze()\n# Names in processed data by axes\nHAND_X_NAMES = LEFT_HAND_NAMES0[HAND_X_IDXS]\nHAND_Y_NAMES = LEFT_HAND_NAMES0[HAND_Y_IDXS]\n\nprint(f'HAND_X_NAMES: {HAND_X_NAMES}')","metadata":{"papermill":{"duration":0.045648,"end_time":"2023-07-02T11:33:49.488024","exception":false,"start_time":"2023-07-02T11:33:49.442376","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:08:58.259174Z","iopub.execute_input":"2023-07-31T07:08:58.259616Z","iopub.status.idle":"2023-07-31T07:08:58.27122Z","shell.execute_reply.started":"2023-07-31T07:08:58.25954Z","shell.execute_reply":"2023-07-31T07:08:58.270142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Mean/STD Loading\n\nWe load the mean and standard deviation of the mean for those values that were different from 0.\n","metadata":{"papermill":{"duration":0.037446,"end_time":"2023-07-02T11:33:49.563172","exception":false,"start_time":"2023-07-02T11:33:49.525726","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Mean/Standard Deviations of data used for normalizing\nMEANS = np.load('/kaggle/input/aslfr-eda-preprocessing-dataset-for-beginners/MEANS.npy').reshape(-1)\nSTDS = np.load('/kaggle/input/aslfr-eda-preprocessing-dataset-for-beginners/STDS.npy').reshape(-1)\n\nprint(f'First 5 values of MEANS: {MEANS[:5]}')\nprint(f'Shape of MEANS: {MEANS.shape}') # 164 values, 1 for every column","metadata":{"papermill":{"duration":0.051188,"end_time":"2023-07-02T11:33:49.648816","exception":false,"start_time":"2023-07-02T11:33:49.597628","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:08:58.272685Z","iopub.execute_input":"2023-07-31T07:08:58.273172Z","iopub.status.idle":"2023-07-31T07:08:58.295746Z","shell.execute_reply.started":"2023-07-31T07:08:58.273139Z","shell.execute_reply":"2023-07-31T07:08:58.294598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's add the preprocessing layer again here, as we will use it in the `TFLite` model.","metadata":{}},{"cell_type":"code","source":"class PreprocessLayer(tf.keras.layers.Layer):\n    def __init__(self):\n        super(PreprocessLayer, self).__init__()\n        self.normalisation_correction = tf.constant(\n                    # Add 0.50 to x coordinates of left hand (original right hand) and substract 0.50 of right hand (original left hand)\n                     [0.50 if 'x' in name else 0.00 for name in LEFT_HAND_NAMES0],\n                dtype=tf.float32,\n            )\n    \n    @tf.function(\n        input_signature=(tf.TensorSpec(shape=[None,N_COLS0], dtype=tf.float32),),\n    )\n    def call(self, data0, resize=True):\n        # Fill NaN Values With 0\n        data = tf.where(tf.math.is_nan(data0), 0.0, data0)\n        \n        # Hacky\n        data = data[None]\n        \n        # Empty Hand Frame Filtering\n        hands = tf.slice(data, [0,0,0], [-1, -1, 84])\n        hands = tf.abs(hands)\n        mask = tf.reduce_sum(hands, axis=2)\n        mask = tf.not_equal(mask, 0)\n        data = data[mask][None]\n        \n        # Pad Zeros\n        N_FRAMES = len(data[0])\n        if N_FRAMES < N_TARGET_FRAMES:\n            data = tf.concat((\n                data,\n                tf.zeros([1,N_TARGET_FRAMES-N_FRAMES,N_COLS], dtype=tf.float32)\n            ), axis=1)\n        # Downsample\n        data = tf.image.resize(\n            data,\n            [1, N_TARGET_FRAMES],\n            method=tf.image.ResizeMethod.BILINEAR,\n        )\n        \n        # Squeeze Batch Dimension\n        data = tf.squeeze(data, axis=[0])\n        \n        return data\n    \npreprocess_layer = PreprocessLayer()","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:08:58.300816Z","iopub.execute_input":"2023-07-31T07:08:58.301114Z","iopub.status.idle":"2023-07-31T07:09:01.924696Z","shell.execute_reply.started":"2023-07-31T07:08:58.301088Z","shell.execute_reply":"2023-07-31T07:09:01.923609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train Dataset\n\nNow, we proceed to create the training dataset through the function `get_train_dataset`, which takes the variables `X`, `y`, and the batch size as arguments.\n\nInside the function, an array called `sample_idxs` is created, containing the indices of the samples in the dataset.\n\nThen, an infinite loop is created using the `while True` statement, allowing continuous iteration over the dataset in batches.\n\nIn each iteration, random indices are obtained from `sample_idxs` using the `np.random.choice` function. `batch_size` random indices are selected without replacement.\n\nNext, two dictionaries are created: `inputs` and `outputs`. In `inputs`, the batches of frames and phrases corresponding to the selected random indices are assigned. In `outputs`, the batch of `y` corresponding to the random indices is assigned.\n\nFinally, the `yield` statement is used to return the dictionaries `inputs` and `outputs` as a pair of values in each iteration of the generator.\n\nThis way, when calling the function `get_train_dataset`, a generator is obtained that can be used in a `for` loop to iterate over the training batches in each training epoch. Each batch contains the input and output data corresponding to the selected random indices.\n\nIn the context of model training, the data generator is used to efficiently provide training samples batches during the training process.\n\nThe data generator is used as the `x` value in the `fit()` method of the model. Instead of passing all the training data directly as a large tensor, which could occupy a lot of memory and be inefficient, a data generator is used to provide smaller batches of samples at each training step.\n\nThe `get_train_dataset` function returns an infinite generator that, in each iteration, takes a random batch of training samples (`batch_size`) from the training dataset X and y. Each batch of samples is a dictionary with two keys: 'frames' and 'phrase'. The input data is in the 'frames' key, containing sequences of numerical data corresponding to the \"frames\" (input data sequences). The output data (labels) is in the 'phrase' key, containing sequences of integer indices representing the \"phrases\" (desired output sequences).\n\nWhen the model's `fit()` method receives this data generator (`train_dataset`) as x, it runs in each training epoch and takes batches of samples from the generator to update the model's weights. This allows training the model using all the training data but without loading them all into memory at the same time, which is useful for large datasets. The generator will continue providing batches of samples indefinitely during training, ensuring that the model is trained on all available training samples.\n","metadata":{"papermill":{"duration":0.023965,"end_time":"2023-07-02T11:33:52.934821","exception":false,"start_time":"2023-07-02T11:33:52.910856","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Train Dataset Iterator\ndef get_train_dataset(X, y, batch_size=BATCH_SIZE):\n    sample_idxs = np.arange(len(X))\n    while True:\n        # Get random indices\n        random_sample_idxs = np.random.choice(sample_idxs, batch_size)\n        \n        inputs = {\n            'frames': X[random_sample_idxs],\n            'phrase': y[random_sample_idxs],\n        }\n        outputs = y[random_sample_idxs]\n        \n        yield inputs, outputs","metadata":{"papermill":{"duration":0.033186,"end_time":"2023-07-02T11:33:52.992213","exception":false,"start_time":"2023-07-02T11:33:52.959027","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:01.929746Z","iopub.execute_input":"2023-07-31T07:09:01.932086Z","iopub.status.idle":"2023-07-31T07:09:01.940167Z","shell.execute_reply.started":"2023-07-31T07:09:01.932049Z","shell.execute_reply":"2023-07-31T07:09:01.939006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train Dataset\ntrain_dataset = get_train_dataset(X_train, y_train)\n\ntrain_dataset","metadata":{"papermill":{"duration":0.03572,"end_time":"2023-07-02T11:33:53.051965","exception":false,"start_time":"2023-07-02T11:33:53.016245","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:01.945434Z","iopub.execute_input":"2023-07-31T07:09:01.948236Z","iopub.status.idle":"2023-07-31T07:09:01.957924Z","shell.execute_reply.started":"2023-07-31T07:09:01.948201Z","shell.execute_reply":"2023-07-31T07:09:01.956893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Training Steps Per Epoch\nTRAIN_STEPS_PER_EPOCH = math.ceil(N_TRAIN_SAMPLES / BATCH_SIZE)\nprint(f'TRAIN_STEPS_PER_EPOCH: {TRAIN_STEPS_PER_EPOCH}')","metadata":{"papermill":{"duration":0.034525,"end_time":"2023-07-02T11:33:53.111793","exception":false,"start_time":"2023-07-02T11:33:53.077268","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:01.962883Z","iopub.execute_input":"2023-07-31T07:09:01.965295Z","iopub.status.idle":"2023-07-31T07:09:01.973148Z","shell.execute_reply.started":"2023-07-31T07:09:01.965261Z","shell.execute_reply":"2023-07-31T07:09:01.972252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Validation Dataset\n\nRepeat the process for the validation data","metadata":{"papermill":{"duration":0.02478,"end_time":"2023-07-02T11:33:53.161179","exception":false,"start_time":"2023-07-02T11:33:53.136399","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Validation Set\ndef get_val_dataset(X, y, batch_size=BATCH_SIZE):\n    offsets = np.arange(0, len(X), batch_size)\n    while True:\n        # Iterate over whole validation set\n        for offset in offsets:\n            inputs = {\n                'frames': X[offset:offset+batch_size],\n                'phrase': y[offset:offset+batch_size],\n            }\n            outputs = y[offset:offset+batch_size]\n\n            yield inputs, outputs","metadata":{"papermill":{"duration":0.0344,"end_time":"2023-07-02T11:33:53.220231","exception":false,"start_time":"2023-07-02T11:33:53.185831","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:01.975687Z","iopub.execute_input":"2023-07-31T07:09:01.977536Z","iopub.status.idle":"2023-07-31T07:09:01.990469Z","shell.execute_reply.started":"2023-07-31T07:09:01.977502Z","shell.execute_reply":"2023-07-31T07:09:01.989205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Validation Dataset\nif USE_VAL:\n    val_dataset = get_val_dataset(X_val, y_val)","metadata":{"papermill":{"duration":0.031785,"end_time":"2023-07-02T11:33:53.276318","exception":false,"start_time":"2023-07-02T11:33:53.244533","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:01.992395Z","iopub.execute_input":"2023-07-31T07:09:01.992803Z","iopub.status.idle":"2023-07-31T07:09:02.001443Z","shell.execute_reply.started":"2023-07-31T07:09:01.99277Z","shell.execute_reply":"2023-07-31T07:09:02.000329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if USE_VAL:\n    N_VAL_STEPS_PER_EPOCH = math.ceil(N_VAL_SAMPLES / BATCH_SIZE)\n    print(f'N_VAL_STEPS_PER_EPOCH: {N_VAL_STEPS_PER_EPOCH}')","metadata":{"papermill":{"duration":0.032228,"end_time":"2023-07-02T11:33:53.33261","exception":false,"start_time":"2023-07-02T11:33:53.300382","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:02.003086Z","iopub.execute_input":"2023-07-31T07:09:02.003492Z","iopub.status.idle":"2023-07-31T07:09:02.016801Z","shell.execute_reply.started":"2023-07-31T07:09:02.00346Z","shell.execute_reply":"2023-07-31T07:09:02.015762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Config\n\nPrior to the modeling process, we will explain what a Transformer model consists of:\n\nUnlike traditional recurrent architectures such as Recurrent Neural Networks (RNN) or Convolutional Neural Networks (CNN), which rely on recurrent connections or convolutions to capture temporal dependencies in sequential data, the Transformer model is based on attention mechanisms to capture these dependencies. The main focus of the attention mechanism is to calculate relationships between all elements of a sequence, assigning weights or importances to each element based on its relevance to the task at hand.\n\nThe Transformer model consists of two main components: the encoder and the decoder. Each component is composed of multiple stacked layers. The encoder is responsible for processing the original text input and capturing relevant information from each word or token in the sequence. This is done through attention mechanisms, allowing the model to focus on the most important parts of the input at each step. In the encoder, multi-head attention layers and fully connected neural network layers are applied to capture and process contextual information from the input.\n\nOn the other hand, the decoder is responsible for generating the desired output, such as translating the original input to another language. The decoder also uses attention mechanisms, but in this case, it focuses on both the original input (through masked attention) and the output generated so far. This allows the model to generate output words or tokens based on the current context and previously generated words.\n\nIn each layer of both the encoder and the decoder, normalization and regularization techniques are applied, such as layer normalization and dropout, to stabilize and regularize the learning process. These techniques help improve the model's generalization and prevent overfitting.\n\nIn summary, the encoder processes the original input and captures contextual information using attention mechanisms, while the decoder utilizes the information from the encoder and the generated output so far to produce the desired output.\n\nNow, let's define several variables related to the architecture and configuration of a Transformer model:\n\n- `LAYER_NORM_EPS`: Represents the epsilon value used in the layer normalization of the Transformer model.\n- `UNITS_ENCODER` and `UNITS_DECODER`: These variables indicate the size of the final output and the embeddings of the encoder and decoder, respectively.\n- `NUM_BLOCKS_ENCODER` and `NUM_BLOCKS_DECODER`: Represent the number of blocks (layers) in the encoder and decoder of the Transformer.\n- `NUM_HEADS`: Indicates the number of attention heads in the multi-head attention mechanism of the Transformer.\n- `MLP_RATIO`: Is the multiplication factor used to calculate the size of the feed-forward layer in the Transformer block.\n- `EMBEDDING_DROPOUT`, `MLP_DROPOUT_RATIO`, `MHA_DROPOUT_RATIO`, and `CLASSIFIER_DROPOUT_RATIO`: These variables define the dropout rates used in different parts of the model, such as the embedding layer, the feed-forward layer, and the classifier layer.\n- `INIT_HE_UNIFORM`, `INIT_GLOROT_UNIFORM`, and `INIT_ZEROS`: Are initializers used to initialize the weights of the model's layers. `INIT_HE_UNIFORM` uses He Uniform initialization, `INIT_GLOROT_UNIFORM` uses Glorot Uniform initialization, and `INIT_ZEROS` initializes the weights with zeros.\n- `GELU`: Is an activation function called GELU (Gaussian Error Linear Unit) used in the model.\n","metadata":{"papermill":{"duration":0.0243,"end_time":"2023-07-02T11:33:53.381866","exception":false,"start_time":"2023-07-02T11:33:53.357566","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Epsilon value for layer normalisation\nLAYER_NORM_EPS = 1e-6\n\n# final embedding and transformer embedding size\nUNITS_ENCODER = 384\nUNITS_DECODER = 256\n\n# Transformer\nNUM_BLOCKS_ENCODER = 4\nNUM_BLOCKS_DECODER = 3\nNUM_HEADS = 4\nMLP_RATIO = 2\n\n# Dropout\nEMBEDDING_DROPOUT = 0.00\nMLP_DROPOUT_RATIO = 0.30\nMHA_DROPOUT_RATIO = 0.20\nCLASSIFIER_DROPOUT_RATIO = 0.10\n\n# Initiailizers\nINIT_HE_UNIFORM = tf.keras.initializers.he_uniform\nINIT_GLOROT_UNIFORM = tf.keras.initializers.glorot_uniform\nINIT_ZEROS = tf.keras.initializers.constant(0.0)\n# Activations\nGELU = tf.keras.activations.gelu","metadata":{"papermill":{"duration":0.036845,"end_time":"2023-07-02T11:33:53.443416","exception":false,"start_time":"2023-07-02T11:33:53.406571","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:02.018565Z","iopub.execute_input":"2023-07-31T07:09:02.018969Z","iopub.status.idle":"2023-07-31T07:09:02.030236Z","shell.execute_reply.started":"2023-07-31T07:09:02.018937Z","shell.execute_reply":"2023-07-31T07:09:02.029098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Landmark Embedding\n\n\n## LandmarkEmbedding Class\n\nIn this step, we will define a class called `LandmarkEmbedding` that inherits from `tf.keras.Model` and is used for embedding landmarks using fully connected layers. The embedding of a landmark refers to transforming a representation of coordinates (landmarks) into a resulting vector by applying non-linear transformations and learning more expressive representations of the landmarks. These embedded representations can capture important features of landmarks, such as their relative position, shape, or any other relevant information for the given task.\n\nThe `LandmarkEmbedding` class has the following components:\n\n- The `__init__` method is responsible for initializing the class and configuring its attributes. It receives two parameters: `units`, which represents the dimension of the embedding, and `name`, which is the name of the embedding.\n- The `build` method is used to construct the model's architecture. It receives the `input_shape`, which defines the shape of the input data. In this case, we build an embedding for landmarks.\n    * Within the `build` method, a weight called `empty_embedding` is defined using `self.add_weight`. This weight represents the embedding for a missing landmark in a frame and is initialized with zeros.\n    * Next, a sequence of dense layers is defined using `tf.keras.Sequential`. These layers are used to perform the embedding of landmarks. The first dense layer has `units` units and uses the GELU (Gaussian Error Linear Unit) activation function. The second dense layer also has `units` units and is initialized with the He Uniform method.\n- The `call` method is the model's call function and is used to perform the embedding of the input data `x`. In this function, `tf.where` is used to apply a condition. If the sum of landmarks in a frame is equal to zero, which means the landmark is null, the `empty_embedding` is used. Otherwise, the embedding of the landmark data is performed using the sequence of dense layers defined earlier.","metadata":{"papermill":{"duration":0.024508,"end_time":"2023-07-02T11:33:53.492645","exception":false,"start_time":"2023-07-02T11:33:53.468137","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Embeds a landmark using fully connected layers\nclass LandmarkEmbedding(tf.keras.Model):\n    def __init__(self, units, name):\n        super(LandmarkEmbedding, self).__init__(name=f'{name}_embedding')\n        self.units = units\n        self.supports_masking = True\n        \n    def build(self, input_shape):\n        # Embedding for missing landmark in frame, initialized with zeros\n        self.empty_embedding = self.add_weight(\n            name=f'{self.name}_empty_embedding',\n            shape=[self.units],\n            initializer=INIT_ZEROS,)\n        # Embedding: 2 dense layers\n        self.dense = tf.keras.Sequential([\n            tf.keras.layers.Dense(self.units, name=f'{self.name}_dense_1', use_bias=False, kernel_initializer=INIT_GLOROT_UNIFORM, activation=GELU),\n            tf.keras.layers.Dense(self.units, name=f'{self.name}_dense_2', use_bias=False, kernel_initializer=INIT_HE_UNIFORM),\n        ], name=f'{self.name}_dense')\n\n    def call(self, x):\n        # if the landmark = 0 -> use empty embedding (return 0s), else use dense embedding\n        return tf.where(\n                # Checks whether landmark is missing in frame\n                tf.reduce_sum(x, axis=2, keepdims=True) == 0,\n                # If so, the empty embedding is used\n                self.empty_embedding,\n                # Otherwise the landmark data is embedded\n                self.dense(x),\n            )","metadata":{"papermill":{"duration":0.03711,"end_time":"2023-07-02T11:33:53.554368","exception":false,"start_time":"2023-07-02T11:33:53.517258","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:02.03226Z","iopub.execute_input":"2023-07-31T07:09:02.032685Z","iopub.status.idle":"2023-07-31T07:09:02.051156Z","shell.execute_reply.started":"2023-07-31T07:09:02.032649Z","shell.execute_reply":"2023-07-31T07:09:02.049846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Ejemplo : We define the landmarks for three frames using TensorFlow constant tensors. Each frame has two landmarks represented as coordinate pairs `(x, y)`. We pass the landmarks through the embedding model using the `call()` function of the model. This will generate the embedded representations of the landmarks. In each print, you will see the embedded representations of the landmarks in the form of tensors. Each tensor will have the shape `(n_batch, num_landmarks, num_dimensions)`, where `num_landmarks` is the number of landmarks in the frame, and `num_dimensions` is the number of dimensions for each landmark (in this case, 2 for the x and y coordinates).","metadata":{}},{"cell_type":"code","source":"# Create an instance of the LandmarkEmbedding model\nembedding_model = LandmarkEmbedding(units=2, name='landmark')\n\n# Define the landmarks of the three frames\nframe_1 = tf.constant([[[1, 2], [3, 4]]])\nframe_2 = tf.constant([[[0, 0], [5, 6]], [[7, 8], [9, 10]]])\n\n# Pass the landmarks through the embedding model\nembedding_1 = embedding_model(frame_1)\nembedding_2 = embedding_model(frame_2)\n\n# Print the embedded representations of the landmarks\nprint(f'Frame 1 shape: {frame_1.shape} ')\nprint(\"Frame 1 embeddings:\")\nprint(embedding_1)\n\nprint(\"\\n\\nFrame 2 embeddings:\")\nprint(embedding_2)","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:02.054529Z","iopub.execute_input":"2023-07-31T07:09:02.055622Z","iopub.status.idle":"2023-07-31T07:09:04.586034Z","shell.execute_reply.started":"2023-07-31T07:09:02.055589Z","shell.execute_reply":"2023-07-31T07:09:04.584954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the output for the first frame, a tensor of shape `(1, 2, 2)` and data type *float32* is displayed. This means there is only one sample in the batch, two landmarks in the frame (two vectors), and two dimensions for each landmark (x, y).\nEach value in the tensor represents the embedded representation of a landmark in the frame. For example, the first value `[0.00237609, -0.14469941]` corresponds to the embedded representation of the first landmark `[1,2]`, and the second value `[0.00064886, -0.10494988]` corresponds to the embedded representation of the second landmark.\n\nThe LandmarkEmbedding model performs a transformation of the original landmark values using the dense layers defined in the `build` method.\n\n- If a landmark value is equal to zero, indicating that the landmark is missing in the frame, an empty embedding (represented by `self.empty_embedding`) is used instead of performing the transformation.\n- If a landmark value is not zero, the transformation is applied using the dense layers (`self.dense`). In this case, a linear operation is followed by the GELU activation function to obtain the final embedding of the landmark.\n\nIn this case, two consecutive dense layers are used. The first layer (`self.dense_1`) performs a matrix multiplication of the input values with a weight matrix. The weights of this layer are initialized using the Glorot uniform initialization. Then, a GELU activation function is applied, which is a smoothed version of the ReLU and introduces non-linearity into the transformation. The second layer (`self.dense_2`) performs a similar operation, multiplying the output of the previous layer by another weight matrix initialized with the He uniform method.\n\nIn the provided example, the number of `units` for the embedding model is set to '2'. This means that the dimension of the embedding output will be of size 2 (thus retaining the same tensor shape as the original).\n\nIn general, the purpose of the embedding is not necessarily to reduce the dimension of the input data but rather to provide a dense and more meaningful representation of the original data. The goal is to capture relevant features and semantic relationships between the input data.\n\nThrough the dense layers of the model, linear and non-linear transformations are applied to the input data, which can capture more complex patterns and relationships between landmarks.\n\nTherefore, even though the size of the output embedding dimension remains the same as the input dimension (in this case, 2), the dense layers introduce weights and biases that allow for non-linear transformation of the data. This can help capture more discriminative features and effectively represent relationships between landmarks.","metadata":{}},{"cell_type":"markdown","source":"# Embedding\n\nThe `Embedding` class is a subclass of `tf.keras.Model` used to create an embedded representation for each input frame. Let's go through the function of this class step by step:\n\n- In the `__init__` method, the `Embedding` class is initialized, and the property `supports_masking=True` is set. This indicates that the model supports the use of masks to ignore certain inputs during training or inference.\n\n- In the `build` method, the model is constructed, and the necessary parameters and layers are defined. In this case, two special layers are created: the \"positional embedding\" layer for each frame and the \"landmark embedding\" layer.\n\n    - The \"positional embedding\" layer is used to add information about the position of each frame in the sequence. It is initialized with a tensor of shape `[N_TARGET_FRAMES, UNITS_ENCODER]` filled with zeros and will be adjusted during training to capture positional patterns.\n\n    - The \"landmark embedding\" layer is used to generate a specialized representation of the landmarks for each frame. This representation captures important features of the landmarks and will be used for specific tasks later on.\n\n- In the `call` method, the forward propagation logic of the model is defined. Here, the processing of the input data is performed:\n    - First, the input data `x` is normalized by subtracting the mean and dividing by the standard deviation. This ensures that the data is on an appropriate scale for the model.\n\n    - Next, the normalized input is passed through the \"landmark embedding\" layer `dominant_hand_embedding`. This means that an embedding is applied to the landmark data. You can think of the landmarks as reference points in a sequence of images. The landmark embedding is similar to a transformation that takes these points and transforms them into a more compact and information-rich representation, making it easier for the model to capture important features of the landmarks and process them effectively.\n\n    - After that, information about the position of each frame in the sequence (video) is added. This is important because the order of frames can be relevant for understanding the context and relationship between them. For example, in a video, the order of frames can represent the temporal flow of actions. The positional embedding adds this positional information to the embedded data. In the code, we use a tensor called `positional_embedding` containing the positional embedding corresponding to each frame in the sequence.\n\n- Finally, the resulting embedded representation `x` is returned.\n\nIn summary, the `Embedding` class is used to generate an embedded representation for each input frame. This involves normalizing the input data, generating an embedded representation of the landmarks, and adding positional information to the resulting embedded representation.","metadata":{"papermill":{"duration":0.023829,"end_time":"2023-07-02T11:33:53.60253","exception":false,"start_time":"2023-07-02T11:33:53.578701","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Creates embedding for each frame\nclass Embedding(tf.keras.Model):\n    def __init__(self):\n        super(Embedding, self).__init__()\n        self.supports_masking = True\n    \n    def build(self, input_shape):\n        # Positional embedding for each frame index\n        self.positional_embedding = tf.Variable(\n            initial_value=tf.zeros([N_TARGET_FRAMES, UNITS_ENCODER], dtype=tf.float32),\n            trainable=True,\n            name='embedding_positional_encoder',)\n        # Embedding layer for Landmarks\n        self.dominant_hand_embedding = LandmarkEmbedding(UNITS_ENCODER, 'dominant_hand')\n\n    def call(self, x, training=False):\n        # Normalize data before aplying embedding\n        x = tf.where(\n                tf.math.equal(x, 0.0),\n                0.0,\n                (x - MEANS) / STDS,)\n        # Dominant Hand: apply landmark embedding to extract information\n        x = self.dominant_hand_embedding(x)\n        # Add Positional Encoding\n        x = x + self.positional_embedding\n        \n        return x","metadata":{"papermill":{"duration":0.036321,"end_time":"2023-07-02T11:33:53.663162","exception":false,"start_time":"2023-07-02T11:33:53.626841","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:04.590709Z","iopub.execute_input":"2023-07-31T07:09:04.593171Z","iopub.status.idle":"2023-07-31T07:09:04.605222Z","shell.execute_reply.started":"2023-07-31T07:09:04.593119Z","shell.execute_reply":"2023-07-31T07:09:04.604188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> Example: *Let's visualize the positional part of the `Embedding` class. For this purpose, we will create a class that performs only positional embedding since the example for the landmark embedding has already been shown. Afterwards, we will observe the combined effect.*","metadata":{}},{"cell_type":"code","source":"# positional Embedding class\nclass PositionalEmbedding(tf.keras.Model):\n    def __init__(self):\n        super(PositionalEmbedding, self).__init__()\n    \n    def build(self, input_shape):\n        self.positional_embedding = tf.Variable(\n            initial_value=tf.zeros([1, 164], dtype=tf.float32),\n            trainable=True,\n            name='positional_embedding')\n    \n    def call(self, x):\n        return x + self.positional_embedding","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:04.607989Z","iopub.execute_input":"2023-07-31T07:09:04.608331Z","iopub.status.idle":"2023-07-31T07:09:04.626053Z","shell.execute_reply.started":"2023-07-31T07:09:04.608295Z","shell.execute_reply":"2023-07-31T07:09:04.625009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create model instance\nembedding_model = PositionalEmbedding()\n\n# Define input data of dimension [1,1,164]\ninput_data = X_train[0,0,:].reshape(1,1,164)\n\n# make the embedding\nembedded_data = embedding_model(input_data)\n\nprint(f'Input Shape: {input_data.shape}')\n\nprint(f'\\nEmbedded Shape: {embedded_data.shape}')\n\n# Imprimir los datos de salida incrustados\nprint(\"\\nInput Data:\")\nprint(input_data)\n\nprint(\"\\n\\nEmbedded Data:\")\nprint(embedded_data)","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:04.631692Z","iopub.execute_input":"2023-07-31T07:09:04.634462Z","iopub.status.idle":"2023-07-31T07:09:04.659758Z","shell.execute_reply.started":"2023-07-31T07:09:04.634428Z","shell.execute_reply":"2023-07-31T07:09:04.65886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The positional embedding is calculated using mathematical functions or predefined patterns that assign a unique numerical value to each position in the sequence. These numerical values are used to represent relative positional information, which the model can utilize to perform specific operations, such as recognizing patterns of spatial or temporal dependency in the sequence.\n\nIn this example, the sequence refers to the input data that you are passing through the model. The shape of the input data is `[1, 164]`, indicating that you have a sequence of length 164.\n\nIn general, it cannot be assumed that the values of elements at the same index in different batches (frames) will be equal in a general context.\n\nIf we had data of type `[N, 164]`, where N represents the number of batches (number of sequences), and 164 is the length of the sequence (number of columns), the values of elements at the same index can be different in different sequences unless there is specific logic in your data that guarantees otherwise.\n\nIt is important to consider that the values of elements at the same index in different batches may vary to capture variability in the data and allow the model to learn patterns and generalize better.\n\nNow, let's see the combined effect of both embeddings:","metadata":{"execution":{"iopub.status.busy":"2023-07-14T10:53:31.205937Z","iopub.execute_input":"2023-07-14T10:53:31.206404Z","iopub.status.idle":"2023-07-14T10:53:31.21732Z","shell.execute_reply.started":"2023-07-14T10:53:31.206367Z","shell.execute_reply":"2023-07-14T10:53:31.216247Z"}}},{"cell_type":"code","source":"# Create model instance\nembedding_model = Embedding()\n\n# Define input data of dimension [1,1,164]\ninput_data = X_train[0,0,:].reshape(1,1,164)\n\n# make the embedding\nembedded_data = embedding_model(input_data)\n\nprint(f'Input Shape: {input_data.shape}')\n\nprint(f'\\nEmbedded Shape: {embedded_data.shape}')\n\n# Imprimir los datos de salida incrustados\nprint(\"\\nInput Data:\")\nprint(input_data)\n\nprint(\"\\n\\nEmbedded Data:\")\nprint(embedded_data)","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:04.661011Z","iopub.execute_input":"2023-07-31T07:09:04.66134Z","iopub.status.idle":"2023-07-31T07:09:04.740442Z","shell.execute_reply.started":"2023-07-31T07:09:04.661308Z","shell.execute_reply":"2023-07-31T07:09:04.739488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see how the shape of the data has changed due to the following transformation:\n\n* Landmark Embedding: The input data is passed through the `dominant_hand_embedding` layer of the `LandmarkEmbedding` class. This layer applies a linear transformation to the input data using dense layers. Since the dimension of the embedding (`UNITS_ENCODER`) is 384, in this case, each individual input element is transformed into a representation of shape `[1, 1, 384]`. This is because the dimension of the sequence is kept at 1 since there is only one input element in the sequence.\n\n* Positional Embedding: After the landmark embedding, a positional embedding is added to each transformed element. The positional embedding, stored in the variable `positional_embedding`, has a shape of `[N_TARGET_FRAMES, UNITS_ENCODER]`. In this case, `N_TARGET_FRAMES` is set to 128, and `UNITS_ENCODER` is set to 384.\n\n* Final Result: By adding the landmark embedding and the positional embedding, the final result with shape `[1, 128, 384]` is obtained. The first dimension (1) refers to the number of elements in the batch, the second dimension (128) represents the target frames, and the third dimension (384) is the dimension of the resulting embedding.\n\nNow, let's see what happens with input data of shape `[1, 128, 164]`:","metadata":{}},{"cell_type":"code","source":"# Create model instance\nembedding_model = Embedding()\n\n# Define input data of dimension [1,128,164]\ninput_data = X_train[0,:,:].reshape(1,128,164) # [128,164] -> [1,128,164]\n\n# Make embedding\nembedded_data = embedding_model(input_data)\n\n\nprint(f'Input Shape: {input_data.shape}')\nprint(f'\\nEmbedded Shape: {embedded_data.shape}')\n\nprint(\"\\nInput Data:\")\nprint(input_data)\n\nprint(\"\\n\\nEmbedded Data:\")\nprint(embedded_data)","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:04.742219Z","iopub.execute_input":"2023-07-31T07:09:04.742604Z","iopub.status.idle":"2023-07-31T07:09:04.794123Z","shell.execute_reply.started":"2023-07-31T07:09:04.742568Z","shell.execute_reply":"2023-07-31T07:09:04.793125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the example, the input data has a shape of `[1, 128, 164]`. When passing through the `Embedding` class, both landmark and positional embeddings are performed, allowing us to capture information about the landmarks while preserving their position in the video.\n\nSpecifically, for the landmark embedding, the `LandmarkEmbedding` class is used, which has a final dense layer with `UNITS_ENCODER` units. In the code, `UNITS_ENCODER=384`. Therefore, after applying the landmark embedding, the data is adjusted to dimension 384.\n\nNext, the positional embedding, which is a tensor of shape `[384,]`, is added, which means it will be added element-wise to each data point without changing the dimensions.\n\nIn the context of natural language processing, embeddings increase the dimensionality of the data, not the other way around.\n\nWhen applying an embedding to text data, such as words or sequences, the goal is to assign each element (word or token) a vector representation in a higher-dimensional space. In other words, it transforms a discrete representation (e.g., a word represented as an index) into a vector representation with a set of real numbers representing the features of the element.\n\nFor example, if you have a vocabulary of 10,000 words and you want to represent each word as a 300-dimensional vector using an embedding, then the dimensionality of the words would increase from 1 (the index of the word in the vocabulary) to 300.\n\nThis increase in dimensionality is beneficial because the embedding can better capture the semantics and relationships between words or tokens, which helps improve performance in subsequent tasks such as text classification, text generation, among others.\n","metadata":{}},{"cell_type":"markdown","source":"# Transformer\n\nNow, a multi-head attention layer is implemented using the scaled dot product operation. Multi-head attention is a key component in Transformer models used in natural language processing tasks and other domains.\n\nIn a Transformer with multi-head attention, the concepts of queries, keys, and values are crucial as they are used to calculate attention between elements in a sequence. The attention mechanism in a Transformer allows it to focus on relevant parts of the input during the encoding or decoding stage. Attention is calculated at each position in the input sequence based on queries, keys, and values.\n\nHere's an explanation of these three components:\n\n- Queries: These are vectors representing the current position in the input sequence for which attention is to be calculated. In other words, they are the positions that are being encoded or decoded. Each query is used to calculate the degree of relevance or similarity between this position and all other positions in the sequence.\n\n- Keys: These are vectors representing all positions in the input sequence. They are used to calculate the degree of relevance between queries and different positions in the sequence. Keys are crucial in determining which parts of the input are important for each query.\n\n- Values: These are vectors containing the actual information at each position in the input sequence. They are used to calculate attention weights and weigh the importance of each position based on its relevance to queries and keys.\n\nThe calculation of attention involves measuring the similarity between queries and keys to obtain attention weights. These attention weights indicate how much importance should be given to each position in the sequence based on its relationship with the current query. The values are then weighted using these attention weights to obtain the attended or contextualized representation of the current query.\n\nIn the context of multi-head attention, this process is performed multiple times, each time with different sets of parameters (queries, keys, and values), allowing the Transformer to capture different patterns and relationships in the input sequence more effectively.\n\nWe divide the explanation of our attention mechanism into steps:\n\n- `scaled_dot_product` is a function that performs the scaled dot product between the queries (q), keys (k), and values (v) of attention. The scaled dot product is a fundamental operation in attention, and it is used to calculate the relevance between queries and keys. In the context of multi-head attention, the scaled dot product is applied in parallel in each attention head to capture different relationships and patterns in the data. The process of the scaled dot product is done in three steps:\n\n    - Dot Product: The dot product is calculated between the queries (q) and the transposed keys (k). This is achieved through the matrix multiplication operation. The result is a matrix representing the relevance between each pair of query and key.\n    - Scaling: After calculating the dot product, it is scaled by dividing it by the square root of the dimension of the queries (q). This scaling operation helps stabilize the attention process and ensures that the attention values are not too large.\n    - Softmax and Weighting: Next, a softmax function is applied to the scaled matrix. The softmax converts the values into a probability distribution, meaning that each attention value represents the relative importance of the corresponding key for a given query. Then, this attention distribution is used to weigh the values (v) and obtain a combined representation of the weighted values.\n    \n    The reason we use this scaled dot product layer in multi-head attention is that it allows us to capture the relationships and dependencies between queries and keys more expressively. By performing the scaled dot product in multiple attention heads, we can capture different patterns and relationships in parallel, which enhances the ability to model relevant information in the data.\n    \n- `MultiHeadAttention` is a multi-head attention layer. It takes as input the queries (q), keys (k), and values (v) of attention. It can also receive an optional attention mask to mask certain inputs. The main parameters of this layer are:\n    - `d_model`: the dimension of the representation space for queries, keys, and values.\n    - `num_of_heads`: the number of attention heads to be used.\n    - `dropout`: the dropout probability to regularize the output.\n    - `d_out` (optional): if provided, specifies the output dimension after the multi-head attention layer.\n- In the `call` method, the following operations are performed:\n    - For each attention head, an independent linear transformation is applied to the queries (q), keys (k), and values (v) using dense layers (wq, wk, wv) to project them into smaller subspaces.\n    - The `scaled_dot_product` function is called to calculate attention with the projections of q, k, and v for each attention head.\n    - The outputs of all attention heads are concatenated along the last dimension to obtain a combined representation.\n    - The combined representation goes through a final linear layer (wo) to obtain the output of the multi-head attention layer.\n    - A dropout layer is applied to regularize the output before returning it.\n    \nIn summary, this implementation performs multi-head attention in parallel, where each head learns different weighted representations of the inputs. Then, it combines the outputs of all heads into a final representation that can be used in natural language processing tasks and other machine learning problems.\n","metadata":{"papermill":{"duration":0.025127,"end_time":"2023-07-02T11:33:53.713795","exception":false,"start_time":"2023-07-02T11:33:53.688668","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"> Example: Explanation of queries, keys, and values in a Transformer with multi-head attention. Let's assume we have the following sentence: \"The cat is sleeping on the carpet.\" The process of multi-head attention would work as follows:\n\n1. First, we tokenize the sentence into word vectors or tokens:\n\n    - **Queries:**\n        - \"The\"\n        - \"cat\"\n        - \"is\"\n        - \"sleeping\"\n        - \"on\"\n        - \"the\"\n        - \"carpet\"\n\n    - **Keys:**\n        - \"The\"\n        - \"cat\"\n        - \"is\"\n        - \"sleeping\"\n        - \"on\"\n        - \"the\"\n        - \"carpet\"\n\n    - **Values:**\n        - Vector of information corresponding to each word.\n\n2. Next, we calculate the similarity (e.g., using the dot product) between each query and each key to obtain the attention weights.\n\n3. The attention weights indicate which words are more relevant to each query. These weights are then applied to the values corresponding to the words in the input to obtain the final result.\n\n4. Finally, the multi-head attention combines the weighted representations of the words to obtain a high-quality representation for the original sentence.\n","metadata":{}},{"cell_type":"code","source":"# based on: https://stackoverflow.com/questions/67342988/verifying-the-implementation-of-multihead-attention-in-transformer\n# replaced softmax with softmax layer to support masked softmax\nclass MultiHeadAttention(tf.keras.layers.Layer):\n    def __init__(self, d_model, n_heads, dropout, d_out=None):\n        super(MultiHeadAttention,self).__init__()\n        # Number of Units in Model\n        self.d_model = d_model\n        # Number of Attention Heads\n        self.n_heads = n_heads\n        # Number of Units in Intermediate Layers\n        self.depth = d_model // 2\n        # Scaling Factor Of Values\n        self.scale = 1.0 / tf.math.sqrt(tf.cast(self.depth, tf.float32))\n        # Learnable Projections to Depth\n        self.wq = self.fused_mha(self.depth)\n        self.wk = self.fused_mha(self.depth)\n        self.wv = self.fused_mha(self.depth)\n        # Output Projection\n        self.wo = tf.keras.layers.Dense(d_model if d_out is None else d_out, use_bias=False)\n        # Softmax Activation Which Supports Masking\n        self.softmax = tf.keras.layers.Softmax()\n        # Reshaping Of Multiple Attention heads to Single Value\n        self.reshape = tf.keras.Sequential([\n            # [attention heads, number of frames, d_model] → [number of frames, n_heads, d_model // n_heads]\n            tf.keras.layers.Permute([2, 1, 3]),\n            # [number of frames, attention heads, d_model] → [number of frames, d_model]\n            tf.keras.layers.Reshape([N_TARGET_FRAMES, self.depth]),\n        ])\n        # Output Dropout\n        self.do = tf.keras.layers.Dropout(dropout)\n        self.supports_masking = True\n        \n    # Single dense layer for all attention heads\n    def fused_mha(self, dim):\n        return tf.keras.Sequential([\n            # Single dense layer\n            tf.keras.layers.Dense(dim, use_bias=False),\n            # Reshape to [number of frames, number of attention head, depth]\n            tf.keras.layers.Reshape([N_TARGET_FRAMES, self.n_heads, dim // self.n_heads]),\n            # Permutate to [number of attention heads, number of frames, depth]\n            tf.keras.layers.Permute([2, 1, 3]),\n        ])\n        \n    def call(self, q, k, v, attention_mask=None, training=False):\n        # Projections to attention heads\n        Q = self.wq(q)\n        K = self.wk(k)\n        V = self.wv(v)\n        # Matrix multiply QxK to acquire attention scores\n        x = tf.matmul(Q, K, transpose_b=True) * self.scale\n        # Softmax attention scores and Multiply with Values\n        x = self.softmax(x, mask=attention_mask) @ V\n        # Reshape to flatten attention heads\n        x = self.reshape(x)\n        # Output projection\n        x = self.wo(x)\n        # Dropout\n        x = self.do(x, training=training)\n        return x","metadata":{"papermill":{"duration":0.04073,"end_time":"2023-07-02T11:33:53.779389","exception":false,"start_time":"2023-07-02T11:33:53.738659","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:04.796811Z","iopub.execute_input":"2023-07-31T07:09:04.79748Z","iopub.status.idle":"2023-07-31T07:09:04.812144Z","shell.execute_reply.started":"2023-07-31T07:09:04.797441Z","shell.execute_reply":"2023-07-31T07:09:04.811187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Encoder\n\n[source](https://keras.io/examples/nlp/neural_machine_translation_with_transformer/)\n\nNow, we will create an Encoder based on transformer blocks. An Encoder takes an input sequence `x` and transforms it into a vector representation, which can then be used for subsequent tasks such as translation or text generation. Let's explain the `Encoder` class:\n\n- We define a class called `Encoder`, which inherits from `tf.keras.Model`. This class represents the Encoder component in a Transformer-based model.\n- In the constructor `__init__`, we initialize the Encoder with the number of attention blocks (`num_blocks`) that will be used in the encoding process.\n- The `build` method is responsible for creating the necessary components for the attention blocks. These components are:\n  - Lists `ln_1s`, `mhas`, `ln_2s`, and `mlps` that will store normalization layers, multi-head attention layers, additional normalization layers, and multi-layer perceptrons, respectively. We create one list for each attention block.\n  - For each attention block, we create a normalization layer (`tf.keras.layers.LayerNormalization`) called `ln_1`, a multi-head attention layer called `mha`, another normalization layer `ln_2`, and a multi-layer perceptron `mlp`.\n  - We also check if the dimension of the Encoder is different from the dimension of the Decoder. If so, we create a projection layer (`dense_out`) to adapt the output dimension of the Encoder to the dimension of the Decoder.\n- The `call` method is used to perform the encoding of the input (`x`). It takes `x` and a variable `x_inp` that contains information about the input sequence (such as attention masks). Here, the processing is done in each attention block:\n    - We create an \"attention mask\" (`attention_mask`) to prevent the model from attending to invalid or padded parts of the input sequence. This mask is calculated from `x_inp`, which is a variable containing information about the input sequence.\n    - We iterate over the attention blocks defined previously (`ln_1`, `mha`, `ln_2`, and `mlp`) and apply them sequentially to the input `x`.\n    - In each iteration, we first apply normalization (`ln_1`) to the input `x` plus the output of the multi-head attention (`mha`). This is known as a \"residual connection,\" where the original input is added to the output of the multi-head attention before applying normalization. This technique helps to avoid gradient vanishing problems and improves information flow.\n    - Next, we apply another normalization (`ln_2`) to the output of the residual connection plus the output of the multi-layer perceptron (`mlp`).\n    - After passing through all the attention blocks and perceptrons, if the dimension of the Encoder is different from the dimension of the Decoder (`UNITS_ENCODER != UNITS_DECODER`), we perform a projection (`dense_out`) to adapt the output dimension of the Encoder to the dimension of the Decoder.\n    - Finally, the function returns the resulting vector representation after applying all the attention blocks and the multi-layer perceptron.\n\n> Note: The multi-layer perceptron and the embedding are different concepts used in the fields of machine learning and natural language processing. Here are their differences:\n\nThe multi-layer perceptron (MLP), or feedforward neural network, is a neural network architecture consisting of multiple layers, including at least one input layer, one or more hidden layers, and one output layer. The neurons in the layer are called nodes or perceptrons.\n\nIn the MLP, each neuron in a layer is connected to all neurons in the previous and subsequent layers via weighted connections. The weighted connections allow a neuron to adjust its \"sensitivity\" to certain patterns in the input data. During the training process, the weights are adjusted to minimize the error between the network's predicted outputs and the known actual outputs, in the case of supervised learning.\n\nThe weights can be positive or negative, meaning they can amplify or reduce the contribution of a particular input. By learning the optimal weights, the neural network can find relevant patterns and relationships in the input data, enabling it to perform tasks such as classification, regression, or other specific tasks for which it has been trained. During the training process, the weights of these connections are adjusted so that the model can learn to represent complex and nonlinear relationships in the data.\n\nThe main purpose of the MLP is to learn nonlinear functions for classification or regression on complex datasets. Through the hidden layers and the nonlinearity introduced by activation functions, the MLP can learn more complex representations than those achieved with linear models.\n\nOn the other hand, the embedding is a continuous and dense vector representation of words or tokens in a low-dimensional vector space. It is specifically used in NLP to map words or tokens from the discrete word space to a continuous vector space. Its goal is to capture the meaning and semantic relationships between words in natural language processing tasks.","metadata":{"papermill":{"duration":0.024245,"end_time":"2023-07-02T11:33:53.828275","exception":false,"start_time":"2023-07-02T11:33:53.80403","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Encoder based on multiple transformer blocks\nclass Encoder(tf.keras.Model):\n    def __init__(self, num_blocks):\n        super(Encoder, self).__init__(name='encoder')\n        self.num_blocks = num_blocks\n        self.supports_masking = True\n    \n    def build(self, input_shape):\n        self.ln_1s = []\n        self.mhas = []\n        self.ln_2s = []\n        self.mlps = []\n        # Make Transformer Blocks\n        for i in range(self.num_blocks):\n            # First Layer Normalisation\n            self.ln_1s.append(tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS))\n            # Multi Head Attention\n            self.mhas.append(MultiHeadAttention(UNITS_ENCODER, NUM_HEADS, MHA_DROPOUT_RATIO))\n            # Second Layer Normalisation\n            self.ln_2s.append(tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS))\n            # Multi Layer Perception\n            self.mlps.append(tf.keras.Sequential([\n                tf.keras.layers.Dense(UNITS_ENCODER * MLP_RATIO, activation=GELU, kernel_initializer=INIT_GLOROT_UNIFORM, use_bias=False),\n                tf.keras.layers.Dropout(MLP_DROPOUT_RATIO),\n                tf.keras.layers.Dense(UNITS_ENCODER, kernel_initializer=INIT_HE_UNIFORM, use_bias=False),\n            ]))\n            # Optional Projection to Decoder Dimension\n            if UNITS_ENCODER != UNITS_DECODER:\n                self.dense_out = tf.keras.layers.Dense(UNITS_DECODER, kernel_initializer=INIT_GLOROT_UNIFORM, use_bias=False)\n                self.apply_dense_out = True\n            else:\n                self.apply_dense_out = False\n                \n    def get_attention_mask(self, x_inp):\n        # Attention Mask\n        attention_mask = tf.math.count_nonzero(x_inp, axis=[2], keepdims=True, dtype=tf.int32)\n        attention_mask = tf.math.count_nonzero(attention_mask, axis=[2], keepdims=False)\n        attention_mask = tf.expand_dims(attention_mask, axis=1)\n        attention_mask = tf.expand_dims(attention_mask, axis=1)\n        return attention_mask\n        \n    def call(self, x, x_inp, training=False):\n        # Attention mask to ignore missing frames\n        attention_mask = self.get_attention_mask(x_inp)\n        # Iterate input over transformer blocks\n        for ln_1, mha, ln_2, mlp in zip(self.ln_1s, self.mhas, self.ln_2s, self.mlps):\n            x = ln_1(x + mha(x, x, x, attention_mask=attention_mask))\n            x = ln_2(x + mlp(x))\n            \n        # Optional Projection to Decoder Dimension\n        if self.apply_dense_out:\n            x = self.dense_out(x)\n    \n        return x","metadata":{"papermill":{"duration":0.041438,"end_time":"2023-07-02T11:33:53.894476","exception":false,"start_time":"2023-07-02T11:33:53.853038","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:04.814651Z","iopub.execute_input":"2023-07-31T07:09:04.815059Z","iopub.status.idle":"2023-07-31T07:09:04.830771Z","shell.execute_reply.started":"2023-07-31T07:09:04.815026Z","shell.execute_reply":"2023-07-31T07:09:04.829916Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Decoder\n\nAfter the `Encoder`, we define the `Decoder` class, which is responsible for using the embeddings obtained by the Encoder to generate an output sequence. Let's break down the different steps:\n\n- In the `__init__` method, we initialize the Decoder model and specify the number of Transformer blocks (`num_blocks`) that will be used in generating the output sequence. Additionally, we indicate that the model supports masking (`supports_masking = True`). Masking is a crucial technique used to handle variable-length sequences during processing.\n\n- In the `build` method, we construct the necessary layers for the Decoder model. This is where we define the different parts of the model that will be used during the output sequence generation. These layers are:\n\n  - `positional_embedding`: This is the positional embedding layer, which is used to assign a numerical representation to each position in the output sequence. It means it assigns a unique representation to each word or token in the sequence.\n  - `char_emb`: This is the character embedding layer, which converts character sequences (phrases) into numerical representations. It is a fundamental step in text processing since neural networks work with numbers, not raw text.\n  - `pos_emb_mha`: This is the causal masked multi-head attention layer for positional embedding. Causal masked attention is a technique that helps the model learn temporal dependency relationships while generating the output sequence. It means it allows the model to \"look\" only at previous words (or itself) during sequence generation.\n  - `pos_emb_ln`: This is the normalization layer applied after the causal masked attention. Normalization is a technique that helps stabilize and speed up the training of the model.\n  - `ln_1s`, `mhas`, `ln_2s`, and `mlps`: These are lists that contain several layers of normalization, multi-head attention, and multi-layer perceptron (MLP) layers, respectively. We build `num_blocks` of these layers to allow the Decoder to perform multiple iterations and capture more complex relationships in the generated output sequence.\n\n- In the `get_causal_attention_mask` method, a causal attention mask is created. This mask is used in the causal masked attention layer to ensure that each position only attends to previous positions (or itself), preventing future information leakage during sequence generation.\n\n- In the `call` method, the processing of the output sequence through the Transformer blocks in the Decoder is performed. The steps are as follows:\n\n1. Preprocessing of the input sequence (`phrase`) is done to convert it into numerical representations (embedding) using the character embedding layer (`char_emb`). Additionally, the special start token (`START_TOKEN`) is added to the beginning of each sequence, and padding tokens (`PAD_TOKEN`) are appended at the end to reach the desired maximum length (`MAX_PHRASE_LENGTH`).\n\n2. The causal attention mask is created using the `get_causal_attention_mask` method to prevent the model from accessing future information during sequence generation.\n\n3. The initial embedding representations are obtained by summing the positional embedding layer (`positional_embedding`) and the character embedding layer (`char_emb`) of the input sequence (`phrase`). This sum combines spatial (positional) and semantic (character embedding) information for each word in the sequence.\n\n4. A normalization layer (`pos_emb_ln`) is applied to the initial embedding representations after the causal masked attention (`pos_emb_mha`). Normalization helps the model train more stably and efficiently.\n\n5. The Transformer blocks (normalization layers, multi-head attention, and MLP) are iterated through a `for` loop. Each block is responsible for processing and improving the embedding representations during sequence generation.\n\n6. Finally, the generated sequence is trimmed to remove any padding tokens (`PAD_TOKEN`), and the output of the Decoder model with a length of `MAX_PHRASE_LENGTH` is obtained.\n\nIn summary, the Decoder model uses multiple Transformer blocks to generate an output sequence based on the information from the Encoder. It applies causal masked attention to capture temporal dependencies and uses normalization and MLP techniques to process and improve the embedding representations during sequence generation. The goal is for the model to generate meaningful and coherent text sequences based on the information from the Encoder.","metadata":{"papermill":{"duration":0.024624,"end_time":"2023-07-02T11:33:53.943518","exception":false,"start_time":"2023-07-02T11:33:53.918894","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Decoder based on multiple transformer blocks\nclass Decoder(tf.keras.Model):\n    def __init__(self, num_blocks):\n        super(Decoder, self).__init__(name='decoder')\n        self.num_blocks = num_blocks\n        self.supports_masking = True\n    \n    def build(self, input_shape):\n        # Causal Mask Batch Size 1\n        self.causal_mask = self.get_causal_attention_mask()\n        # Positional Embedding, initialized with zeros\n        self.positional_embedding = tf.Variable(\n            initial_value=tf.zeros([N_TARGET_FRAMES, UNITS_DECODER], dtype=tf.float32),\n            trainable=True,\n            name='embedding_positional_encoder',\n        )\n        # Character Embedding\n        self.char_emb = tf.keras.layers.Embedding(N_UNIQUE_CHARACTERS, UNITS_DECODER, embeddings_initializer=INIT_ZEROS)\n        # Positional Encoder MHA\n        self.pos_emb_mha = MultiHeadAttention(UNITS_DECODER, NUM_HEADS, MHA_DROPOUT_RATIO)\n        self.pos_emb_ln = tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS)\n        # First Layer Normalisation\n        self.ln_1s = []\n        self.mhas = []\n        self.ln_2s = []\n        self.mlps = []\n        # Make Transformer Blocks\n        for i in range(self.num_blocks):\n            # First Layer Normalisation\n            self.ln_1s.append(tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS))\n            # Multi Head Attention\n            self.mhas.append(MultiHeadAttention(UNITS_DECODER, NUM_HEADS, MHA_DROPOUT_RATIO))\n            # Second Layer Normalisation\n            self.ln_2s.append(tf.keras.layers.LayerNormalization(epsilon=LAYER_NORM_EPS))\n            # Multi Layer Perception\n            self.mlps.append(tf.keras.Sequential([\n                tf.keras.layers.Dense(UNITS_DECODER * MLP_RATIO, activation=GELU, kernel_initializer=INIT_GLOROT_UNIFORM, use_bias=False),\n                tf.keras.layers.Dropout(MLP_DROPOUT_RATIO),\n                tf.keras.layers.Dense(UNITS_DECODER, kernel_initializer=INIT_HE_UNIFORM, use_bias=False),\n            ]))\n            \n    def get_causal_attention_mask(self):\n        i = tf.range(N_TARGET_FRAMES)[:, tf.newaxis]\n        j = tf.range(N_TARGET_FRAMES)\n        mask = tf.cast(i >= j, dtype=tf.int32)\n        mask = tf.reshape(mask, (1, N_TARGET_FRAMES, N_TARGET_FRAMES))\n        mult = tf.concat(\n            [tf.expand_dims(1, -1), tf.constant([1, 1], dtype=tf.int32)],\n            axis=0,\n        )\n        mask = tf.tile(mask, mult)\n        mask = tf.cast(mask, tf.float32)\n        return mask\n    \n    def get_attention_mask(self, x_inp):\n        # Attention Mask\n        attention_mask = tf.math.count_nonzero(x_inp, axis=[2], keepdims=True, dtype=tf.int32)\n        attention_mask = tf.math.count_nonzero(attention_mask, axis=[2], keepdims=False)\n        attention_mask = tf.expand_dims(attention_mask, axis=1)\n        attention_mask = tf.expand_dims(attention_mask, axis=1)\n        return attention_mask\n        \n    def call(self, encoder_outputs, phrase, x_inp, training=False):\n        # Batch Size\n        B = tf.shape(encoder_outputs)[0]\n        # Cast to INT32\n        phrase = tf.cast(phrase, tf.int32)\n        # Prepend SOS Token\n        phrase = tf.pad(phrase, [[0,0], [1,0]], constant_values=START_TOKEN, name='prepend_sos_token')\n        # Pad With PAD Token\n        phrase = tf.pad(phrase, [[0,0], [0,N_TARGET_FRAMES-MAX_PHRASE_LENGTH-1]], constant_values=PAD_TOKEN, name='append_pad_token')\n        # Positional Embedding\n        x = self.positional_embedding + self.char_emb(phrase)\n        # Causal Attention\n        x = self.pos_emb_ln(x + self.pos_emb_mha(x, x, x, attention_mask=self.causal_mask))\n        # Attention mask to ignore missing frames\n        attention_mask = self.get_attention_mask(x_inp)\n        # Iterate input over transformer blocks\n        for ln_1, mha, ln_2, mlp in zip(self.ln_1s, self.mhas, self.ln_2s, self.mlps):\n            x = ln_1(x + mha(x, encoder_outputs, encoder_outputs, attention_mask=attention_mask))\n            x = ln_2(x + mlp(x))\n        # Slice 31 Characters\n        x = tf.slice(x, [0, 0, 0], [-1, MAX_PHRASE_LENGTH, -1])\n    \n        return x","metadata":{"papermill":{"duration":0.046582,"end_time":"2023-07-02T11:33:54.014771","exception":false,"start_time":"2023-07-02T11:33:53.968189","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:04.83405Z","iopub.execute_input":"2023-07-31T07:09:04.83483Z","iopub.status.idle":"2023-07-31T07:09:04.857571Z","shell.execute_reply.started":"2023-07-31T07:09:04.834794Z","shell.execute_reply":"2023-07-31T07:09:04.856741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This function takes a single parameter `B`. The purpose of the causal attention mask is to prevent the Decoder model from accessing future information during sequence generation. In tasks like text generation, machine translation, or other autoregressive tasks, it's essential that the model only has access to information from previous positions while predicting the current position. Otherwise, the model might use future information for predictions, leading to inconsistent and unrealistic results (data leakage).\n\n1. The function creates a numeric sequence of length `N_TARGET_FRAMES` using `tf.range(N_TARGET_FRAMES)`. Then, additional dimensions are added to this sequence using `[:, tf.newaxis]` to convert it into a column vector.\n\n2. A Cartesian product is taken between the sequence created in step 1 and the original sequence using `i >= j`. This creates a matrix of zeros and ones, where the elements are zero in all positions corresponding to future positions (columns) relative to the current position (row).\n\n3. The matrix of zeros and ones is reshaped to have a shape of `(1, N_TARGET_FRAMES, N_TARGET_FRAMES)`. The additional dimension at the beginning is necessary for its later use in calculating the final mask.\n\n4. The value of `B` is concatenated with a constant `[1, 1]` using `tf.concat` to create a tensor that specifies how many times the mask will be replicated to handle the batch size in the model.\n\n5. `tf.tile` is used to replicate the causal attention mask `B` times, where `B` is the batch size. This is done so that the mask can be applied to all elements in the batch in the model.\n\n6. The resulting attention mask is converted to a tensor of type `float32` using `tf.cast` to ensure compatibility with the attention computation in TensorFlow. Finally, the function returns the resulting causal attention mask.\n","metadata":{}},{"cell_type":"code","source":"# Causal Attention to make decoder not attent to future characters which it needs to predict\ndef get_causal_attention_mask(B):\n    i = tf.range(N_TARGET_FRAMES)[:, tf.newaxis]\n    j = tf.range(N_TARGET_FRAMES)\n    mask = tf.cast(i >= j, dtype=tf.int32)\n    mask = tf.reshape(mask, (1, N_TARGET_FRAMES, N_TARGET_FRAMES))\n    mult = tf.concat(\n        [tf.expand_dims(B, -1), tf.constant([1, 1], dtype=tf.int32)],\n        axis=0,\n    )\n    mask = tf.tile(mask, mult)\n    mask = tf.cast(mask, tf.float32)\n    return mask\n\nget_causal_attention_mask(1)","metadata":{"papermill":{"duration":0.094639,"end_time":"2023-07-02T11:33:54.137348","exception":false,"start_time":"2023-07-02T11:33:54.042709","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:04.867687Z","iopub.execute_input":"2023-07-31T07:09:04.868073Z","iopub.status.idle":"2023-07-31T07:09:04.909622Z","shell.execute_reply.started":"2023-07-31T07:09:04.868049Z","shell.execute_reply":"2023-07-31T07:09:04.908669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We proceed to see step by step the process of the attention mask:","metadata":{}},{"cell_type":"code","source":"B = 1\ni = tf.range(N_TARGET_FRAMES)[:, tf.newaxis]\nj = tf.range(N_TARGET_FRAMES)\nprint(i[:5])\nprint(j)","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:04.912823Z","iopub.execute_input":"2023-07-31T07:09:04.913085Z","iopub.status.idle":"2023-07-31T07:09:04.926656Z","shell.execute_reply.started":"2023-07-31T07:09:04.913062Z","shell.execute_reply":"2023-07-31T07:09:04.925758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, when performing the operation `tf.cast(i >= j, dtype=tf.int32)`, for each element `i` out of the 128 elements, we will obtain an array of size 128 with one 1 (when `i=j`) and the rest zeros. By repeating this operation for all `i`, we will have 128 elements of size 128.","metadata":{}},{"cell_type":"code","source":"mask = tf.cast(i >= j, dtype=tf.int32)\nprint(mask)","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:04.929653Z","iopub.execute_input":"2023-07-31T07:09:04.92997Z","iopub.status.idle":"2023-07-31T07:09:04.939217Z","shell.execute_reply.started":"2023-07-31T07:09:04.929944Z","shell.execute_reply":"2023-07-31T07:09:04.938248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We reshape the tensor by adding a dimension to the left and create `mult` to replicate mask `B` times. As in this case `B=1`, it will not change anything. `expand_dims` expands `B` by one dimension at the position specified by the second argument. In this case, it adds a dimension to the right of `B`, so it becomes an array `[1]` instead of a single element `1`.","metadata":{}},{"cell_type":"code","source":"tf.expand_dims(B, -1)","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:04.940939Z","iopub.execute_input":"2023-07-31T07:09:04.941199Z","iopub.status.idle":"2023-07-31T07:09:04.949462Z","shell.execute_reply.started":"2023-07-31T07:09:04.941168Z","shell.execute_reply":"2023-07-31T07:09:04.948474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mask = tf.reshape(mask, (1, N_TARGET_FRAMES, N_TARGET_FRAMES)) # reshape [128,128] to [1,128,128]  \nmult = tf.concat(\n        [tf.expand_dims(B, -1), tf.constant([1, 1], dtype=tf.int32)],\n        axis=0,)\n\nprint(mult)\n\nmask = tf.tile(mask, mult)\nprint(mask)","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:04.950947Z","iopub.execute_input":"2023-07-31T07:09:04.951549Z","iopub.status.idle":"2023-07-31T07:09:04.962087Z","shell.execute_reply.started":"2023-07-31T07:09:04.951513Z","shell.execute_reply":"2023-07-31T07:09:04.961003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mask = tf.cast(mask, tf.float32) # turn into float\nmask","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:04.963844Z","iopub.execute_input":"2023-07-31T07:09:04.964296Z","iopub.status.idle":"2023-07-31T07:09:04.97454Z","shell.execute_reply.started":"2023-07-31T07:09:04.964188Z","shell.execute_reply":"2023-07-31T07:09:04.973226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Non Pad/START/END Token Accuracy\n\nThis is a custom metric that inherits from `tf.keras.metrics.Metric`. Its purpose is to calculate the top-k accuracy of a model with multi-dimensional output.\n\nInstead of simply comparing the predicted class by the model with the true class, the Top-K Accuracy considers whether the true class is among the K most probable classes predicted by the model.\n\nImagine you have a classification model that needs to predict the correct category of an image. If the Top-K is 1 (Top-1 Accuracy), then the model must predict the correct category as the one with the highest probability. However, if the Top-K is 5 (Top-5 Accuracy), the model is considered correct if the correct category is among the top 5 categories with the highest probabilities predicted by the model.\n\nThe metric is calculated as follows:\n\n1. For each sample or example in the dataset, the model generates a probability distribution over all possible classes.\n2. The model selects the K classes with the highest probabilities and checks if the true class is among these K classes.\n3. If the true class is among the K most probable classes, then the model has correctly predicted that sample for the calculation of the Top-K Accuracy.\n4. The Top-K Accuracy metric is calculated by taking the percentage of samples where the model has correctly predicted the true class within the K most probable classes.\n5. The value of K in the Top-K Accuracy allows adjusting the model's level of strictness. A high Top-K value (e.g., Top-5 or Top-10) can be useful when some classification tasks are difficult and there are several possible classes close in terms of probability. On the other hand, a low Top-K value (e.g., Top-1) is more stringent and only considers the class with the highest predicted probability.\n\nThe steps for implementing the `TopKAccuracy` class are as follows:\n\n- In the `__init__` method, the class is initialized and it receives a parameter `k`, which indicates the value of k in the top-k accuracy. For example, if `k=1`, it will calculate the top-1 accuracy (also known as the accuracy in the top-ranked probability). Inside the `__init__` method, an instance of `tf.keras.metrics.SparseTopKCategoricalAccuracy` is created, which is a built-in metric in TensorFlow that calculates top-k accuracy for sparse categorical outputs. This instance is saved in the attribute `self.top_k_acc` for later use.\n\n- In the `update_state` method, the calculation of the top-k accuracy is performed by updating the state of the metric. This method takes three arguments:\n    - `y_true`: The true label of the model, which contains the indices of the correct classes for each sample.\n    - `y_pred`: The output predicted by the model, which contains the probabilities of the classes for each sample.\n    - `sample_weight`: Optional weights for the samples.\n\n    - A series of transformations is performed to prepare the data for the top-k accuracy calculation:\n        - Both `y_true` and `y_pred` are reshaped to become one-dimensional tensors. This is done to ensure that the data is in the correct format for the top-k accuracy metric.\n        - The indices of those `y_true` that are less than `N_UNIQUE_CHARACTERS` (remember that there were a total of 62 distinct tokens or characters) are obtained. This is done to filter only the valid indices in the metric calculation, as there might be indices outside the valid range in certain cases.\n        - `tf.gather` is used to select only the samples corresponding to the valid indices in both `y_true` and `y_pred`.\n\n    - Finally, the `update_state` method of `self.top_k_acc` is called, passing the preprocessed data `y_true` and `y_pred` to perform the actual calculation of the top-k accuracy.\n\n- The `result` and `reset_state` methods simply return the current top-k accuracy result and reset the state of the metric, respectively.","metadata":{"papermill":{"duration":0.024935,"end_time":"2023-07-02T11:33:54.189242","exception":false,"start_time":"2023-07-02T11:33:54.164307","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# TopK accuracy for multi dimensional output\nclass TopKAccuracy(tf.keras.metrics.Metric):\n    def __init__(self, k, **kwargs):\n        super(TopKAccuracy, self).__init__(name=f'top{k}acc', **kwargs)\n        self.top_k_acc = tf.keras.metrics.SparseTopKCategoricalAccuracy(k=k)\n\n    def update_state(self, y_true, y_pred, sample_weight=None):\n        y_true = tf.reshape(y_true, [-1])\n        y_pred = tf.reshape(y_pred, [-1, N_UNIQUE_CHARACTERS])\n        character_idxs = tf.where(y_true < N_UNIQUE_CHARACTERS0)\n        y_true = tf.gather(y_true, character_idxs, axis=0)\n        y_pred = tf.gather(y_pred, character_idxs, axis=0)\n        self.top_k_acc.update_state(y_true, y_pred)\n\n    def result(self):\n        return self.top_k_acc.result()\n    \n    def reset_state(self):\n        self.top_k_acc.reset_state()","metadata":{"papermill":{"duration":0.037227,"end_time":"2023-07-02T11:33:54.251228","exception":false,"start_time":"2023-07-02T11:33:54.214001","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:04.976519Z","iopub.execute_input":"2023-07-31T07:09:04.976952Z","iopub.status.idle":"2023-07-31T07:09:04.986403Z","shell.execute_reply.started":"2023-07-31T07:09:04.97692Z","shell.execute_reply":"2023-07-31T07:09:04.985201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loss Weights\n\n","metadata":{"papermill":{"duration":0.024945,"end_time":"2023-07-02T11:33:54.301243","exception":false,"start_time":"2023-07-02T11:33:54.276298","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Create Initial Loss Weights All Set To 1\nloss_weights = np.ones(N_UNIQUE_CHARACTERS, dtype=np.float32)\n# Set Loss Weight Of Pad Token To 0\nloss_weights[PAD_TOKEN] = 0","metadata":{"papermill":{"duration":0.033995,"end_time":"2023-07-02T11:33:54.360318","exception":false,"start_time":"2023-07-02T11:33:54.326323","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:04.988282Z","iopub.execute_input":"2023-07-31T07:09:04.988994Z","iopub.status.idle":"2023-07-31T07:09:05.000636Z","shell.execute_reply.started":"2023-07-31T07:09:04.98896Z","shell.execute_reply":"2023-07-31T07:09:04.999732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Sparse Categorical Crossentropy With Label Smoothing¶\n\n\nThis function calculates the loss using the `categorical_crossentropy` loss function with label smoothing support. The purpose of label smoothing is to avoid the model becoming too confident and \"overconfident\" in its predictions.\n\nWhen training a classification model, we typically use one-hot encoded labels or target outputs. In one-hot format, each sample has a label that is a binary vector with a \"1\" at the position corresponding to the true class and \"0\" in all other positions. For example, if we have a classification problem with 3 classes, a sample belonging to class 2 would have a label `[0, 1, 0]`.\n\nLabel smoothing involves reducing the absolute confidence of the model in individual labels by redistributing a small amount of probability from the true class to other classes. Instead of using \"1\" at the position corresponding to the true class, a value close to \"1\" is used, and the rest of the value is distributed among the other classes.\n\nFor example, if we apply label smoothing with a smoothing value of 0.1 in the previous problem, the label `[0, 1, 0]` could become `[0.05, 0.9, 0.05]`. This means that instead of being completely certain that the sample belongs to class 2, the model becomes slightly less confident and distributes a small amount of probability to other classes. This effect generally results in improved model generalization and prevents overfitting.\n\nThe function contains the following elements:\n\n1. `idxs = tf.where(y_true != PAD_TOKEN)`: In this line, `tf.where` is used to find the positions in `y_true` where the value is not equal to the padding token (`PAD_TOKEN`). The goal is to filter out indices that do not correspond to padding tokens, as we do not want to consider them in the loss calculation.\n\n2. `y_true = tf.gather_nd(y_true, idxs)`: Here, `tf.gather_nd` is used to obtain the label values from `y_true` that are not padding tokens.\n\n3. `y_pred = tf.gather_nd(y_pred, idxs)`: Similarly, `tf.gather_nd` is used to obtain the model predictions corresponding to the filtered indices from `y_true`.\n\n4. `y_true = tf.cast(y_true, tf.int32)`: For the calculation of the loss with `tf.keras.losses.categorical_crossentropy`, `y_true` needs to be of type int32, so a conversion is performed.\n\n5. `y_true = tf.one_hot(y_true, N_UNIQUE_CHARACTERS, axis=1)`: The `tf.one_hot` function is used to convert the labels `y_true` into a one-hot representation.\n\n6. `loss = tf.keras.losses.categorical_crossentropy(y_true, y_pred, label_smoothing=0.25, from_logits=True)`: Here, the loss is calculated using the `categorical_crossentropy` loss function. The function has an optional parameter `label_smoothing`, which is set to 0.25 in this case. The value 0.25 indicates how much \"confidence\" should be added to the true class and subtracted from the true class probability while adding to the other classes in the loss. `from_logits=True` indicates that `y_pred` has not been passed through a softmax activation function before calculating the loss.\n\n7. `loss = tf.math.reduce_mean(loss)`: Finally, the loss calculated in the previous step is reduced to a single value by taking its average using `tf.math.reduce_mean`. This is necessary to obtain a single loss measure representing the overall model performance on the dataset.\n","metadata":{"papermill":{"duration":0.024662,"end_time":"2023-07-02T11:33:54.409907","exception":false,"start_time":"2023-07-02T11:33:54.385245","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# source:: https://stackoverflow.com/questions/60689185/label-smoothing-for-sparse-categorical-crossentropy\ndef scce_with_ls(y_true, y_pred):\n    # Filter Pad Tokens\n    idxs = tf.where(y_true != PAD_TOKEN)\n    y_true = tf.gather_nd(y_true, idxs)\n    y_pred = tf.gather_nd(y_pred, idxs)\n    # One Hot Encode Sparsely Encoded Target Sign\n    y_true = tf.cast(y_true, tf.int32)\n    y_true = tf.one_hot(y_true, N_UNIQUE_CHARACTERS, axis=1)\n    # Categorical Crossentropy with native label smoothing support\n    loss = tf.keras.losses.categorical_crossentropy(y_true, y_pred, label_smoothing=0.25, from_logits=True)\n    loss = tf.math.reduce_mean(loss)\n    return loss","metadata":{"papermill":{"duration":0.034321,"end_time":"2023-07-02T11:33:54.468872","exception":false,"start_time":"2023-07-02T11:33:54.434551","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:05.002314Z","iopub.execute_input":"2023-07-31T07:09:05.00265Z","iopub.status.idle":"2023-07-31T07:09:05.012702Z","shell.execute_reply.started":"2023-07-31T07:09:05.002619Z","shell.execute_reply":"2023-07-31T07:09:05.011618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model\n\nOnce the different sub-elements that make up the model are defined, we create the model.\n\n- The model has two input elements `tf.keras.layers.Input` corresponding to the elements in the dictionary `X_batch`. The first input is for the input data called \"frames,\" which we have defined as sequences of numerical data with shape `[N_TARGET_FRAMES, N_COLS] = [128, 164]`. The second input is for the input data called \"phrase,\" which consists of sequences of words represented by integer indices with shape `[MAX_PHRASE_LENGTH] = 32`.\n\n- The construction of the model starts with the \"frames\" input data using `x = frames_inp`. Next, a `Masking` layer is applied to ignore null values in the \"frames\" input data. This is done to deal with variable-length sequences where null values represent time points without valid information.\n\n- Subsequently, the `Embedding` layer is applied. We will end up with a dimension of `[N, 128, 384]`.\n\n- Next, the \"frames\" data is passed through Transformer blocks using the `Encoder(NUM_BLOCKS_ENCODER)(x, frames_inp)` layer. The Transformer blocks are responsible for processing and extracting relevant features from the input sequences.\n","metadata":{"papermill":{"duration":0.025057,"end_time":"2023-07-02T11:33:54.518693","exception":false,"start_time":"2023-07-02T11:33:54.493636","status":"completed"},"tags":[]}},{"cell_type":"code","source":"print(f' # elements in \"X_batch\": {len(X_batch)}') # frames and phrase\nprint(f' Shape of the first element \"frames\": {X_batch[\"frames\"].shape}') # [N_batch, TARGET_FRAMES, COLS] \nprint(f' Shape of the second element \"phrase\": {X_batch[\"phrase\"].shape}') # [N_batch, MAX_PHRASE_LENGTH]","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:05.014352Z","iopub.execute_input":"2023-07-31T07:09:05.01469Z","iopub.status.idle":"2023-07-31T07:09:05.025325Z","shell.execute_reply.started":"2023-07-31T07:09:05.014658Z","shell.execute_reply":"2023-07-31T07:09:05.024188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We begin with two tensors that have the same shape as 'frames' and 'phrase'. We apply the mask to handle null values in the data and proceed to create a dense representation using Landmark Embedding and Positional Embedding. Subsequently, we apply the encoder to enable the model to learn and study the underlying patterns in the data.","metadata":{}},{"cell_type":"code","source":"# Inputs\nframes_inp = tf.keras.layers.Input([N_TARGET_FRAMES, N_COLS], dtype=tf.float32, name='frames')\nphrase_inp = tf.keras.layers.Input([MAX_PHRASE_LENGTH], dtype=tf.int32, name='phrase')\n# Frames\nx = frames_inp\n\n# Masking to eliminate NaN values in frames\nx = tf.keras.layers.Masking(mask_value=0.0, input_shape=(N_TARGET_FRAMES, N_COLS))(x)\nprint(f' Input shape: {x.shape}')\n# Embedding\nx = Embedding()(x)\n\nprint(f' Shape post Embedding: {x.shape}')\n    \n# Encoder Transformer Blocks\nx = Encoder(NUM_BLOCKS_ENCODER)(x, frames_inp)\n\nprint(f' Shape post encoder: {x.shape}')\nprint(f' Units Decoder: {UNITS_DECODER}')","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:05.026518Z","iopub.execute_input":"2023-07-31T07:09:05.027494Z","iopub.status.idle":"2023-07-31T07:09:06.742087Z","shell.execute_reply.started":"2023-07-31T07:09:05.027459Z","shell.execute_reply":"2023-07-31T07:09:06.741081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- After passing through the encoder blocks, the data is then fed through the `Decoder(NUM_BLOCKS_DECODER)(x, phrase_inp)` layer. This layer is responsible for generating text sequences based on the features extracted by the encoder.\n\n- Next, a `Sequential` layer is added, consisting of a `Dropout` layer to prevent overfitting, followed by a `Dense` layer serving as the output layer of the model. The number of neurons in this `Dense` layer is equal to `N_UNIQUE_CHARACTERS`, which corresponds to the total number of possible classes or tokens in the text generation problem.\n\n- The outputs of the model are assigned to `x`.","metadata":{}},{"cell_type":"code","source":"# Decoder\nx = Decoder(NUM_BLOCKS_DECODER)(x, phrase_inp, frames_inp)\nprint(f'Shape post Decoder: {x.shape}')    \nprint(f'# of unique characters: {N_UNIQUE_CHARACTERS}')\n# Classifier\nx = tf.keras.Sequential([\n    # Dropout\n    tf.keras.layers.Dropout(CLASSIFIER_DROPOUT_RATIO), # CLASSIFIER_DROPOUT_RATIO = 0.1\n    # Output Neurons: 62 classes (different tokens)\n    tf.keras.layers.Dense(N_UNIQUE_CHARACTERS, activation=tf.keras.activations.linear,\n                          kernel_initializer=INIT_HE_UNIFORM, use_bias=False),\n    ], name='classifier')(x)\n    \noutputs = x","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:06.743429Z","iopub.execute_input":"2023-07-31T07:09:06.744988Z","iopub.status.idle":"2023-07-31T07:09:08.437738Z","shell.execute_reply.started":"2023-07-31T07:09:06.744951Z","shell.execute_reply":"2023-07-31T07:09:08.43677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- The model is constructed using `tf.keras.models.Model`, taking the previously defined inputs and outputs.\n\n- The loss function `scce_with_ls`, which is the Categorical Crossentropy Loss with support for Label Smoothing, is defined as the model's loss function.\n\n- The optimizer for the model is defined using `RectifiedAdam`, which is a variant of the Adam optimizer.\n\n- Two evaluation metrics for the model are defined: Top-1 accuracy (`TopKAccuracy(1)`) and Top-5 accuracy (`TopKAccuracy(5)`).\n\n- The model is then compiled using `model.compile`, where the loss function, optimizer, metrics, and loss weights are specified. The previously defined loss weights are used to adjust the relative contribution of each class in the calculation of the model's total loss.\n\nThe model we have constructed is a seq2seq model with attention, commonly used in tasks such as text generation and machine translation.","metadata":{}},{"cell_type":"code","source":"# Create Tensorflow Model\nmodel = tf.keras.models.Model(inputs=[frames_inp, phrase_inp], outputs=outputs)\n    \n# Categorical Crossentropy Loss With Label Smoothing\nloss = scce_with_ls\n    \n# Adam Optimizer\noptimizer = tfa.optimizers.RectifiedAdam(sma_threshold=4)\noptimizer = tfa.optimizers.Lookahead(optimizer, sync_period=5)\n\n# TopK Metrics\nmetrics = [\n        TopKAccuracy(1),\n        TopKAccuracy(5),]\n    \nmodel.compile(\n    loss=loss,\n    optimizer=optimizer,\n    metrics=metrics,\n    loss_weights=loss_weights,\n    )\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:08.439094Z","iopub.execute_input":"2023-07-31T07:09:08.440162Z","iopub.status.idle":"2023-07-31T07:09:08.525187Z","shell.execute_reply.started":"2023-07-31T07:09:08.440127Z","shell.execute_reply":"2023-07-31T07:09:08.524303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After all the necessary elements for creating the model have been seen, we define the `get_model` function that contains all of them grouped together.","metadata":{}},{"cell_type":"code","source":"def get_model():\n    # Inputs\n    frames_inp = tf.keras.layers.Input([N_TARGET_FRAMES, N_COLS], dtype=tf.float32, name='frames')\n    phrase_inp = tf.keras.layers.Input([MAX_PHRASE_LENGTH], dtype=tf.int32, name='phrase')\n    # Frames\n    x = frames_inp\n\n    # Masking\n    x = tf.keras.layers.Masking(mask_value=0.0, input_shape=(N_TARGET_FRAMES, N_COLS))(x)\n    \n    # Embedding\n    x = Embedding()(x)\n    \n    # Encoder Transformer Blocks\n    x = Encoder(NUM_BLOCKS_ENCODER)(x, frames_inp)\n    \n    # Decoder\n    x = Decoder(NUM_BLOCKS_DECODER)(x, phrase_inp, frames_inp)\n    \n    # Classifier\n    x = tf.keras.Sequential([\n        # Dropout\n        tf.keras.layers.Dropout(CLASSIFIER_DROPOUT_RATIO),\n        # Output Neurons\n        tf.keras.layers.Dense(N_UNIQUE_CHARACTERS, activation=tf.keras.activations.linear, kernel_initializer=INIT_HE_UNIFORM, use_bias=False),\n    ], name='classifier')(x)\n    \n    outputs = x\n    \n    # Create Tensorflow Model\n    model = tf.keras.models.Model(inputs=[frames_inp, phrase_inp], outputs=outputs)\n    \n    # Categorical Crossentropy Loss With Label Smoothing\n    loss = scce_with_ls\n    \n    # Adam Optimizer\n    optimizer = tfa.optimizers.RectifiedAdam(sma_threshold=4)\n    optimizer = tfa.optimizers.Lookahead(optimizer, sync_period=5)\n\n    # TopK Metrics\n    metrics = [\n        TopKAccuracy(1),\n        TopKAccuracy(5),\n    ]\n    \n    model.compile(\n        loss=loss,\n        optimizer=optimizer,\n        metrics=metrics,\n        loss_weights=loss_weights,\n    )\n    \n    return model","metadata":{"papermill":{"duration":0.03826,"end_time":"2023-07-02T11:33:54.581784","exception":false,"start_time":"2023-07-02T11:33:54.543524","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:08.526416Z","iopub.execute_input":"2023-07-31T07:09:08.526697Z","iopub.status.idle":"2023-07-31T07:09:08.538055Z","shell.execute_reply.started":"2023-07-31T07:09:08.526672Z","shell.execute_reply":"2023-07-31T07:09:08.536898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Input data: X_batch['frames'] & X_batch['phrase']\nfor k, v in X_batch.items():\n    print(f'{k}: {v.shape}')","metadata":{"papermill":{"duration":0.035009,"end_time":"2023-07-02T11:33:54.641801","exception":false,"start_time":"2023-07-02T11:33:54.606792","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:08.539822Z","iopub.execute_input":"2023-07-31T07:09:08.540221Z","iopub.status.idle":"2023-07-31T07:09:08.551639Z","shell.execute_reply.started":"2023-07-31T07:09:08.540189Z","shell.execute_reply":"2023-07-31T07:09:08.550499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.keras.backend.clear_session()\n\nmodel = get_model()","metadata":{"papermill":{"duration":3.159973,"end_time":"2023-07-02T11:33:57.826278","exception":false,"start_time":"2023-07-02T11:33:54.666305","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:08.552902Z","iopub.execute_input":"2023-07-31T07:09:08.553346Z","iopub.status.idle":"2023-07-31T07:09:11.102964Z","shell.execute_reply.started":"2023-07-31T07:09:08.553314Z","shell.execute_reply":"2023-07-31T07:09:11.101962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot model summary\nmodel.summary(expand_nested=True)","metadata":{"papermill":{"duration":0.237637,"end_time":"2023-07-02T11:33:58.088939","exception":false,"start_time":"2023-07-02T11:33:57.851302","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:11.104227Z","iopub.execute_input":"2023-07-31T07:09:11.10666Z","iopub.status.idle":"2023-07-31T07:09:11.253488Z","shell.execute_reply.started":"2023-07-31T07:09:11.106621Z","shell.execute_reply":"2023-07-31T07:09:11.252742Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot Model Architecture\ntf.keras.utils.plot_model(model, show_shapes=True, show_dtype=True, show_layer_names=True, expand_nested=True, show_layer_activations=True)","metadata":{"papermill":{"duration":0.307642,"end_time":"2023-07-02T11:33:58.432321","exception":false,"start_time":"2023-07-02T11:33:58.124679","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:11.2545Z","iopub.execute_input":"2023-07-31T07:09:11.254931Z","iopub.status.idle":"2023-07-31T07:09:11.579756Z","shell.execute_reply.started":"2023-07-31T07:09:11.254898Z","shell.execute_reply":"2023-07-31T07:09:11.578789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Verify Training Flag\n\nNow, we will check if it works correctly:\n\nIn the first `assert`, we obtain the model's predictions using the input dataset `X_batch_small` with the `training=False` indicator. This means that we are performing static inference, where the model will not be updated or changed during this process. The predictions are stored in the variable `pred`.\n\nNext, we run a `for` loop ten times to verify if the predictions are consistent across multiple inference runs. We do this by comparing the original predictions with new predictions obtained again using the input dataset `X_batch_small` and the `training=False` indicator. For this comparison, we use the `tf.cast` function to ensure that the predictions are of type int8 (8-bit integers) and then reduce the resulting tensor using `tf.reduce_min`.\n\nThe `tf.reduce_min` function is used to find the minimum value in the resulting tensor, which should be equal to 1. This indicates that all the original predictions and new predictions should be identical, indicating that the model produces consistent static results during inference.\n\nIn the second `assert`, we again obtain the predictions using the input dataset `X_batch_small`, but this time with the `training=True` indicator. This simulates the training process, where the model utilizes dropout technique to reduce overfitting and improve generalization. Dropout involves randomly deactivating some neurons during training to prevent the model from becoming too reliant on specific features of the training data.\n\nThen, we run a `for` loop ten times to verify if the predictions significantly vary across multiple runs during training. For this, we compare the tensor of original predictions (calculated in step 1) with new predictions obtained with the `X_batch_small` dataset and the `training=True` indicator. We use the `tf.cast` function to ensure that the predictions are of type float32 (floating-point numbers) and then calculate the mean using `tf.reduce_mean`.\n\nWe expect the resulting mean value to be greater than 0.99, indicating that at least 99% of the predictions in each run during training are different from each other due to the applied dropout. This indicates that the model is using different subsets of neurons in different runs, which is desirable for improving generalization and reducing overfitting.","metadata":{"papermill":{"duration":0.037937,"end_time":"2023-07-02T11:33:58.509486","exception":false,"start_time":"2023-07-02T11:33:58.471549","status":"completed"},"tags":[]}},{"cell_type":"code","source":"print(X_batch_small['frames'].shape)\nprint(X_batch_small['phrase'].shape)","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:11.581075Z","iopub.execute_input":"2023-07-31T07:09:11.581512Z","iopub.status.idle":"2023-07-31T07:09:11.587553Z","shell.execute_reply.started":"2023-07-31T07:09:11.581477Z","shell.execute_reply":"2023-07-31T07:09:11.586604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def verify_correct_training_flag():\n    # Verify static output for inference\n    pred = model(X_batch_small, training=False)\n    for _ in tqdm(range(10)):\n        assert tf.reduce_min(tf.cast(pred == model(X_batch_small, training=False), tf.int8)) == 1\n\n    # Verify at least 99% varying output due to dropout during training\n    for _ in tqdm(range(10)):\n        assert tf.reduce_mean(tf.cast(pred != model(X_batch_small, training=True), tf.float32)) > 0.99\n        \nverify_correct_training_flag()","metadata":{"papermill":{"duration":5.527641,"end_time":"2023-07-02T11:34:04.075359","exception":false,"start_time":"2023-07-02T11:33:58.547718","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:11.588531Z","iopub.execute_input":"2023-07-31T07:09:11.589254Z","iopub.status.idle":"2023-07-31T07:09:15.361643Z","shell.execute_reply.started":"2023-07-31T07:09:11.58922Z","shell.execute_reply":"2023-07-31T07:09:15.360655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Verify No NaN Predictions\n\nNext, we check if there are any NaN values in the predictions made by the model.\n\nThe code starts by using the \"predict\" method of the model to obtain predictions for the training dataset (\"train_dataset\"), as no validation dataset is used in this context. The number of steps to be performed is specified using the `steps` parameter, which is set to `N_VAL_STEPS_PER_EPOCH` if the validation dataset is used, or 100 if no validation dataset is used (we are using 100). The \"verbose\" parameter is set to 1.\n\nOnce the predictions are obtained, the code counts how many NaN values are present in the results using the function `np.isnan(y_pred).sum()`. If the result of this count is greater than zero, it indicates that there are NaN values in the predictions.\n\nNext, the code displays a histogram of the predictions to visualize the distribution of logits. In this context, logits refer to the output values of a neural network before applying an activation function, such as the softmax function. Logits represent the scores or scores associated with each class in a classification problem.\n\nFor example, let's assume you have a classification problem with three classes: \"dog,\" \"cat,\" and \"bird.\" After training a neural network, it will produce a set of numerical values as output for a given input, which will correspond to the logits for each class. For instance, you could obtain [2.5, 1.8, 0.1] as logits for a given image. These values represent the scores or confidence levels that the network assigns to each class. In this case, the network is more confident that the image is a \"dog\" because the highest value is 2.5, followed by \"cat\" with 1.8, and \"bird\" with 0.1.\n\nAfter obtaining the logits, we typically apply an activation function, such as the softmax function, to convert these values into probabilities, so they sum up to 1 and can be interpreted as the probabilities of the input belonging to each class. In this example, after applying the softmax function, we could obtain [0.64, 0.29, 0.07], indicating that the network has a 64% probability of the image being a \"dog,\" 29% of being a \"cat,\" and 7% of being a \"bird.\"\n\nThis visualization is helpful to understand the distribution of predictions and to detect any extreme high or low values that could indicate a problem.\nPlease note that the translation is based on the provided explanation and may require adjustments or additional context depending on the overall context of your work.","metadata":{"papermill":{"duration":0.036771,"end_time":"2023-07-02T11:34:04.151722","exception":false,"start_time":"2023-07-02T11:34:04.114951","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Verify No NaN predictions\ndef verify_no_nan_predictions():\n    y_pred = model.predict(\n        val_dataset if USE_VAL else train_dataset,\n        steps=N_VAL_STEPS_PER_EPOCH if USE_VAL else 100,\n        verbose=VERBOSE,\n    )\n\n    print(f'# NaN Values In Predictions: {np.isnan(y_pred).sum()}')\n    \n    plt.figure(figsize=(15,8))\n    plt.title(f'Logit Predictions Initialized Model')\n    pd.Series(y_pred.flatten()).plot(kind='hist', bins=128)\n    plt.xlabel('Logits')\n    plt.grid()\n    plt.show()\n    \nverify_no_nan_predictions()","metadata":{"papermill":{"duration":9.257248,"end_time":"2023-07-02T11:34:13.446151","exception":false,"start_time":"2023-07-02T11:34:04.188903","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:15.363161Z","iopub.execute_input":"2023-07-31T07:09:15.364334Z","iopub.status.idle":"2023-07-31T07:09:24.564269Z","shell.execute_reply.started":"2023-07-31T07:09:15.364298Z","shell.execute_reply":"2023-07-31T07:09:24.563299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Learning Rate Scheduler\n\nNow, we define the function called `lrfn`, which we will use to calculate the learning rate during the model training.\n\nThe `lrfn` function takes several arguments:\n\n- `current_step`: The current training step, indicating how many steps have been completed so far in the model training.\n- `num_warmup_steps`: The number of training steps used for `warm-up` of the `learning rate`. During warm-up, the `learning rate` gradually increases from a very small value to its maximum value.\n- `lr_max`: The maximum value that the `learning rate` can reach during training. After warm-up, the `learning rate` will oscillate between 0 and this maximum value.\n- `num_cycles`: The number of complete cycles of `learning rate` oscillation. A complete cycle means that the `learning rate` goes from its maximum value to 0 and then back to its maximum value. A value of 0.50 indicates that there will be half a complete cycle.\n- `num_training_steps`: The total number of training steps that will be performed during the entire model training.\n\nThe main purpose of the `lrfn` function is to calculate the `learning rate` for each training step. This is divided into two parts:\n\n1. Learning Rate Warm-up: If the current training step is less than the number of warm-up steps (`num_warmup_steps`), then the `learning rate` will gradually increase from a very small value to its maximum value. There are two possible warm-up methods: logarithmic (`'log'`) or exponential (`2 ** -`). This allows the model to adapt more smoothly to the data at the beginning of the training.\n\n2. Learning Rate Oscillation: After the warm-up steps, the `learning rate` will oscillate between 0 and its maximum value (`lr_max`). The shape of this oscillation follows a modified cosine function, which goes from 0 to 1 and then back to 0 as the cycles progress.\n\nThis learning rate function allows for an effective and adaptive way to adjust the `learning rate` during the model training process, ensuring better convergence and optimization performance.","metadata":{"papermill":{"duration":0.037744,"end_time":"2023-07-02T11:34:13.521856","exception":false,"start_time":"2023-07-02T11:34:13.484112","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def lrfn(current_step, num_warmup_steps, lr_max, num_cycles=0.50, num_training_steps=N_EPOCHS):\n    \n    if current_step < num_warmup_steps:\n        if WARMUP_METHOD == 'log':\n            return lr_max * 0.10 ** (num_warmup_steps - current_step)\n        else:\n            return lr_max * 2 ** -(num_warmup_steps - current_step)\n    else:\n        progress = float(current_step - num_warmup_steps) / float(max(1, num_training_steps - num_warmup_steps))\n\n        return max(0.0, 0.5 * (1.0 + math.cos(math.pi * float(num_cycles) * 2.0 * progress))) * lr_max","metadata":{"papermill":{"duration":0.048796,"end_time":"2023-07-02T11:34:13.609084","exception":false,"start_time":"2023-07-02T11:34:13.560288","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:24.565835Z","iopub.execute_input":"2023-07-31T07:09:24.566197Z","iopub.status.idle":"2023-07-31T07:09:24.573156Z","shell.execute_reply.started":"2023-07-31T07:09:24.566164Z","shell.execute_reply":"2023-07-31T07:09:24.57222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we define a function called `plot_lr_schedule`, which will be used to visualize the learning rate schedule. The function takes two arguments:\n\n- `lr_schedule`: A list containing the learning rate values for each training step (epoch).\n- `epochs`: The total number of training epochs.\n\nThe function performs the following steps to plot the learning rate schedule:\n\n- A line is plotted for the scheduled learning rate (`lr_schedule`). A `None` value is added at the beginning and end to prevent the line from reaching the edges of the graph.\n- Points are plotted on the graph for each learning rate in `lr_schedule`. Depending on the total number of epochs, only some points are shown to avoid the graph becoming too dense. The points are labeled with the corresponding learning rate values.\n\nThen, a list called `LR_SCHEDULE` is created using the `lrfn` function, which calculates the learning rate for each step in the number of epochs (in this case, 20). We plot the graph to visualize the learning rate for each epoch.\n\nFinally, a Keras callback called `lr_callback` is created, which will be used during the model training. This callback will automatically adjust the learning rate in each training epoch using the lambda function `step: LR_SCHEDULE[step]`. The learning rate value for each epoch will be obtained from the `LR_SCHEDULE` list that was created earlier.","metadata":{}},{"cell_type":"code","source":"print(f'N_EPOCHS: {N_EPOCHS}')\nprint(f'N_WARMUP_EPOCHS: {N_WARMUP_EPOCHS}')","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:09:24.574946Z","iopub.execute_input":"2023-07-31T07:09:24.575764Z","iopub.status.idle":"2023-07-31T07:09:24.587682Z","shell.execute_reply.started":"2023-07-31T07:09:24.57571Z","shell.execute_reply":"2023-07-31T07:09:24.586717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_lr_schedule(lr_schedule, epochs):\n    fig = plt.figure(figsize=(20, 10))\n    plt.plot([None] + lr_schedule + [None])\n    # X Labels\n    x = np.arange(1, epochs + 1)\n    x_axis_labels = [i if epochs <= 40 or i % 5 == 0 or i == 1 else None for i in range(1, epochs + 1)]\n    plt.xlim([1, epochs])\n    plt.xticks(x, x_axis_labels) # set tick step to 1 and let x axis start at 1\n    \n    # Increase y-limit for better readability\n    plt.ylim([0, max(lr_schedule) * 1.1])\n    \n    # Title\n    schedule_info = f'start: {lr_schedule[0]:.1E}, max: {max(lr_schedule):.1E}, final: {lr_schedule[-1]:.1E}'\n    plt.title(f'Step Learning Rate Schedule, {schedule_info}', size=18, pad=12)\n    \n    # Plot Learning Rates\n    for x, val in enumerate(lr_schedule):\n        if epochs <= 40 or x % 5 == 0 or x is epochs - 1:\n            if x < len(lr_schedule) - 1:\n                if lr_schedule[x - 1] < val:\n                    ha = 'right'\n                else:\n                    ha = 'left'\n            elif x == 0:\n                ha = 'right'\n            else:\n                ha = 'left'\n            plt.plot(x + 1, val, 'o', color='black');\n            offset_y = (max(lr_schedule) - min(lr_schedule)) * 0.02\n            plt.annotate(f'{val:.1E}', xy=(x + 1, val + offset_y), size=12, ha=ha)\n    \n    plt.xlabel('Epoch', size=16, labelpad=5)\n    plt.ylabel('Learning Rate', size=16, labelpad=5)\n    plt.grid()\n    plt.show()\n\n# Learning rate for encoder\nLR_SCHEDULE = [lrfn(step, num_warmup_steps=N_WARMUP_EPOCHS, lr_max=LR_MAX, num_cycles=0.50) for step in range(N_EPOCHS)]\n# Plot Learning Rate Schedule\nplot_lr_schedule(LR_SCHEDULE, epochs=N_EPOCHS)\n# Learning Rate Callback\nlr_callback = tf.keras.callbacks.LearningRateScheduler(lambda step: LR_SCHEDULE[step], verbose=0)","metadata":{"papermill":{"duration":0.885914,"end_time":"2023-07-02T11:34:14.533559","exception":false,"start_time":"2023-07-02T11:34:13.647645","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:24.590933Z","iopub.execute_input":"2023-07-31T07:09:24.591274Z","iopub.status.idle":"2023-07-31T07:09:24.981814Z","shell.execute_reply.started":"2023-07-31T07:09:24.591238Z","shell.execute_reply":"2023-07-31T07:09:24.980797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Weight Decay Callback","metadata":{"papermill":{"duration":0.038978,"end_time":"2023-07-02T11:34:14.612776","exception":false,"start_time":"2023-07-02T11:34:14.573798","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"\nThis code fragment defines a custom callback class called `WeightDecayCallback`, which will be used during the model training to update the \"weight decay\" term in the optimizer based on the learning rate.\n\nWeight decay, also known as \"L2 regularization,\" is a technique used in training machine learning models to prevent overfitting and improve model generalization.\n\nIn simple terms, weight decay involves adding a penalty term to the model's loss function that is related to the values of the model's weights. This penalty term is proportional to the square of the weight values and is added to the loss function during the optimization process.\n\nThe goal of weight decay is to penalize large model weights, meaning that the model is less likely to adjust its weights to extremely large values that could cause overfitting. By penalizing large weights, the model is encouraged to use smaller weights, which can improve generalization and prevent the model from memorizing the training data instead of learning general patterns.\n\nThe `WeightDecayCallback` callback has an `__init__` constructor that takes an optional parameter `wd_ratio`, representing the weight decay ratio with respect to the learning rate (default value is equal to `WD_RATIO`).\n\nThe callback has an `on_epoch_begin` method, which is executed at the beginning of each epoch during the model training. In this method, the weight decay in the optimizer is updated by multiplying the model's current learning rate by the `wd_ratio` value. This allows the weight decay to automatically adjust based on the learning rate, which can help improve model regularization and prevent overfitting.\n\nAdditionally, the `on_epoch_begin` method prints the current values of the learning rate and weight decay to the screen, providing useful information during training to monitor how these values are changing throughout the epochs.\nPlease note that the translation is based on the provided explanation and may require adjustments or additional context depending on the overall context of your work.","metadata":{}},{"cell_type":"code","source":"# Custom callback to update weight decay with learning rate\nclass WeightDecayCallback(tf.keras.callbacks.Callback):\n    def __init__(self, wd_ratio=WD_RATIO):\n        self.step_counter = 0\n        self.wd_ratio = wd_ratio\n    \n    def on_epoch_begin(self, epoch, logs=None):\n        model.optimizer.weight_decay = model.optimizer.learning_rate * self.wd_ratio\n        print(f'learning rate: {model.optimizer.learning_rate.numpy():.2e}, weight decay: {model.optimizer.weight_decay.numpy():.2e}')","metadata":{"papermill":{"duration":0.051216,"end_time":"2023-07-02T11:34:14.702691","exception":false,"start_time":"2023-07-02T11:34:14.651475","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:24.983413Z","iopub.execute_input":"2023-07-31T07:09:24.984496Z","iopub.status.idle":"2023-07-31T07:09:24.992171Z","shell.execute_reply.started":"2023-07-31T07:09:24.984459Z","shell.execute_reply":"2023-07-31T07:09:24.991051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Evaluate Initialized Model\n\nNow the evaluate the model with the training dataset","metadata":{"papermill":{"duration":0.038896,"end_time":"2023-07-02T11:34:14.781663","exception":false,"start_time":"2023-07-02T11:34:14.742767","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Evaluate Initialized Model On Validation Data\nmodel.evaluate(\n    val_dataset if USE_VAL else train_dataset,\n    steps=N_VAL_STEPS_PER_EPOCH if USE_VAL else TRAIN_STEPS_PER_EPOCH,\n    verbose=VERBOSE,\n)","metadata":{"papermill":{"duration":41.4984,"end_time":"2023-07-02T11:34:56.318989","exception":false,"start_time":"2023-07-02T11:34:14.820589","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:09:24.993695Z","iopub.execute_input":"2023-07-31T07:09:24.994115Z","iopub.status.idle":"2023-07-31T07:10:19.186185Z","shell.execute_reply.started":"2023-07-31T07:09:24.994078Z","shell.execute_reply":"2023-07-31T07:10:19.185063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Baseline\n\nThis code fragment calculates the \"baseline accuracy\" for the model when predicting only the padding token (`PAD_TOKEN`).\n\nIn this particular case, the \"baseline accuracy\" refers to the accuracy of the model when predicting only the padding token (`PAD_TOKEN`) for all samples in the dataset. This accuracy is calculated as the proportion of samples in the dataset that contain only the padding token.\n\nIt is important to calculate the \"baseline accuracy\" to get an idea of how a simple or naive model that always predicts the most common value (the padding token) would perform. If the model we are building does not significantly surpass this baseline accuracy, it indicates that the model is not learning relevant patterns in the data and needs improvement to be useful in the task at hand.\n\nIf the variable `USE_VAL` is true, it means that the validation set is being used to calculate the baseline accuracy. Otherwise, the training set is used.\n\nThe baseline accuracy is calculated by comparing the actual labels (`y_val` or `y_train`) with the value of the padding token (`PAD_TOKEN`). The padding token is used in sequences to indicate empty or irrelevant positions.\n\nThe calculation of the baseline accuracy is done as follows:\n\n- If `USE_VAL=True`, the baseline accuracy is calculated using the validation set: Each element in the validation set (`y_val`) is compared with the padding token value (`PAD_TOKEN`).\n- A boolean array is obtained that indicates whether each element is equal to the padding token or not.\n- The mean of this boolean array is calculated to obtain the proportion of elements that are equal to the padding token. This proportion represents the baseline accuracy in the validation set.\n- If `USE_VAL=False`, the same process is performed but using the training set (`y_train`) instead.\n\nA high baseline accuracy could indicate that the model needs improvement to be useful in the task being addressed.","metadata":{"papermill":{"duration":0.044168,"end_time":"2023-07-02T11:34:56.412094","exception":false,"start_time":"2023-07-02T11:34:56.367926","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# baseline accuracy when only pad token is predicted\nif USE_VAL:\n    baseline_accuracy = np.mean(y_val == PAD_TOKEN)\nelse:\n    baseline_accuracy = np.mean(y_train == PAD_TOKEN)\nprint(f'Baseline Accuracy: {baseline_accuracy:.4f}')","metadata":{"papermill":{"duration":0.056578,"end_time":"2023-07-02T11:34:56.511966","exception":false,"start_time":"2023-07-02T11:34:56.455388","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:10:19.188215Z","iopub.execute_input":"2023-07-31T07:10:19.188615Z","iopub.status.idle":"2023-07-31T07:10:19.200513Z","shell.execute_reply.started":"2023-07-31T07:10:19.188577Z","shell.execute_reply":"2023-07-31T07:10:19.199522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train","metadata":{"papermill":{"duration":0.039868,"end_time":"2023-07-02T11:34:56.591346","exception":false,"start_time":"2023-07-02T11:34:56.551478","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# We manually call `gc.collect()` to release unused memory.\ngc.collect()","metadata":{"papermill":{"duration":0.33862,"end_time":"2023-07-02T11:34:56.969069","exception":false,"start_time":"2023-07-02T11:34:56.630449","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:10:19.202232Z","iopub.execute_input":"2023-07-31T07:10:19.202644Z","iopub.status.idle":"2023-07-31T07:10:19.529663Z","shell.execute_reply.started":"2023-07-31T07:10:19.202611Z","shell.execute_reply":"2023-07-31T07:10:19.528699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We proceed to train the model using the `train_dataset`.\n\nThe steps involved in training the model are as follows:\n\n1. First, we clear all previous models from the GPU to free up memory. This is done using `tf.keras.backend.clear_session()`.\n\n2. We obtain a new model using the `get_model()` function. We print a summary of the model using `model.summary()` to verify that it has been created correctly.\n\n3. We start the training process of the model using the `fit()` method. In this method, the following arguments are provided:\n   - `x`: the training data (`train_dataset`).\n   - `steps_per_epoch`: The number of steps that will be performed in each training epoch. Each step is an update of the model's weights based on a batch of samples.\n   - `epochs`: The total number of epochs that will be performed during the entire training process.\n   - `validation_data`: the validation data (`val_dataset`). In our case, `USE_VAL=False`, so it is set to `None`.\n   - `validation_steps`: The number of steps that will be performed during evaluation in each validation epoch. We do not use this parameter.\n   - `callbacks`: A list of callbacks that are executed during training. In this case, we use the `lr_callback` to update the learning rate and the `WeightDecayCallback()` to update the weight decay during training.\n   - `verbose`: An integer value that controls the amount of information displayed during training. A value of 0 shows nothing, a value of 1 shows a progress bar, and a value of 2 shows a progress bar and a summary of each epoch.","metadata":{}},{"cell_type":"code","source":"print(f'TRAIN_STEPS_PER_EPOCH: {TRAIN_STEPS_PER_EPOCH}')\nprint(f'N_EPOCHS: {N_EPOCHS}')\nprint(f'VERBOSE: {VERBOSE}')","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:10:19.531548Z","iopub.execute_input":"2023-07-31T07:10:19.531914Z","iopub.status.idle":"2023-07-31T07:10:19.537336Z","shell.execute_reply.started":"2023-07-31T07:10:19.53188Z","shell.execute_reply":"2023-07-31T07:10:19.536381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if TRAIN_MODEL:\n    # Clear all models in GPU\n    tf.keras.backend.clear_session()\n\n    # Get new fresh model\n    model = get_model()\n\n    # Sanity Check\n    model.summary()\n    print('\\n\\n')\n    # Actual Training\n    history = model.fit(\n            x=train_dataset,\n            steps_per_epoch=TRAIN_STEPS_PER_EPOCH,\n            epochs=N_EPOCHS,\n            # Only used for validation data since training data is a generator\n            validation_data=val_dataset if USE_VAL else None,\n            validation_steps=N_VAL_STEPS_PER_EPOCH if USE_VAL else None,\n            callbacks=[\n                lr_callback,\n                WeightDecayCallback(),\n            ],\n            verbose = VERBOSE,\n        )","metadata":{"papermill":{"duration":12244.208248,"end_time":"2023-07-02T14:59:01.217198","exception":false,"start_time":"2023-07-02T11:34:57.00895","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:10:19.538798Z","iopub.execute_input":"2023-07-31T07:10:19.539381Z","iopub.status.idle":"2023-07-31T07:15:21.27165Z","shell.execute_reply.started":"2023-07-31T07:10:19.539348Z","shell.execute_reply":"2023-07-31T07:15:21.270619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load Weights\nif LOAD_WEIGHTS: # LOAD_WEIGHTS is set to False\n    model.load_weights('/kaggle/input/aslfr-training-python37/model.h5')\n    print(f'Successfully Loaded Pretrained Weights')","metadata":{"papermill":{"duration":0.069569,"end_time":"2023-07-02T14:59:01.347383","exception":false,"start_time":"2023-07-02T14:59:01.277814","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:15:21.273402Z","iopub.execute_input":"2023-07-31T07:15:21.274227Z","iopub.status.idle":"2023-07-31T07:15:21.279287Z","shell.execute_reply.started":"2023-07-31T07:15:21.27419Z","shell.execute_reply":"2023-07-31T07:15:21.278322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save Model Weights\nmodel.save_weights('model.h5')","metadata":{"papermill":{"duration":0.206194,"end_time":"2023-07-02T14:59:01.616525","exception":false,"start_time":"2023-07-02T14:59:01.410331","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:15:21.28063Z","iopub.execute_input":"2023-07-31T07:15:21.281086Z","iopub.status.idle":"2023-07-31T07:15:21.423192Z","shell.execute_reply.started":"2023-07-31T07:15:21.281047Z","shell.execute_reply":"2023-07-31T07:15:21.421933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we evaluate the loaded model on the specified dataset and calculate the loss and metrics of the model on that dataset. This step helps us verify that the model is loaded correctly and ready for use. Since `train_dataset` is a generator object, each time we access it, we obtain a new batch of different data. Therefore, we are not evaluating with the same data that we used for training.\n\nThe `evaluate()` function prints three values to the screen in the form of a list:\n\n- Loss (Sparse Categorical Cross Entropy with Label Smoothing in our case): This measures how well the model's predictions fit the true labels during training. A lower loss indicates a better fit of the model to the training data.\n\n- Top-1 Accuracy (TopKAccuracy(1)): This is the top-1 accuracy, also known as the accuracy in the ranking of the highest probability. It represents the fraction of samples in which the model correctly predicts the true class as the class with the highest probability. In other words, it is the accuracy for the case where only the most probable class is taken as the prediction.\n\n- Top-5 Accuracy (TopKAccuracy(5)): This is the top-5 accuracy, which represents the fraction of samples in which the model correctly predicts the true class as one of the five classes with the highest probabilities. In other words, it considers the top five most probable classes and checks if the true class is present in those five classes.","metadata":{"execution":{"iopub.status.busy":"2023-07-26T17:48:41.772628Z","iopub.execute_input":"2023-07-26T17:48:41.77303Z","iopub.status.idle":"2023-07-26T17:48:42.5113Z","shell.execute_reply.started":"2023-07-26T17:48:41.772989Z","shell.execute_reply":"2023-07-26T17:48:42.509845Z"}}},{"cell_type":"code","source":"# Verify Model is Loaded Correctly\nmodel.evaluate(\n    val_dataset if USE_VAL else train_dataset,\n    steps=N_VAL_STEPS_PER_EPOCH if USE_VAL else TRAIN_STEPS_PER_EPOCH,\n    batch_size=BATCH_SIZE,\n    verbose=VERBOSE,\n)","metadata":{"papermill":{"duration":41.357991,"end_time":"2023-07-02T14:59:43.031109","exception":false,"start_time":"2023-07-02T14:59:01.673118","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:15:21.424841Z","iopub.execute_input":"2023-07-31T07:15:21.4252Z","iopub.status.idle":"2023-07-31T07:16:45.463159Z","shell.execute_reply.started":"2023-07-31T07:15:21.425165Z","shell.execute_reply":"2023-07-31T07:16:45.46192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Levenshtein Distance","metadata":{"papermill":{"duration":0.056044,"end_time":"2023-07-02T14:59:43.142922","exception":false,"start_time":"2023-07-02T14:59:43.086878","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"We define a function called `outputs2phrase`, which is used to convert the model outputs (`outputs`) into a sequence of text (phrase). The function takes the model outputs as input, which can be a two-dimensional tensor or a one-dimensional tensor. The function performs the following steps:\n\n1. If the outputs have two dimensions, it takes the index of the class with the highest probability for each sample using `np.argmax(outputs, axis=1)`. This assumes that the outputs are class probabilities, and the model has predicted the class with the highest probability for each sample.\n\n2. Then, the function iterates over the indices of the predicted classes and uses a dictionary called `ORD2CHAR` to convert the indices into characters. `ORD2CHAR` is a dictionary that maps class indices to the corresponding characters in the text sequence. For example, index 0 could be mapped to the character \"a\", index 1 to the character \"b\", and so on.\n\n3. Finally, the function concatenates all the characters to form the complete text sequence and returns it as the output.\n\nIn summary, this function is useful for converting the model outputs, which are class indices or probabilities, into a readable text sequence to interpret the model predictions in text form.","metadata":{}},{"cell_type":"code","source":"# Output Predictions to string\ndef outputs2phrase(outputs):\n    if outputs.ndim == 2:\n        outputs = np.argmax(outputs, axis=1)\n    \n    return ''.join([ORD2CHAR.get(s, '') for s in outputs])","metadata":{"papermill":{"duration":0.066802,"end_time":"2023-07-02T14:59:43.266519","exception":false,"start_time":"2023-07-02T14:59:43.199717","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:16:45.464895Z","iopub.execute_input":"2023-07-31T07:16:45.465364Z","iopub.status.idle":"2023-07-31T07:16:45.471488Z","shell.execute_reply.started":"2023-07-31T07:16:45.465318Z","shell.execute_reply":"2023-07-31T07:16:45.470537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We define a function called `predict_phrase` that is used to generate text predictions based on an input sequence of frames using a pre-trained language model.\n\nThe steps involved in the `predict_phrase` function are as follows:\n\n1. The function is decorated with `@tf.function()` to convert it into a TensorFlow graph and gain performance benefits during execution.\n\n2. The function takes the `frames` tensor as input, which represents the input sequence of frames (similar to what we've seen previously as `X_batch['frames']`).\n\n3. A batch dimension is added to the `frames` tensor using `tf.expand_dims()`. This is necessary because the model expects input with a batch dimension, even if we are making a prediction for a single example.\n\n4. A variable called `phrase` is initialized with shape `[1, MAX_PHRASE_LENGTH]` and filled with the padding token (`PAD_TOKEN`). This variable is used to store the generated text sequence.\n\n5. A loop is initiated that will run `MAX_PHRASE_LENGTH` times, which is the maximum number of tokens in the generated text sequence.\n\n6. Inside the loop, the following steps are performed:\n\n   * The `phrase` tensor is converted to `int8` data type.\n   * The model is called with inputs `X = frames` and `y = phrase` using the `model()` method to get the model predictions for the next token in the text sequence. These predictions are stored in the variable `outputs`.\n   * `phrase = tf.cast(phrase, tf.int32)`: The `phrase` tensor is converted to data type `int32`. This is necessary to ensure that we can use the `tf.where()` function correctly as it requires the arguments to have the same data type.\n   * `tf.range(MAX_PHRASE_LENGTH) < idx + 1`: A boolean tensor of shape `[MAX_PHRASE_LENGTH]=32` is created, where each element is `True` if its index is less than `idx + 1`, and `False` otherwise. This creates a mask that indicates which positions in the text sequence have not been predicted yet.\n   * `tf.argmax(outputs, axis=2, output_type=tf.int32)`: The `tf.argmax()` function is used to find the index of the token with the highest probability in the predictions (`outputs`). The argument `axis=2` indicates that the maximum search will be done along the third axis of `outputs`, which corresponds to the different classes or tokens possible in the text sequence. The argument `output_type=tf.int32` ensures that the result of `tf.argmax()` will be of integer data type `int32`.\n   * `tf.where(condition, x, y)`: This function performs a \"conditional selection\" operation. It takes three arguments: `condition`, `x`, and `y`. If `condition` is `True` at a position, the value of `x` at that position is selected; otherwise, the value of `y` is selected. In this case, `condition` is the mask created in step 2, `x` is the result of `tf.argmax()` in step 3 (i.e., the index of the predicted token), and `y` is the current text sequence (`phrase`). This means that if a position in the mask is `True`, the predicted token at that position is selected; otherwise, the current token is retained at that position.\n\n   * At the end of this operation, the `phrase` variable has been updated with the predicted token at the corresponding position, allowing the loop to iterate and add the next token to the next position in the text sequence. This way, the text sequence is gradually constructed step by step until it is completed with `MAX_PHRASE_LENGTH` tokens, and the generated text sequence is obtained from the language model.\n\n7. Once the loop is finished, the `phrase` tensor is squeezed to remove the previously added batch dimension, as we are only generating a text sequence for a single example.\n\n8. The numeric values in `phrase` are converted into a \"one-hot\" representation using `tf.one_hot()`. This \"one-hot\" representation is useful for obtaining the discrete labels (tokens) in numerical format.\n\n9. Finally, the function returns a dictionary with the key \"outputs\" and the resulting output tensor. This tensor contains the generated text sequence in \"one-hot\" format with shape `[MAX_PHRASE_LENGTH, N_UNIQUE_CHARACTERS]`.\n\nIn summary, the `predict_phrase` function takes an input sequence of frames and uses the language model to predict the next word in the generated text sequence. It then iterates to predict all the words in the sequence. The function returns the generated text sequence in \"one-hot\" format to represent the tokens numerically.\n\nPlease note that the specific details and the mappings for tokens in the `ORD2CHAR` dictionary will depend on your specific use case and the data you are working with.\n","metadata":{}},{"cell_type":"code","source":"print(tf.fill([1,MAX_PHRASE_LENGTH], PAD_TOKEN)) # MAX_PHRASE_LENGTH = 32\nprint(tf.cast(tf.fill([1,MAX_PHRASE_LENGTH], PAD_TOKEN), tf.int8))","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:16:45.473159Z","iopub.execute_input":"2023-07-31T07:16:45.473909Z","iopub.status.idle":"2023-07-31T07:16:45.494578Z","shell.execute_reply.started":"2023-07-31T07:16:45.473874Z","shell.execute_reply":"2023-07-31T07:16:45.493414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"@tf.function()\ndef predict_phrase(frames):\n    # Add Batch Dimension\n    frames = tf.expand_dims(frames, axis=0)\n    # Start Phrase\n    phrase = tf.fill([1,MAX_PHRASE_LENGTH], PAD_TOKEN)\n\n    for idx in tf.range(MAX_PHRASE_LENGTH):\n        # Cast phrase to int8\n        phrase = tf.cast(phrase, tf.int8)\n        # Predict Next Token\n        outputs = model({\n            'frames': frames,\n            'phrase': phrase,\n        })\n\n        # Add predicted token to input phrase\n        phrase = tf.cast(phrase, tf.int32)\n        phrase = tf.where( # where its True search for max prob, if its False keep Pad token\n            tf.range(MAX_PHRASE_LENGTH) < idx + 1, # create a mask of Trues and Falses\n            tf.argmax(outputs, axis=2, output_type=tf.int32), # search for the max probability\n            phrase,\n        )\n\n    # Squeeze outputs\n    outputs = tf.squeeze(phrase, axis=0) # drop first dimension\n    outputs = tf.one_hot(outputs, N_UNIQUE_CHARACTERS) # one-hot encoding of the numbers\n\n    # Return a dictionary with the output tensor\n    return outputs","metadata":{"papermill":{"duration":0.068395,"end_time":"2023-07-02T14:59:43.392955","exception":false,"start_time":"2023-07-02T14:59:43.32456","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:16:45.496281Z","iopub.execute_input":"2023-07-31T07:16:45.496628Z","iopub.status.idle":"2023-07-31T07:16:45.506783Z","shell.execute_reply.started":"2023-07-31T07:16:45.496597Z","shell.execute_reply":"2023-07-31T07:16:45.505783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Levenstein Distance Train\n\nWe define a function called `get_ld_train` that is used to calculate the Levenshtein distance between the real phrases and the phrases predicted by the model on the training set.\n\nThe function takes as input the training data `X_train` and the real labels `y_train`, which contain sequences of frames and phrases represented by integer indices, respectively.\n\nWe create an empty list called `LD_TRAIN` that will be used to store the Levenshtein distances for each pair of real and predicted phrases.\n\nThen, the function iterates through the training data using a `for` loop. For each pair of frames and phrase, it does the following:\n\n1. It uses the `predict_phrase` function to predict the phrase from the input frames. The `predict_phrase` function takes the frames as input and returns the predicted sequence of characters in \"one-hot\" format. It then uses the `outputs2phrase` function to convert the predicted sequence of characters into a readable string.\n\n2. It converts the real phrase from integer indices to a readable string using the `outputs2phrase` function.\n\n3. It calculates the Levenshtein distance between the real phrase and the predicted phrase using the `levenshtein` function. The Levenshtein distance is a measure of the difference between two strings, which is calculated by counting the minimum number of operations (insertions, deletions, or substitutions of characters) required to transform one string into the other.\n\n4. It appends the real phrase, the predicted phrase, and the Levenshtein distance to the `LD_TRAIN` list.\n\nAfter completing the loop, the function converts the `LD_TRAIN` list into a pandas DataFrame named `LD_TRAIN_DF` and returns it as the result.","metadata":{"papermill":{"duration":0.056358,"end_time":"2023-07-02T14:59:43.504839","exception":false,"start_time":"2023-07-02T14:59:43.448481","status":"completed"},"tags":[]}},{"cell_type":"code","source":"print(f'Shape X_train: {X_train.shape}')\nprint(f'Shape y_train: {y_train.shape}')","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:16:45.508194Z","iopub.execute_input":"2023-07-31T07:16:45.50858Z","iopub.status.idle":"2023-07-31T07:16:45.520109Z","shell.execute_reply.started":"2023-07-31T07:16:45.508548Z","shell.execute_reply":"2023-07-31T07:16:45.518887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In each iteration of the loop, `zip(tqdm(X_train, total=N))` pairs an element from `X_train` with its corresponding element from `tqdm`, a progress-tracking function with a visual progress bar. Consequently, we obtain 3 elements in the `enumerate` function: the index, the element from `tqdm`, and the element from `X_train` and `y_train`.","metadata":{}},{"cell_type":"code","source":"i = 0\nN = 100 if IS_INTERACTIVE else 1000\nfor idx, (frames, phrase_true) in enumerate(zip(tqdm(X_train, total=N), y_train)):\n    print(idx)\n    print(phrase_true)\n    i+=1\n    if i == 5:\n        break","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:16:45.521678Z","iopub.execute_input":"2023-07-31T07:16:45.522392Z","iopub.status.idle":"2023-07-31T07:16:45.55297Z","shell.execute_reply.started":"2023-07-31T07:16:45.522358Z","shell.execute_reply":"2023-07-31T07:16:45.552068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Compute Levenstein Distances\ndef get_ld_train():\n    N = 100 if IS_INTERACTIVE else 1000\n    LD_TRAIN = []\n    for idx, (frames, phrase_true) in enumerate(zip(tqdm(X_train, total=N), y_train)):\n        # Predict Phrase and Convert to String\n        phrase_pred = predict_phrase(frames).numpy()\n        phrase_pred = outputs2phrase(phrase_pred)\n        # True Phrase Ordinal to String\n        phrase_true = outputs2phrase(phrase_true)\n        # Add Levenstein Distance\n        LD_TRAIN.append({\n            'phrase_true': phrase_true,\n            'phrase_pred': phrase_pred,\n            'levenshtein_distance': levenshtein(phrase_pred, phrase_true),\n        })\n        # Take subset in interactive mode\n        if idx == N:\n            break\n            \n    # Convert to DataFrame\n    LD_TRAIN_DF = pd.DataFrame(LD_TRAIN)\n    \n    return LD_TRAIN_DF","metadata":{"papermill":{"duration":0.071676,"end_time":"2023-07-02T14:59:43.633777","exception":false,"start_time":"2023-07-02T14:59:43.562101","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:16:45.55434Z","iopub.execute_input":"2023-07-31T07:16:45.555264Z","iopub.status.idle":"2023-07-31T07:16:45.563456Z","shell.execute_reply.started":"2023-07-31T07:16:45.555231Z","shell.execute_reply":"2023-07-31T07:16:45.562456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"LD_TRAIN_DF = get_ld_train()\n\n# add column to see the length of the true phrase\nLD_TRAIN_DF['len_char'] = LD_TRAIN_DF['phrase_true'].apply(lambda x: len(x))\n\n# Display Errors\ndisplay(LD_TRAIN_DF.head(30))","metadata":{"papermill":{"duration":126.418133,"end_time":"2023-07-02T15:01:50.108598","exception":false,"start_time":"2023-07-02T14:59:43.690465","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:16:45.564884Z","iopub.execute_input":"2023-07-31T07:16:45.565693Z","iopub.status.idle":"2023-07-31T07:16:58.961425Z","shell.execute_reply.started":"2023-07-31T07:16:45.565656Z","shell.execute_reply":"2023-07-31T07:16:58.959677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we calculate the distribution of Levenshtein distances in the training set and create a bar chart to visualize how these distances are distributed in the dataset.","metadata":{}},{"cell_type":"code","source":"# Value Counts\nLD_TRAIN_VC = dict([(i, 0) for i in range(LD_TRAIN_DF['levenshtein_distance'].max()+1)])\nfor ld in LD_TRAIN_DF['levenshtein_distance']:\n    LD_TRAIN_VC[ld] += 1\n\nplt.figure(figsize=(15,8))\npd.Series(LD_TRAIN_VC).plot(kind='bar', width=1)\nplt.title(f'Train Levenstein Distance Distribution | Mean: {LD_TRAIN_DF.levenshtein_distance.mean():.4f}')\nplt.xlabel('Levenstein Distance')\nplt.ylabel('Sample Count')\nplt.xlim(-0.50, LD_TRAIN_DF.levenshtein_distance.max()+0.50)\nplt.grid(axis='y')\nplt.savefig('temp.png')\nplt.show()","metadata":{"papermill":{"duration":0.724484,"end_time":"2023-07-02T15:01:50.890212","exception":false,"start_time":"2023-07-02T15:01:50.165728","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:16:58.963178Z","iopub.execute_input":"2023-07-31T07:16:58.963546Z","iopub.status.idle":"2023-07-31T07:16:59.622211Z","shell.execute_reply.started":"2023-07-31T07:16:58.963511Z","shell.execute_reply":"2023-07-31T07:16:59.621145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Levenstein Distance Evaluation\n\nWe do the same now with the validation dataset:","metadata":{"papermill":{"duration":0.057603,"end_time":"2023-07-02T15:01:51.005493","exception":false,"start_time":"2023-07-02T15:01:50.94789","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Compute Levenstein Distances\ndef get_ld_val():\n    N = 100 if IS_INTERACTIVE else 1000\n    LD_VAL = []\n    for idx, (frames, phrase_true) in enumerate(zip(tqdm(X_val, total=N), y_val)):\n        # Predict Phrase and Convert to String\n        phrase_pred = predict_phrase(frames).numpy()\n        phrase_pred = outputs2phrase(phrase_pred)\n        # True Phrase Ordinal to String\n        phrase_true = outputs2phrase(phrase_true)\n        # Add Levenstein Distance\n        LD_VAL.append({\n            'phrase_true': phrase_true,\n            'phrase_pred': phrase_pred,\n            'levenshtein_distance': levenshtein(phrase_pred, phrase_true),\n        })\n        # Take subset in interactive mode\n        if idx == N:\n            break\n            \n    # Convert to DataFrame\n    LD_VAL_DF = pd.DataFrame(LD_VAL)\n    \n    return LD_VAL_DF","metadata":{"papermill":{"duration":0.068889,"end_time":"2023-07-02T15:01:51.131182","exception":false,"start_time":"2023-07-02T15:01:51.062293","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:16:59.623849Z","iopub.execute_input":"2023-07-31T07:16:59.624239Z","iopub.status.idle":"2023-07-31T07:16:59.632942Z","shell.execute_reply.started":"2023-07-31T07:16:59.624203Z","shell.execute_reply":"2023-07-31T07:16:59.632037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if USE_VAL:\n    LD_VAL_DF = get_ld_val()\n\n    # Display Errors\n    display(LD_VAL_DF.head(30))","metadata":{"papermill":{"duration":0.064943,"end_time":"2023-07-02T15:01:51.25391","exception":false,"start_time":"2023-07-02T15:01:51.188967","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:16:59.634434Z","iopub.execute_input":"2023-07-31T07:16:59.635275Z","iopub.status.idle":"2023-07-31T07:16:59.647936Z","shell.execute_reply.started":"2023-07-31T07:16:59.635249Z","shell.execute_reply":"2023-07-31T07:16:59.646749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Value Counts\nif USE_VAL:\n    LD_VAL_VC = dict([(i, 0) for i in range(LD_VAL_DF['levenshtein_distance'].max()+1)])\n    for ld in LD_VAL_DF['levenshtein_distance']:\n        LD_VAL_VC[ld] += 1\n\n    plt.figure(figsize=(15,8))\n    pd.Series(LD_VAL_VC).plot(kind='bar', width=1)\n    plt.title(f'Validation Levenstein Distance Distribution | Mean: {LD_VAL_DF.levenshtein_distance.mean():.4f}')\n    plt.xlabel('Levenstein Distance')\n    plt.ylabel('Sample Count')\n    plt.xlim(0-0.50, LD_VAL_DF.levenshtein_distance.max()+0.50)\n    plt.grid(axis='y')\n    plt.savefig('temp.png')\n    plt.show()","metadata":{"papermill":{"duration":0.068365,"end_time":"2023-07-02T15:01:51.379464","exception":false,"start_time":"2023-07-02T15:01:51.311099","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:16:59.649325Z","iopub.execute_input":"2023-07-31T07:16:59.649904Z","iopub.status.idle":"2023-07-31T07:16:59.660448Z","shell.execute_reply.started":"2023-07-31T07:16:59.64987Z","shell.execute_reply":"2023-07-31T07:16:59.659457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training History\n\nOnce the model is trained and predictions are created for our data, a function called `plot_history_metric` is defined to plot the evolution of different metrics used in the training.\n\nLet's go step by step to understand what each part of the code does:\n\n- `if not TRAIN_MODEL: return`: This line checks if the variable `TRAIN_MODEL` is `True` or `False`. If it is `False`, the function will not take any action and will stop here with `return`.\n\n- `values = history.history[metric]`: It retrieves the values of the specified metric (`metric`) from the training history of the model, which is stored in the `history.history` dictionary.\n\n- `N_EPOCHS = len(values)`: It obtains the total number of training epochs by counting the length of the metric values.\n\n- `val = 'val' in ''.join(history.history.keys())`: It checks if the word 'val' is present in the keys of the model's history (if we used validation). If so, it sets `val=True`, otherwise, it sets it to `False`.\n\n- `if N_EPOCHS <= 20: x = np.arange(1, N_EPOCHS + 1)`: If the number of epochs is less than or equal to 20, it creates an arange `x` with values ranging from 1 to the total number of epochs.\n\n- `else: x = [1, 5] + [10 + 5 * idx for idx in range((N_EPOCHS - 10) // 5 + 1)]`: If the number of epochs is greater than 20, it creates a list `x` with specific values spaced to clearly show the evolution of the metric in the graph.\n\n- `x_ticks = np.arange(1, N_EPOCHS+1)`: Another arange `x_ticks` is created, containing values from 1 to the total number of epochs. This is used to define the x-axis labels in the graph.\n\n- If `val=True`:\n    - `val_values = history.history[f'val_{metric}']`: It retrieves the values of the specific validation metric (`val_{metric}`) from the model's history.\n    - `val_argmin = f_best(val_values)`: It finds the index of the validation metric that has the best value, using the function `f_best` specified as an argument (default is `np.argmax`, which finds the index with the maximum value).\n    - `plt.plot(x_ticks, val_values, label=f'val')`: It plots the evolution of the validation metric on the graph with the label 'val'.\n\n- In case of `val=False`:\n    - `plt.plot(x_ticks, values, label=f'train')`: It plots the evolution of the training metric on the graph with the label 'train'.\n    - `argmin = f_best(values)`: It finds the index of the training metric that has the best value, using the function `f_best` specified as an argument (default is `np.argmax`, which finds the index with the maximum value).\n    - `plt.scatter(argmin + 1, values[argmin], color='red', s=75, marker='o', label=f'train_best')`: It adds a red point on the graph indicating the best training metric and its corresponding value.\n\n- If there are validation metrics (`val`):\n    - `plt.scatter(val_argmin + 1, val_values[val_argmin], color='purple', s=75, marker='o', label=f'val_best')`: If there are validation metrics, it adds a purple point on the graph indicating the best validation metric and its corresponding value.","metadata":{"papermill":{"duration":0.056463,"end_time":"2023-07-02T15:01:51.49247","exception":false,"start_time":"2023-07-02T15:01:51.436007","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def plot_history_metric(metric, f_best=np.argmax, ylim=None, yscale=None, yticks=None):\n    # Only plot when training\n    if not TRAIN_MODEL:\n        return\n    \n    plt.figure(figsize=(20, 10))\n    \n    values = history.history[metric]\n    N_EPOCHS = len(values)\n    val = 'val' in ''.join(history.history.keys())\n    # Epoch Ticks\n    if N_EPOCHS <= 20:\n        x = np.arange(1, N_EPOCHS + 1)\n    else:\n        x = [1, 5] + [10 + 5 * idx for idx in range((N_EPOCHS - 10) // 5 + 1)]\n\n    x_ticks = np.arange(1, N_EPOCHS+1)\n\n    # Validation\n    if val:\n        val_values = history.history[f'val_{metric}']\n        val_argmin = f_best(val_values)\n        plt.plot(x_ticks, val_values, label=f'val')\n\n    # summarize history for accuracy\n    plt.plot(x_ticks, values, label=f'train')\n    argmin = f_best(values)\n    plt.scatter(argmin + 1, values[argmin], color='red', s=75, marker='o', label=f'train_best')\n    if val:\n        plt.scatter(val_argmin + 1, val_values[val_argmin], color='purple', s=75, marker='o', label=f'val_best')\n\n    plt.title(f'Model {metric}', fontsize=24, pad=10)\n    plt.ylabel(metric, fontsize=20, labelpad=10)\n\n    if ylim:\n        plt.ylim(ylim)\n\n    if yscale is not None:\n        plt.yscale(yscale)\n        \n    if yticks is not None:\n        plt.yticks(yticks, fontsize=16)\n\n    plt.xlabel('epoch', fontsize=20, labelpad=10)        \n    plt.tick_params(axis='x', labelsize=8)\n    plt.xticks(x, fontsize=16) # set tick step to 1 and let x axis start at 1\n    plt.yticks(fontsize=16)\n    \n    plt.legend(prop={'size': 10})\n    plt.grid()\n    plt.show()","metadata":{"papermill":{"duration":0.072449,"end_time":"2023-07-02T15:01:51.621731","exception":false,"start_time":"2023-07-02T15:01:51.549282","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:16:59.663469Z","iopub.execute_input":"2023-07-31T07:16:59.664264Z","iopub.status.idle":"2023-07-31T07:16:59.677248Z","shell.execute_reply.started":"2023-07-31T07:16:59.664227Z","shell.execute_reply":"2023-07-31T07:16:59.676281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history.history.keys() # there's not 'val'","metadata":{"execution":{"iopub.status.busy":"2023-07-31T07:16:59.679813Z","iopub.execute_input":"2023-07-31T07:16:59.680274Z","iopub.status.idle":"2023-07-31T07:16:59.694511Z","shell.execute_reply.started":"2023-07-31T07:16:59.680242Z","shell.execute_reply":"2023-07-31T07:16:59.693395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_history_metric('loss', f_best=np.argmin)","metadata":{"papermill":{"duration":0.534279,"end_time":"2023-07-02T15:01:52.213146","exception":false,"start_time":"2023-07-02T15:01:51.678867","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:16:59.696072Z","iopub.execute_input":"2023-07-31T07:16:59.696403Z","iopub.status.idle":"2023-07-31T07:17:00.102977Z","shell.execute_reply.started":"2023-07-31T07:16:59.696371Z","shell.execute_reply":"2023-07-31T07:17:00.1019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_history_metric('top1acc', ylim=[0,1], yticks=np.arange(0.0, 1.1, 0.1))","metadata":{"papermill":{"duration":0.542072,"end_time":"2023-07-02T15:01:52.812958","exception":false,"start_time":"2023-07-02T15:01:52.270886","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:17:00.104586Z","iopub.execute_input":"2023-07-31T07:17:00.104955Z","iopub.status.idle":"2023-07-31T07:17:00.482576Z","shell.execute_reply.started":"2023-07-31T07:17:00.104922Z","shell.execute_reply":"2023-07-31T07:17:00.481641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_history_metric('top5acc', ylim=[0,1], yticks=np.arange(0.0, 1.1, 0.1))","metadata":{"papermill":{"duration":0.542339,"end_time":"2023-07-02T15:01:53.417063","exception":false,"start_time":"2023-07-02T15:01:52.874724","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:17:00.484233Z","iopub.execute_input":"2023-07-31T07:17:00.484936Z","iopub.status.idle":"2023-07-31T07:17:00.863866Z","shell.execute_reply.started":"2023-07-31T07:17:00.4849Z","shell.execute_reply":"2023-07-31T07:17:00.862668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Inference\n\nThe following code defines the `TFLiteModel` class, which represents a TensorFlow Lite (TFLite) model constructed from the previously trained model. This is useful when deploying the model on devices with limited resources, such as mobile phones, IoT devices, or embedded systems. TFLite is an optimized version of TensorFlow designed for fast and efficient inference on devices with limited computational capabilities.\n\n### Paso a paso:\n\n1. `TFLiteModel(tf.Module)`: Defines a class named `TFLiteModel` that inherits from `tf.Module`, indicating that it is a trainable entity in TensorFlow.\n\n2. `def __init__(self, model)`: The `__init__` method is the constructor of the class and is automatically executed when creating an instance of `TFLiteModel`. It takes the argument `model`, which is the previously trained model that will be used as the base for the TFLite model.\n\n3. `self.preprocess_layer = preprocess_layer`: The `self.preprocess_layer` is assigned the preprocessing layer used to process the input data before passing it to the base model.\n\n4. `self.model = model`: The `self.model` is assigned the previously trained model that will be used as the base for the TFLite model.\n\n5. `@tf.function(jit_compile=True)`: The following methods are decorated with `@tf.function`, indicating that they will be compiled by TensorFlow for faster performance.\n\n6. `def encoder(self, x, frames_inp)`: This method implements the part of the model corresponding to the encoder. Similar to when constructing the model, the encoder takes two arguments; `x`, which represents the input data, and `frames_inp`, which is the shape of the data. In this function, embedding and then encoding is performed.\n\n7. `def decoder(self, x, phrase_inp)`: This method implements the part of the model corresponding to the decoder. Similar to when constructing the model, the decoder takes two arguments; `x`, which represents the output of the encoder, and `phrase_inp`, which is the shape of the \"phrase\" data.\n\n8. `@tf.function(input_signature=[tf.TensorSpec(shape=[None, N_COLS0], dtype=tf.float32, name='inputs')])`: The `__call__` method is defined as a TensorFlow function that will be used for inference with the TFLite model. The shape and type of the input data are specified using `tf.TensorSpec`.\n\n9. `N_INPUT_FRAMES = tf.shape(inputs)[0]`: The number of rows in the input `inputs` is obtained, representing the number of frames.\n\n10. `frames_inp = self.preprocess_layer(inputs)`: The input data is processed using the preprocessing layer `preprocess_layer`.\n\n11. `frames_inp = tf.expand_dims(frames_inp, axis=0)`: A batch dimension is added to `frames_inp` to match the input of the original model.\n\n12. `encoding = self.encoder(frames_inp, frames_inp)`: The encoded representation of the input data is obtained using the `encoder` method defined earlier.\n\n13. `phrase = tf.fill([1,MAX_PHRASE_LENGTH], PAD_TOKEN)`: Similarly to before, a tensor of shape `[1, MAX_PHRASE_LENGTH]` filled with the value `PAD_TOKEN` is created. This tensor represents the output phrase that will be constructed during inference.\n\n14. `for idx in tf.range(MAX_PHRASE_LENGTH)`: It iterates over the range of indices of the maximum length of the phrase (`MAX_PHRASE_LENGTH=32`).\n\n15. `phrase = tf.cast(phrase, tf.int8)`: The tensor `phrase` is cast to type `int8`.\n\n16. `outputs = tf.cond(stop, ...)`: A TensorFlow conditional (`tf.cond`) is used to determine if the generation of the phrase should stop. If `stop=True`, which means the end of sentence token (`END_TOKEN`) was predicted, the tensor `phrase` converted to one-hot encoding is returned, and the generation of the phrase stops. If `stop=False`, the next word of the phrase is generated using the `decoder` method.\n\n17. `phrase = tf.where(tf.range(MAX_PHRASE_LENGTH) < idx + 1, ...)`: The `phrase` tensor is updated by replacing the padding token (`PAD_TOKEN`) with the predicted word up to the current index (`idx`) in the phrase generation.\n\n18. `predicted_token = phrase[0,idx]`: The predicted token at position `idx` is obtained.\n\n19. `if not stop: stop = predicted_token == END_TOKEN`: The `stop` variable is updated by checking if the predicted token is equal to the end of sentence token (`END_TOKEN`). If so, `stop` is set to true, indicating that phrase generation should stop.\n\n20. `outputs = tf.squeeze(phrase, axis=0)`: The added batch dimension is removed to obtain the complete generated phrase.\n\n21. `outputs = tf.one_hot(outputs, N_UNIQUE_CHARACTERS)`: The generated phrase is converted to one-hot encoding.\n\n22. Finally, the `__call__` method returns a dictionary with the key `'outputs'` containing the generated phrase.\n\nAfter defining the `TFLiteModel` class, an instance of it named `tflite_keras_model` is created, and a demonstration of inference is performed using input data `demo_raw_data`. Inference is done by calling the `__call__` method of `tflite_keras_model`, and the generated phrase, the true phrase, and the input data in tensor format are displayed.","metadata":{"papermill":{"duration":0.059194,"end_time":"2023-07-02T15:01:53.537697","exception":false,"start_time":"2023-07-02T15:01:53.478503","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Model Layer Names\nfor l in model.layers:\n    print(l.name)","metadata":{"papermill":{"duration":0.070811,"end_time":"2023-07-02T15:01:53.667988","exception":false,"start_time":"2023-07-02T15:01:53.597177","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:17:00.866075Z","iopub.execute_input":"2023-07-31T07:17:00.866633Z","iopub.status.idle":"2023-07-31T07:17:00.876306Z","shell.execute_reply.started":"2023-07-31T07:17:00.866593Z","shell.execute_reply":"2023-07-31T07:17:00.875319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# TFLite model for submission\nclass TFLiteModel(tf.Module):\n    def __init__(self, model):\n        super(TFLiteModel, self).__init__()\n\n        # Load the feature generation and main models\n        self.preprocess_layer = preprocess_layer\n        self.model = model\n    \n    @tf.function(jit_compile=True)\n    def encoder(self, x, frames_inp):\n        x = self.model.get_layer('embedding')(x)\n        x = self.model.get_layer('encoder')(x, frames_inp)\n        \n        return x\n        \n    @tf.function(jit_compile=True)\n    def decoder(self, x, phrase_inp, frames_inp):\n        x = self.model.get_layer('decoder')(x, phrase_inp, frames_inp)\n        x = self.model.get_layer('classifier')(x)\n        \n        return x\n    \n    @tf.function(input_signature=[tf.TensorSpec(shape=[None, N_COLS0], dtype=tf.float32, name='inputs')])\n    def __call__(self, inputs):\n        # Number Of Input Frames\n        N_INPUT_FRAMES = tf.shape(inputs)[0]\n        # Preprocess Data\n        frames_inp = self.preprocess_layer(inputs)        \n        # Add Batch Dimension\n        frames_inp = tf.expand_dims(frames_inp, axis=0)\n        # Get Encoding\n        encoding = self.encoder(frames_inp, frames_inp)\n        # Make Prediction\n        phrase = tf.fill([1,MAX_PHRASE_LENGTH], PAD_TOKEN)\n        # Predict One Token At A Time\n        stop = False\n        for idx in tf.range(MAX_PHRASE_LENGTH):\n            # Cast phrase to int8\n            phrase = tf.cast(phrase, tf.int8)\n            # If EOS token is predicted, stop predicting\n            outputs = tf.cond(\n                stop,\n                lambda: tf.one_hot(tf.cast(phrase, tf.int32), N_UNIQUE_CHARACTERS),\n                lambda: self.decoder(encoding, phrase, frames_inp)\n            )\n            # Add predicted token to input phrase\n            phrase = tf.cast(phrase, tf.int32)\n            # Replcae PAD token with predicted token up to idx\n            phrase = tf.where(\n                tf.range(MAX_PHRASE_LENGTH) < idx + 1,\n                tf.argmax(outputs, axis=2, output_type=tf.int32),\n                phrase,\n            )\n            # Predicted Token\n            predicted_token = phrase[0,idx]\n            # If EOS (End Of Sentence) token is predicted stop\n            if not stop:\n                stop = predicted_token == END_TOKEN\n            \n        # Squeeze outputs\n        outputs = tf.squeeze(phrase, axis=0)\n        outputs = tf.one_hot(outputs, N_UNIQUE_CHARACTERS)\n            \n        # Return a dictionary with the output tensor\n        return {'outputs': outputs }\n\n# Define TF Lite Model\ntflite_keras_model = TFLiteModel(model)\n\n# Sanity Check\n# demo_sequence_id = 1816796431\ndemo_sequence_id = example_parquet_df.index.unique()[0]\ndemo_raw_data = example_parquet_df.loc[demo_sequence_id, COLUMNS0].values\ndemo_phrase_true = train_sequence_id.loc[demo_sequence_id, 'phrase']\nprint(f'demo_raw_data shape: {demo_raw_data.shape}, dtype: {demo_raw_data.dtype}')\ndemo_output = tflite_keras_model(demo_raw_data)['outputs'].numpy()\nprint(f'demo_output shape: {demo_output.shape}, dtype: {demo_output.dtype}')\nprint(f'demo_outputs phrase decoded: {outputs2phrase(demo_output)}')\nprint(f'phrase true: {demo_phrase_true}')","metadata":{"papermill":{"duration":6.507954,"end_time":"2023-07-02T15:02:00.234987","exception":false,"start_time":"2023-07-02T15:01:53.727033","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:17:00.877832Z","iopub.execute_input":"2023-07-31T07:17:00.878368Z","iopub.status.idle":"2023-07-31T07:17:07.34859Z","shell.execute_reply.started":"2023-07-31T07:17:00.878333Z","shell.execute_reply":"2023-07-31T07:17:07.347423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create Model Converter\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)\n# Convert Model\ntflite_model = keras_model_converter.convert()\n# Write Model\nwith open('/kaggle/working/model.tflite', 'wb') as f:\n    f.write(tflite_model)","metadata":{"papermill":{"duration":63.286967,"end_time":"2023-07-02T15:03:03.582231","exception":false,"start_time":"2023-07-02T15:02:00.295264","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:17:07.350176Z","iopub.execute_input":"2023-07-31T07:17:07.350836Z","iopub.status.idle":"2023-07-31T07:18:21.650372Z","shell.execute_reply.started":"2023-07-31T07:17:07.350794Z","shell.execute_reply":"2023-07-31T07:18:21.649295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add selected_columns json to only select specific columns from input frames\nwith open('inference_args.json', 'w') as f:\n     json.dump({ 'selected_columns': COLUMNS0.tolist() }, f)","metadata":{"papermill":{"duration":0.070965,"end_time":"2023-07-02T15:03:03.713519","exception":false,"start_time":"2023-07-02T15:03:03.642554","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:18:21.651803Z","iopub.execute_input":"2023-07-31T07:18:21.652398Z","iopub.status.idle":"2023-07-31T07:18:21.660476Z","shell.execute_reply.started":"2023-07-31T07:18:21.652361Z","shell.execute_reply":"2023-07-31T07:18:21.659565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Zip Model\n!zip submission.zip /kaggle/working/model.tflite /kaggle/working/inference_args.json","metadata":{"papermill":{"duration":1.931415,"end_time":"2023-07-02T15:03:05.70449","exception":false,"start_time":"2023-07-02T15:03:03.773075","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-07-31T07:18:21.661989Z","iopub.execute_input":"2023-07-31T07:18:21.66235Z","iopub.status.idle":"2023-07-31T07:18:24.077417Z","shell.execute_reply.started":"2023-07-31T07:18:21.662307Z","shell.execute_reply":"2023-07-31T07:18:24.076261Z"},"trusted":true},"execution_count":null,"outputs":[]}]}