{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🔎 Dataset Description\n\nThe objective of this competition is to detect and translate American Sign Language (ASL) to text.\n\nThis competition requires submissions in the form of TensorFlow Lite models. You can train your model using the framework of your choice as long as you convert the model checkpoint to the tflite format before submission. Please refer to the evaluation page for more details.\n\n# 📄 Files\n\n## [train/supplemental_metadata].csv\n\n* `path` - The path to the reference file.\n* `file_id` - A unique identifier for the data file.\n* `participant_id` - A unique identifier for the data contributor.\n* `sequence_id` - A unique identifier for the reference sequence. Each data file can contain many sequences.\n* `phrase` - The labels for the reference sequence. The training and test datasets contain **randomly generated addresses, phone numbers, and URLs derived from real addresses/phone numbers/URLs**. Any matches with real addresses, phone numbers, or URLs are purely coincidental. **The supplemental dataset consists of finger-spelled sentences**. Please note that some of the URLs include adult content. The goal of this competition is to support the deaf and hard-of-hearing community to engage with technology on equal terms with other adults.\n\n## character_to_prediction_index.json\n\nA dictionary mapping different symbols to an integer number.\n\n## [train/supplemental]_landmarks/\n\nThe reference data. The landmarks were extracted from raw videos using the MediaPipe model. **Not all frames necessarily had visible hands or hands that could be detected by the model**. The reference files contain the same data as in the ASL Signs competition (except for the row ID column) but with a wide structure. **This allows you to leverage the Parquet format to completely skip loading landmarks you're not using**.\n\n* `sequence_id` - A unique identifier for the reference sequence. The reference files contain approximately 1,000 sequences. The sequence ID is used as the index of the dataframe.\n* `frame` - The frame number within a reference sequence.\n* `[x/y/z]_[type]_[landmark_index]` - There are now 1,629 columns of spatial coordinates for the x, y, and z coordinates of each of the 543 landmarks. The landmark type can be one of `['face', 'left_hand', 'pose', 'right_hand']`. Details about the landmark locations for hands can be found here. The spatial coordinates have already been normalized by MediaPipe. Please note that the MediaPipe model is not fully trained to predict depth, so you may want to disregard the z-values. The landmarks have been converted to float32.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport matplotlib as mpl\nimport seaborn as sn\nimport tensorflow as tf\n\nfrom tqdm.notebook import tqdm\nfrom sklearn.model_selection import train_test_split, GroupShuffleSplit\nimport Levenshtein as lev\n\nimport glob\nimport sys\nimport os\nimport math\nimport gc\nimport sys\nimport sklearn\nimport time\nimport json\n\n# TQDM Progress Bar With Pandas Apply Function\ntqdm.pandas()\n\n# Supress Warnings\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nprint(f'Tensorflow Version {tf.__version__}')\nprint(f'Python Version: {sys.version}')\nprint(f'Levenshtein Version : {lev.__version__}')","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-07-03T08:11:50.531091Z","iopub.execute_input":"2023-07-03T08:11:50.531537Z","iopub.status.idle":"2023-07-03T08:11:50.542902Z","shell.execute_reply.started":"2023-07-03T08:11:50.531504Z","shell.execute_reply":"2023-07-03T08:11:50.54166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Character_to_prediction_index\n\nConvert the dictionary into a pandas dataframe","metadata":{}},{"cell_type":"markdown","source":"# Hyperparameters\n\nDefine the hyperparameters that will be used later\n\n- `IS_INTERACTIVE`: This variable indicates whether the notebook is running in interactive mode, which means it is being run in an environment where you can make changes and experiment interactively. In this mode, you can execute code cells individually, modify the code, and see the results immediately. This is useful when you are developing and debugging your code as it allows you to make quick changes and observe the effects of those changes in real-time.\n\n- `SEED`: This is a global random seed used for reproducibility. By setting a seed, it ensures that the results are consistent across different runs.\n\n- `N_TARGET_FRAMES`: This variable indicates the number of frames the recordings will be resized to. This is done to **standardize the length of input sequences**.\n\n- `DEBUG`: This variable indicates whether it is running in debug mode. If set to `True`, it will run in debug mode, which may involve running a subset of the data for faster execution and easier debugging.\n\n- `N_UNIQUE_CHARACTERS0`: This variable represents the number of **unique characters to be predicted in the model**. It does not include the padding token, start of sentence token, and end of sentence token.\n\n- `N_UNIQUE_CHARACTERS`: This variable is similar to `N_UNIQUE_CHARACTERS0`, but it includes the padding token, start of sentence token, and end of sentence token. Therefore, it represents the total number of classes that the model needs to predict.\n\n- `PAD_TOKEN`, `START_TOKEN`, and `END_TOKEN`: These variables represent the indices of the padding token, start of sentence token, and end of sentence token in the vocabulary. They are used during data processing and generation of target text sequences.\n","metadata":{}},{"cell_type":"code","source":"# If Notebook Is Run By Committing or In Interactive Mode For Development\nIS_INTERACTIVE = os.environ['KAGGLE_KERNEL_RUN_TYPE'] == 'Interactive'\n# Global Random Seed\nSEED = 42\n# Number of Frames to resize recording to\nN_TARGET_FRAMES = 128\n# Global debug flag, takes subset of train\nDEBUG = False\n# Length of Phrase + EOS Token\n# MAX_PHRASE_LENGTH = 31 + 1","metadata":{"collapsed":false,"ExecuteTime":{"start_time":"2023-06-30T11:02:29.959637Z","end_time":"2023-06-30T11:02:29.972515Z"},"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-07-03T08:11:50.545046Z","iopub.execute_input":"2023-07-03T08:11:50.546381Z","iopub.status.idle":"2023-07-03T08:11:50.564907Z","shell.execute_reply.started":"2023-07-03T08:11:50.546333Z","shell.execute_reply":"2023-07-03T08:11:50.56366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read Character to Ordinal Encoding Mapping\nwith open('/kaggle/input/asl-fingerspelling/character_to_prediction_index.json') as json_file:\n    CHAR2ORD = json.load(json_file)\n    \n# Number of Unique Characters To Predict + Pad Token + SOS Token + EOS Token\nN_UNIQUE_CHARACTERS0 = len(CHAR2ORD)\nN_UNIQUE_CHARACTERS = len(CHAR2ORD) + 1 + 1 + 1\nPAD_TOKEN = len(CHAR2ORD) # Padding # idx 59\nSTART_TOKEN = len(CHAR2ORD) + 1 # Start Of Sentence # idx 60\nEND_TOKEN = len(CHAR2ORD) + 2 # End Of Sentence # idx 61\n\n# assign tokens to the padding, start of the sentence and end of the sentence\nCHAR2ORD['P'] = PAD_TOKEN   # padding\nCHAR2ORD['S'] = START_TOKEN # start\nCHAR2ORD['E'] = END_TOKEN   # end\n\n# convert dictionary to pandas dataframe\nCHAR2ORD_DF = pd.DataFrame(CHAR2ORD.values(),index=CHAR2ORD.keys(),columns=['Ordinal Encoding'])\nCHAR2ORD_DF.head()","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-07-03T08:11:50.619931Z","iopub.execute_input":"2023-07-03T08:11:50.620944Z","iopub.status.idle":"2023-07-03T08:11:50.643011Z","shell.execute_reply.started":"2023-07-03T08:11:50.620888Z","shell.execute_reply":"2023-07-03T08:11:50.640779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Parquet Files: Working with Landmarks\n\nWe will select features from the *parquet* data for further modeling. To do this, we need to understand some notions about the provided data. The *pose landmarker* model tracks 33 body landmark locations, representing the approximate location of the following body parts:\n\n![Landmark Indexes](http://developers.google.com/static/mediapipe/images/solutions/pose_landmarks_index.png)","metadata":{}},{"cell_type":"code","source":"import pyarrow.parquet as pq\n\n# Read parquet file\ntable = pq.read_table('/kaggle/input/asl-fingerspelling/supplemental_landmarks/1176508147.parquet')\n\n# Convert to pd dataframe\ndf = table.to_pandas()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T08:11:50.645256Z","iopub.execute_input":"2023-07-03T08:11:50.646651Z","iopub.status.idle":"2023-07-03T08:11:56.164093Z","shell.execute_reply.started":"2023-07-03T08:11:50.646599Z","shell.execute_reply":"2023-07-03T08:11:56.163164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# z component of the variable pose ==> range(33)\ndf.iloc[:, -54:-21].head()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T08:11:56.166768Z","iopub.execute_input":"2023-07-03T08:11:56.167739Z","iopub.status.idle":"2023-07-03T08:11:56.218492Z","shell.execute_reply.started":"2023-07-03T08:11:56.1677Z","shell.execute_reply":"2023-07-03T08:11:56.216746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, the *pose* indexes are in the range `[0,32]`. In our case, since we are only interested in describing sign language gestures, we will only use the indexes corresponding to the hands. For the left hand, we will use the indexes `[13, 15, 17, 19, 21]`, and for the right hand, we will use the indexes `[14, 16, 18, 20, 22]`. Therefore, we will use a simplified model.\n\nWe will preprocess the *parquet* data for further modeling. First, let's examine the structure of the *parquet* format:","metadata":{}},{"cell_type":"code","source":"# z component of the variable right hand ==> range(21)\ndf.iloc[:, -21:].head()","metadata":{"execution":{"iopub.status.busy":"2023-07-03T08:11:56.222227Z","iopub.execute_input":"2023-07-03T08:11:56.223606Z","iopub.status.idle":"2023-07-03T08:11:56.271049Z","shell.execute_reply.started":"2023-07-03T08:11:56.223533Z","shell.execute_reply":"2023-07-03T08:11:56.269659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once the *parquet* format and the simplification we are going to use have been clarified, we proceed with the preprocessing.\n\n# Preprocessing","metadata":{}},{"cell_type":"code","source":"# read train.csv\ninpdir = \"/kaggle/input/asl-fingerspelling\"\ndf = pd.read_csv(f'{inpdir}/train.csv')\n\ndf.head()","metadata":{"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-07-03T08:11:56.272712Z","iopub.execute_input":"2023-07-03T08:11:56.27305Z","iopub.status.idle":"2023-07-03T08:11:56.450073Z","shell.execute_reply.started":"2023-07-03T08:11:56.273021Z","shell.execute_reply":"2023-07-03T08:11:56.448593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We define the following actions:\n\n- We define the list `LPOSE` that contains the indices of the positions corresponding to the left hand in our dataset.\n- We define the list `RPOSE` that contains the indices of the positions corresponding to the right hand in our dataset.\n- We combine the `LPOSE` and `RPOSE` lists into the `POSE` list, which contains all the indices of the hand positions.\n- We define the list `RHAND_LBLS` that contains the labels corresponding to the X, Y, and Z coordinates of the right hand in our dataset.\n- We define the list `LHAND_LBLS` that contains the labels corresponding to the X, Y, and Z coordinates of the left hand in our dataset.\n- We define the list `POSE_LBLS` that contains the labels corresponding to the X, Y, and Z coordinates of the hand positions in our dataset.\n- We define the list `X` that combines the X coordinate labels of the right hand, left hand, and hand positions.\n- We define the list `Y` that combines the Y coordinate labels of the right hand, left hand, and hand positions.\n- We define the list `Z` that combines the Z coordinate labels of the right hand, left hand, and hand positions.\n- We define the list `SEL_COLS` that combines the `X`, `Y`, and `Z` lists, representing all the selected columns for our analysis.\n- We define the variable `FRAME_LEN` with the value 128, indicating the length of each data frame.\n- We define the list `X_IDX` that contains the indices of the columns in `SEL_COLS` that contain \"x_\" in their name.\n- We define the list `Y_IDX` that contains the indices of the columns in `SEL_COLS` that contain \"y_\" in their name.\n- We define the list `Z_IDX` that contains the indices of the columns in `SEL_COLS` that contain \"z_\" in their name.\n- We define the list `RHAND_IDX` that contains the indices of the columns in `SEL_COLS` that contain \"right\" in their name, corresponding to the coordinates of the right hand.\n- We define the list `LHAND_IDX` that contains the indices of the columns in `SEL_COLS` that contain \"left\" in their name, corresponding to the coordinates of the left hand.\n- We define the list `RPOSE_IDX` that contains the indices of the columns in `SEL_COLS` that contain \"pose\" in their name and their last digit is in `RPOSE`, corresponding to the positions of the right hand.\n- We define the list `LPOSE_IDX` that contains the indices of the columns in `SEL_COLS` that contain \"pose\" in their name and their last digit is in `LPOSE`, corresponding to the positions of the left hand.","metadata":{}},{"cell_type":"code","source":"LPOSE = [13, 15, 17, 19, 21] # idx of the the left hand positions\nRPOSE = [14, 16, 18, 20, 22] # idx of the the right hand positions\nPOSE = LPOSE + RPOSE # all posible indexes (both hands)\n\n# all posibly labels for the right and left hand\n## hand labels:\nRHAND_LABELS = [f'x_right_hand_{i}' for i in range(21)] + [f'y_right_hand_{i}' for i in range(21)] + [f'z_right_hand_{i}' for i in range(21)]\nLHAND_LABELS = [ f'x_left_hand_{i}' for i in range(21)] + [ f'y_left_hand_{i}' for i in range(21)] + [ f'z_left_hand_{i}' for i in range(21)]\n## hand position labels:\nPOSE_LABELS = [f'x_pose_{i}' for i in POSE] + [f'y_pose_{i}' for i in POSE] + [f'z_pose_{i}' for i in POSE]\n\n# all posibly [x,y,z] labels\nX = [f'x_right_hand_{i}' for i in range(21)] + [f'x_left_hand_{i}' for i in range(21)] + [f'x_pose_{i}' for i in POSE]\nY = [f'y_right_hand_{i}' for i in range(21)] + [f'y_left_hand_{i}' for i in range(21)] + [f'y_pose_{i}' for i in POSE]\nZ = [f'z_right_hand_{i}' for i in range(21)] + [f'z_left_hand_{i}' for i in range(21)] + [f'z_pose_{i}' for i in POSE]\n\nSEL_COLS = X + Y + Z\nFRAME_LEN = 128\n\nX_IDX = [i for i, col in enumerate(SEL_COLS)  if \"x_\" in col]\nY_IDX = [i for i, col in enumerate(SEL_COLS)  if \"y_\" in col]\nZ_IDX = [i for i, col in enumerate(SEL_COLS)  if \"z_\" in col]\n\nRHAND_IDX = [i for i, col in enumerate(SEL_COLS)  if \"right\" in col]\nLHAND_IDX = [i for i, col in enumerate(SEL_COLS)  if  \"left\" in col]\nRPOSE_IDX = [i for i, col in enumerate(SEL_COLS)  if  \"pose\" in col and int(col[-2:]) in RPOSE]\nLPOSE_IDX = [i for i, col in enumerate(SEL_COLS)  if  \"pose\" in col and int(col[-2:]) in LPOSE]","metadata":{"execution":{"iopub.status.busy":"2023-07-03T08:11:56.451383Z","iopub.execute_input":"2023-07-03T08:11:56.451726Z","iopub.status.idle":"2023-07-03T08:11:56.469172Z","shell.execute_reply.started":"2023-07-03T08:11:56.451698Z","shell.execute_reply":"2023-07-03T08:11:56.467746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we aim to resize the data sequence in a standardized manner. For this purpose, we utilize the functions `resize_pad(x)` and `pre_process(x)`.\n\n1) The `resize_pad(x)` function is responsible for resizing and padding a data sequence `x` to ensure it has a specific length defined by `FRAME_LEN`.\n\n- If the length of the sequence `x` is shorter than `FRAME_LEN`, padding is added to the sequence using the `tf.pad()` function (zeros are appended to the sequence until it reaches the desired length).\n\n- On the other hand, if the length of the sequence `x` is longer than `FRAME_LEN`, resizing is performed on the sequence using the `tf.image.resize()` function. This ensures the sequence fits the desired length without losing important information.\n\n2) The `pre_process(x)` function is responsible for preprocessing the input data `x` before using it in a model.\n\n- The `tf.gather` function is used to perform an element extraction operation from a tensor based on the provided indices. Here, the columns corresponding to the coordinates of the right hand, left hand, right hand positions, and left hand positions are selected using the variables `RHAND_IDX`, `LHAND_IDX`, `RPOSE_IDX`, and `LPOSE_IDX`.\n\n- Next, NaN values in the right hand and left hand data are checked using the `tf.reduce_any(tf.math.is_nan())` function. This returns a boolean index (True/False) indicating whether there are any NaN values in each row of the right hand and left hand data. The `tf.reduce_any` function is used to reduce a tensor to a single boolean value indicating if at least one condition is met. In this case, it is applied to the expression `tf.math.is_nan(rhand)`, which checks if there are any NaN (Not a Number) values in the rhand tensor. By specifying axis=1, the reduction is performed along the column axis, meaning a boolean value is obtained for each row of the tensor. The result will be True for a row if at least one value in that row is NaN, and False otherwise.\n\n- Then, the count of NaN values is obtained using the `tf.math.count_nonzero()` function, giving the number of NaN values in the right hand and left hand data.\n\n- The NaN count in the right hand and left hand data is compared to determine which hand is dominant. If there are more NaN values in the right hand, it is assumed that the left hand is dominant, and the corresponding data is assigned to the variables `hand` and `pose`. Otherwise, the right hand data is assigned.\n\n- Create the hand tensor with new dimensions:\n  * The `tf.concat` function is used to combine the `hand_x`, `hand_y`, and `hand_z` coordinates along the last axis (the one we created), thus creating a new hand tensor with an additional dimension.\n  * The `tf.newaxis` function is used to add a new dimension to each coordinate. In the code, the `hand_x`, `hand_y`, and `hand_z` coordinates represent the three spatial dimensions of the hand, and we want to combine them into a three-dimensional tensor `hand` with shape `(batch_size, num_points, 3)`. However, for `tf.concat` to work, the tensors must have the same shape in all dimensions except the dimension along which we want to concatenate. By adding a new axis using `[..., tf.newaxis]`, the dimension of the individual coordinates is expanded from `(batch_size, num_points) => (batch_size, num_points, 1)`. This allows the coordinates to be concatenated properly along the last axis using `tf.concat`, thus creating the three-dimensional `hand` tensor.\n\n- Normalize the hand values: The mean and standard deviation of the hand tensor are calculated along axis 1 using `tf.math.reduce_mean` and `tf.math.reduce_std`, respectively. Then, the mean is subtracted and divided by the standard deviation to normalize the hand values.\n\n- Select the hand position coordinates: `pose_x`, `pose_y`, and `pose_z` are tensors containing the X, Y, and Z coordinates of the hand position, respectively. These coordinates are obtained by splitting the `pose` tensor into three equal parts based on the length of `LPOSE_IDX` (position indices).\n\n- Create the pose tensor with new dimensions: The `tf.concat` function is used similarly to hand, combining the `pose_x`, `pose_y`, and `pose_z` coordinates into a new pose tensor with an additional dimension.\n\n- Combine hand and pose into a single tensor `x`: The `tf.concat` function is used to concatenate hand and pose along the second axis, creating a new tensor `x`.\n\n- Perform resizing and size adjustment on `x`: A `resize_pad` function is applied to the `x` tensor to adjust its size according to some specific criterion that is not shown in the provided code. The `resize_pad` function likely changes the size of the tensor or adds padding values based on some dimension or format requirement.\n\n- Replace NaN values with zeros: The `tf.where` function is used to replace any NaN values in the `x` tensor with zero, creating a modified tensor `x`.\n\n- Reshape `x`: Finally, the `x` tensor is reshaped using the `tf.reshape` function to have a specific shape of (`FRAME_LEN, len(LHAND_IDX) + len(LPOSE_IDX)`).","metadata":{}},{"cell_type":"code","source":"def resize_pad(x):\n    if tf.shape(x)[0] < FRAME_LEN:\n        # if its < then we use padding to resize it\n        x = tf.pad(x, ([[0, FRAME_LEN-tf.shape(x)[0]], [0, 0], [0, 0]]))\n    else:\n        # if its > we resize it\n        x = tf.image.resize(x, (FRAME_LEN, tf.shape(x)[1]))\n    return x\n\ndef pre_process(x):\n    # right and left hand indexes\n    rhand = tf.gather(x, RHAND_IDX, axis=1) # extract right hand indexes on the tensor 'x'\n    lhand = tf.gather(x, LHAND_IDX, axis=1)\n    # right and left hand position indexes\n    rpose = tf.gather(x, RPOSE_IDX, axis=1) # extract right hand position indexes on the tensor 'x'\n    lpose = tf.gather(x, LPOSE_IDX, axis=1)\n    \n    # search if this row got a NaN\n    rnan_idx = tf.reduce_any(tf.math.is_nan(rhand), axis=1) \n    lnan_idx = tf.reduce_any(tf.math.is_nan(lhand), axis=1)\n    \n    # count number of NaNs \n    rnans = tf.math.count_nonzero(rnan_idx)\n    lnans = tf.math.count_nonzero(lnan_idx)\n    \n    # For dominant hand\n    if rnans > lnans:\n        # if the number of NaNs if greater in one right hand then the dominant is the left one\n        # we assign the hand as the dominant one\n        hand = lhand\n        # positions then are the left positions\n        pose = lpose\n        \n        hand_x = hand[:, 0*(len(LHAND_IDX)//3) : 1*(len(LHAND_IDX)//3)]\n        hand_y = hand[:, 1*(len(LHAND_IDX)//3) : 2*(len(LHAND_IDX)//3)]\n        hand_z = hand[:, 2*(len(LHAND_IDX)//3) : 3*(len(LHAND_IDX)//3)]\n        hand = tf.concat([1-hand_x, hand_y, hand_z], axis=1) # this is to make it right handed\n        \n        pose_x = pose[:, 0*(len(LPOSE_IDX)//3) : 1*(len(LPOSE_IDX)//3)]\n        pose_y = pose[:, 1*(len(LPOSE_IDX)//3) : 2*(len(LPOSE_IDX)//3)]\n        pose_z = pose[:, 2*(len(LPOSE_IDX)//3) : 3*(len(LPOSE_IDX)//3)]\n        pose = tf.concat([1-pose_x, pose_y, pose_z], axis=1) # this is to make it right handed\n        \n        \n    else:\n        # if the number of NaNs if lower in one right hand then the dominant is the right one\n        hand = rhand\n        pose = rpose\n    \n    # select the coordinates of the hand \n    hand_x = hand[:, 0*(len(LHAND_IDX)//3) : 1*(len(LHAND_IDX)//3)]\n    hand_y = hand[:, 1*(len(LHAND_IDX)//3) : 2*(len(LHAND_IDX)//3)]\n    hand_z = hand[:, 2*(len(LHAND_IDX)//3) : 3*(len(LHAND_IDX)//3)]\n    #   create the hand in the coordinates x,y,z \n    hand = tf.concat([hand_x[..., tf.newaxis], hand_y[..., tf.newaxis], hand_z[..., tf.newaxis]], axis=-1)\n    #hand = tf.concat([hand_x, hand_y, hand_z], axis=1) # ESTO DA ERROR, ERA UNA PRUEBA\n    \n    mean = tf.math.reduce_mean(hand, axis=1)[:, tf.newaxis, :]\n    std = tf.math.reduce_std(hand, axis=1)[:, tf.newaxis, :]\n    hand = (hand - mean) / std\n    # select the coordinates of the position of the hand \n    pose_x = pose[:, 0*(len(LPOSE_IDX)//3) : 1*(len(LPOSE_IDX)//3)]\n    pose_y = pose[:, 1*(len(LPOSE_IDX)//3) : 2*(len(LPOSE_IDX)//3)]\n    pose_z = pose[:, 2*(len(LPOSE_IDX)//3) : 3*(len(LPOSE_IDX)//3)]\n    #  create position of the hand in the coordinates x,y,z \n    pose = tf.concat([pose_x[..., tf.newaxis], pose_y[..., tf.newaxis], pose_z[..., tf.newaxis]], axis=-1)\n    # pose = tf.concat([pose_x, pose_y, pose_z], axis=1) # ESTO DA ERROR, ERA UNA PRUEBA\n    # join the hand with the position\n    x = tf.concat([hand, pose], axis=1)\n    x = resize_pad(x)\n    \n    x = tf.where(tf.math.is_nan(x), tf.zeros_like(x), x)\n    x = tf.reshape(x, (FRAME_LEN, len(LHAND_IDX) + len(LPOSE_IDX)))\n    return x","metadata":{"execution":{"iopub.status.busy":"2023-07-03T08:15:24.006088Z","iopub.execute_input":"2023-07-03T08:15:24.006539Z","iopub.status.idle":"2023-07-03T08:15:24.029283Z","shell.execute_reply.started":"2023-07-03T08:15:24.006508Z","shell.execute_reply":"2023-07-03T08:15:24.02819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, once the preprocessing function is defined, we proceed to prepare the dataset for further training:\n\n1) A lookup table `(StaticHashTable)` is defined using the `tf.lookup.StaticHashTable` class. This table is used to map characters (keys) to numeric values (values). It is initialized with an initializer that specifies the keys and values using a `KeyValueTensorInitializer`. Additionally, a `default_value` is specified, which is used when a key does not have an associated value in the table. We can think of `tf.lookup.StaticHashTable` as a scalable dictionary in TensorFlow. While a conventional dictionary in Python may work well for a small number of keys and values, when dealing with larger datasets, the lookup table provides a more efficient way to perform mappings and lookups.\n\n2) A `preprocess_fn` function is defined, which takes `landmarks` and `phrase` as inputs. The function performs the following operations:\n\n- Adds a start token `(start_token)` and an end token `(end_token)` to the phrase (already defined previously).\n- Splits the phrase into bytes using `tf.strings.bytes_split`.\n- Performs a lookup in the table using `table.lookup` to convert the bytes of the phrase into numeric values.\n- Pads the phrase with zeros (pad_token_idx) to have a length of 64 elements (all of them have the same length).\n- Calls the pre_process function to preprocess the landmarks.\n\n3) A `decode_fn` function is defined, which takes a byte record as input and performs the decoding process.\n\n- A variable named `schema` is defined, which specifies the structure of the data stored in a record. In this case, a dictionary is used where each key `(COL)` represents a data column, and its associated value is a `VarLenFeature` object that specifies that the data in that column is of type *float32* and has a variable length.\n- An additional entry is added to the `schema` called `\"phrase\"`, which is expected to be a text string and has a fixed length of size 0 (i.e., a single element).\n- The `parse_single_example` function is used to decode a record stored in the *TFRecord* format. The function takes two arguments: `record_bytes`, which represents the record in bytes form, and `schema`, which defines the structure of the data in the record. The `parse_single_example` function uses `schema` to interpret and extract the data from the `record_bytes` record. After calling `parse_single_example`, a dictionary named `features` is obtained, which contains the decoded data from the record. Each key in `features` corresponds to a feature defined in the schema, and the associated values are the decoded data for those features.\n- The data for the remaining columns is extracted using a list comprehension. It iterates over the keys `(COL)` in `SEL_COLS` and uses `tf.sparse.to_dense` to convert sparse data into a dense representation. This data is stored in a list.\n- Finally, the data stored in the `landmarks` list is transposed so that the rows represent data instances and the columns represent features.\n- The function returns the `landmarks` data (landmarks) and the phrase, which will be used later in the preprocessing and model training process.","metadata":{}},{"cell_type":"code","source":"table = tf.lookup.StaticHashTable(\n    initializer=tf.lookup.KeyValueTensorInitializer(\n        keys=list(CHAR2ORD.keys()),\n        values=list(CHAR2ORD.values()),\n    ),\n    default_value=tf.constant(-1),\n    name=\"class_weight\"\n)\n\ndef preprocess_fn(landmarks, phrase):\n    # normalize prhase length\n    phrase = 'S' + phrase + 'E' # add start token and end token\n    phrase = tf.strings.bytes_split(phrase) # split into bytes\n    phrase = table.lookup(phrase) # convert into numeric using StaticHashTable\n    phrase = tf.pad(phrase, paddings=[[0, 64 - tf.shape(phrase)[0]]], mode = 'CONSTANT',\n                    constant_values = PAD_TOKEN) # normalize\n    return pre_process(landmarks), phrase\n\ndef decode_fn(record_bytes):\n    # decodify\n    schema = {COL: tf.io.VarLenFeature(dtype=tf.float32) for COL in SEL_COLS} # set data to be float32 \n    schema[\"phrase\"] = tf.io.FixedLenFeature([], dtype=tf.string) # fixed length \n    features = tf.io.parse_single_example(record_bytes, schema) # transform record_bytes into tfrecord using fixed schema\n    phrase = features[\"phrase\"] \n    landmarks = ([tf.sparse.to_dense(features[COL]) for COL in SEL_COLS]) # select the rest of the features\n    landmarks = tf.transpose(landmarks)\n    \n    return landmarks, phrase","metadata":{"execution":{"iopub.status.busy":"2023-07-03T08:14:20.895751Z","iopub.execute_input":"2023-07-03T08:14:20.896133Z","iopub.status.idle":"2023-07-03T08:14:20.912073Z","shell.execute_reply.started":"2023-07-03T08:14:20.896103Z","shell.execute_reply":"2023-07-03T08:14:20.910527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we finally create the train and validation datasets:\n\n- A list of full paths to the corresponding TFRecord files for each unique value of `file_id` is created.\n- The list of paths is converted into a set using the `unique()` method, which removes any duplicates that may exist.\n- The `batch_size` variable is defined with a value of 32, representing the size of each batch of data to be used during training.\n- The training `(train_dataset)` and validation `(val_dataset)` datasets are created using the `TFRecordDataset` class from TensorFlow. The paths of the TFRecord files corresponding to the values starting from the `val_len` index in the `tffiles` list are passed as an argument. Then, the decoding function `(decode_fn)` and the preprocessing function `(preprocess_fn)` are applied. After that, the data is randomly shuffled using `shuffle` with a buffer size of 30000 and with the `reshuffle_each_iteration` parameter set to True. Next, the data is batched using `batch` with a size equal to `batch_size`. Finally, `prefetch` is used to efficiently load the data into the model, with a buffer size automatically configured by TensorFlow using `tf.data.AUTOTUNE`.","metadata":{}},{"cell_type":"code","source":"inpdir = \"/kaggle/input/aslfr-parquets-to-tfrecords-cleaned\"\ntffiles = df['file_id'].map(lambda x: f'{inpdir}/tfds/{x}.tfrecord').unique() # number of files tfrecord\n\nBATCH_SIZE = 32\nval_len = int(0.05 * len(tffiles)) # length of the validation dataset\n\ntrain_dataset = tf.data.TFRecordDataset(tffiles[val_len:]).map(decode_fn).map(preprocess_fn).shuffle(30000, reshuffle_each_iteration=True).batch(BATCH_SIZE).prefetch(buffer_size=tf.data.AUTOTUNE)\nval_dataset = tf.data.TFRecordDataset(tffiles[:val_len]).map(decode_fn).map(preprocess_fn).batch(BATCH_SIZE).prefetch(buffer_size=tf.data.AUTOTUNE)","metadata":{"execution":{"iopub.status.busy":"2023-07-03T08:15:27.125265Z","iopub.execute_input":"2023-07-03T08:15:27.125798Z","iopub.status.idle":"2023-07-03T08:15:29.087356Z","shell.execute_reply.started":"2023-07-03T08:15:27.125751Z","shell.execute_reply":"2023-07-03T08:15:29.086222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dataset","metadata":{"execution":{"iopub.status.busy":"2023-07-03T08:21:59.32966Z","iopub.execute_input":"2023-07-03T08:21:59.330187Z","iopub.status.idle":"2023-07-03T08:21:59.338974Z","shell.execute_reply.started":"2023-07-03T08:21:59.330142Z","shell.execute_reply":"2023-07-03T08:21:59.338098Z"},"trusted":true},"execution_count":null,"outputs":[]}]}