{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Hello Fellow Kagglers,\n\nThis notebook gives an analysis of the competition dataset and demonstrates the preprocessing of the dataset, giving the X/y data to use for training.\n\nThe processing is as follows:\n\n1) Select dominant hand based on most number of non empty hand frames\n\n2) Filter out all frames with missing dominant hand coordinates\n\n3) Resize video to 256 frames\n\nThis is a work in progress and updates will follow.\n\nThe processed data could help with making a baseline.\n\nSoon, a training notebook will follow with the corresponding inference.\n\nExited to continue working on sign language!\n\n> This is an explained version of the notebook of [MARK WIJKHUIZEN](https://www.kaggle.com/code/markwijkhuizen/aslfr-eda-preprocessing-dataset/notebook)\n\n# 🔎 Dataset Description\n\nThe objective of this competition is to detect and translate American Sign Language (ASL) to text.\n\nThis competition requires submissions in the form of TensorFlow Lite models. You can train your model using the framework of your choice as long as you convert the model checkpoint to the tflite format before submission. Please refer to the evaluation page for more details.\n\n# 📄 Files\n\n## [train/supplemental_metadata].csv\n\n* `path` - The path to the reference file.\n* `file_id` - A unique identifier for the data file.\n* `participant_id` - A unique identifier for the data contributor.\n* `sequence_id` - A unique identifier for the reference sequence. Each data file can contain many sequences.\n* `phrase` - The labels for the reference sequence. The training and test datasets contain **randomly generated addresses, phone numbers, and URLs derived from real addresses/phone numbers/URLs**. Any matches with real addresses, phone numbers, or URLs are purely coincidental. **The supplemental dataset consists of finger-spelled sentences**. Please note that some of the URLs include adult content. The goal of this competition is to support the deaf and hard-of-hearing community to engage with technology on equal terms with other adults.\n\n## character_to_prediction_index.json\n\nA dictionary mapping different symbols to an integer number.\n\n## [train/supplemental]_landmarks/\n\nThe reference data. The landmarks were extracted from raw videos using the MediaPipe model. **Not all frames necessarily had visible hands or hands that could be detected by the model**. The reference files contain the same data as in the ASL Signs competition (except for the row ID column) but with a wide structure. **This allows you to leverage the Parquet format to completely skip loading landmarks you're not using**.\n\n* `sequence_id` - A unique identifier for the reference sequence. The reference files contain approximately 1,000 sequences. The sequence ID is used as the index of the dataframe.\n* `frame` - The frame number within a reference sequence.\n* `[x/y/z]_[type]_[landmark_index]` - There are now 1,629 columns of spatial coordinates for the x, y, and z coordinates of each of the 543 landmarks. The landmark type can be one of `['face', 'left_hand', 'pose', 'right_hand']`. Details about the landmark locations for hands can be found here. The spatial coordinates have already been normalized by MediaPipe. Please note that the MediaPipe model is not fully trained to predict depth, so you may want to disregard the z-values. The landmarks have been converted to float32.\n","metadata":{"execution":{"iopub.status.busy":"2023-07-03T12:37:40.032765Z","iopub.execute_input":"2023-07-03T12:37:40.03309Z","iopub.status.idle":"2023-07-03T12:37:40.045785Z","shell.execute_reply.started":"2023-07-03T12:37:40.033065Z","shell.execute_reply":"2023-07-03T12:37:40.044177Z"}}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport matplotlib as mpl\nimport seaborn as sn\nimport tensorflow as tf\n\nfrom tqdm.notebook import tqdm\nfrom sklearn.model_selection import train_test_split, GroupShuffleSplit\nfrom pathlib import Path\n\nimport glob\nimport sys\nimport os\nimport math\nimport gc\nimport sys\nimport sklearn\nimport time\nimport json\nimport re\n\n# TQDM Progress Bar With Pandas Apply Function\ntqdm.pandas()","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:46:46.051247Z","iopub.execute_input":"2023-07-11T08:46:46.051653Z","iopub.status.idle":"2023-07-11T08:46:57.749322Z","shell.execute_reply.started":"2023-07-11T08:46:46.051622Z","shell.execute_reply":"2023-07-11T08:46:57.748289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Character_to_prediction_index\n\nConvertimos the dictionary into a pandas dataframe ","metadata":{"execution":{"iopub.status.busy":"2023-07-03T12:40:21.75744Z","iopub.execute_input":"2023-07-03T12:40:21.757777Z","iopub.status.idle":"2023-07-03T12:40:21.763205Z","shell.execute_reply.started":"2023-07-03T12:40:21.757749Z","shell.execute_reply":"2023-07-03T12:40:21.762033Z"}}},{"cell_type":"code","source":"# Read Character to Ordinal Encoding Mapping\nwith open('/kaggle/input/asl-fingerspelling/character_to_prediction_index.json') as json_file:\n    CHAR2ORD = json.load(json_file)\n\n# convert dictionary to pandas dataframe\nCHAR2ORD_DF = pd.DataFrame(CHAR2ORD.values(),index=CHAR2ORD.keys(),columns=['Ordinal Encoding'])\nCHAR2ORD_DF.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:46:57.751572Z","iopub.execute_input":"2023-07-11T08:46:57.753307Z","iopub.status.idle":"2023-07-11T08:46:57.798199Z","shell.execute_reply.started":"2023-07-11T08:46:57.753259Z","shell.execute_reply":"2023-07-11T08:46:57.797194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Hyperparameters\n\nDefine the hyperparameters that will be used later\n\n- `IS_INTERACTIVE`: This variable indicates whether the notebook is running in interactive mode, which means it is being run in an environment where you can make changes and experiment interactively. In this mode, you can execute code cells individually, modify the code, and see the results immediately. This is useful when you are developing and debugging your code as it allows you to make quick changes and observe the effects of those changes in real-time.\n\n- `SEED`: This is a global random seed used for reproducibility. By setting a seed, it ensures that the results are consistent across different runs.\n\n- `N_TARGET_FRAMES`: This variable indicates the number of frames the recordings will be resized to. This is done to **standardize the length of input sequences**.\n\n- `DEBUG`: This variable indicates whether it is running in debug mode. If set to `True`, it will run in debug mode, which may involve running a subset of the data for faster execution and easier debugging.\n\n- `N_UNIQUE_CHARACTERS0`: This variable represents the number of **unique characters to be predicted in the model**. It does not include the padding token, start of sentence token, and end of sentence token.\n\n- `N_UNIQUE_CHARACTERS`: This variable is similar to `N_UNIQUE_CHARACTERS0`, but it includes the padding token, start of sentence token, and end of sentence token. Therefore, it represents the total number of classes that the model needs to predict.\n\n- `N_UNIQUE_CHARACTERSPAD_TOKEN`, `START_TOKEN`, and `END_TOKEN`: These variables represent the indices of the padding token, start of sentence token, and end of sentence token in the vocabulary. They are used during data processing and generation of target text sequences.","metadata":{}},{"cell_type":"code","source":"# Number of Unique Characters\nN_UNIQUE_CHARACTERS = len(CHAR2ORD)\nprint(f'N_UNIQUE_CHARACTERS: {N_UNIQUE_CHARACTERS}') # ends in number 58","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:46:57.799548Z","iopub.execute_input":"2023-07-11T08:46:57.799956Z","iopub.status.idle":"2023-07-11T08:46:57.806098Z","shell.execute_reply.started":"2023-07-11T08:46:57.799918Z","shell.execute_reply":"2023-07-11T08:46:57.804915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# If Notebook Is Run By Committing or In Interactive Mode For Development\nIS_INTERACTIVE = os.environ['KAGGLE_KERNEL_RUN_TYPE'] == 'Interactive'\n# Describe Statistics Percentiles\nPERCENTILES = [0.01, 0.10, 0.05, 0.25, 0.50, 0.75, 0.90, 0.95, 0.99, 0.999]\n# Global Random Seed\nSEED = 42\n# Number of Frames to resize recording to\nN_TARGET_FRAMES = 128\n# Global debug flag, takes subset of train\nDEBUG = False\n# Fast Processing\nFAST= False\n# Number of Unique Characters To Predict + Pad Token + Start of sentence and end of sentence Token \nN_UNIQUE_CHARACTERSPAD_TOKEN = len(CHAR2ORD) # Padding = 59\nSTART_TOKEN = len(CHAR2ORD) + 1 # Start Of Sentence = 60\nEND_TOKEN = len(CHAR2ORD) + 2 # End Of Sentence = 61","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:46:57.809108Z","iopub.execute_input":"2023-07-11T08:46:57.80949Z","iopub.status.idle":"2023-07-11T08:46:57.820585Z","shell.execute_reply.started":"2023-07-11T08:46:57.809461Z","shell.execute_reply":"2023-07-11T08:46:57.819352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Preparation","metadata":{}},{"cell_type":"code","source":"# Read Train DataFrame\nif DEBUG:\n    train = pd.read_csv('/kaggle/input/asl-fingerspelling/train.csv').head(5000)\nelse:\n    train = pd.read_csv('/kaggle/input/asl-fingerspelling/train.csv')\n\n# Get complete file path to file\ndef get_file_path(path):\n    return f'/kaggle/input/asl-fingerspelling/{path}'\n\ntrain['file_path'] = train['path'].apply(get_file_path)\n\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:46:57.824236Z","iopub.execute_input":"2023-07-11T08:46:57.824653Z","iopub.status.idle":"2023-07-11T08:46:58.061665Z","shell.execute_reply.started":"2023-07-11T08:46:57.824623Z","shell.execute_reply":"2023-07-11T08:46:58.060249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We define the `get_phrase_type` function, which takes a phrase as input:\n\n- First, it checks if the phrase consists solely of digits, plus signs (+), or hyphens (-) using the regex expression `re.match(r'^[\\d+-]+$', phrase)`. If the phrase matches this pattern, it is considered a \"Phone Number\".\n- If the first condition is not met, the function checks if the phrase contains any of the substrings `['www', '.', '/']` and does not contain whitespace using the expression `any([substr in phrase for substr in ['www', '.', '/']]) and ' ' not in phrase`. If this condition is met, it is considered a \"URL\".\n- If neither of the two conditions above is met, it is assumed that the phrase is an \"Address\".\n- Finally, the function returns the identifier.","metadata":{}},{"cell_type":"code","source":"def get_phrase_type(phrase):\n    # Phone Number\n    if re.match(r'^[\\d+-]+$', phrase):\n        return 'phone_number'\n    # url\n    elif any([substr in phrase for substr in ['www', '.', '/']]) and ' ' not in phrase:\n        return 'url'\n    # Address\n    else:\n        return 'address'\n    \ntrain['phrase_type'] = train['phrase'].apply(get_phrase_type)\n\ntrain.head(10)","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:46:58.063097Z","iopub.execute_input":"2023-07-11T08:46:58.063527Z","iopub.status.idle":"2023-07-11T08:46:58.27034Z","shell.execute_reply.started":"2023-07-11T08:46:58.063498Z","shell.execute_reply":"2023-07-11T08:46:58.269018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We extract the characters and the length of the phrase from the sentence.","metadata":{}},{"cell_type":"code","source":"# Split Phrase To Char Tuple\ntrain['phrase_char'] = train['phrase'].apply(tuple)\n# Character Length of Phrase\ntrain['phrase_char_len'] = train['phrase_char'].apply(len)\n\n# Maximum Input Length\nMAX_PHRASE_LENGTH = train['phrase_char_len'].max()\nprint(f'MAX_PHRASE_LENGTH: {MAX_PHRASE_LENGTH}')\n\n# Train DataFrame indexed by sequence_id to convenientlyy lookup recording data\ntrain_sequence_id = train.set_index('sequence_id')\n\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:46:58.271614Z","iopub.execute_input":"2023-07-11T08:46:58.271957Z","iopub.status.idle":"2023-07-11T08:46:58.420988Z","shell.execute_reply.started":"2023-07-11T08:46:58.271921Z","shell.execute_reply":"2023-07-11T08:46:58.41983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Character Count Occurance\nplt.figure(figsize=(15,8))\nplt.title('Character Length Occurance of Phrases')\ntrain['phrase_char_len'].value_counts().sort_index().plot(kind='bar')\nplt.xlim(-0.50, train['phrase_char_len'].max() - 1.50)\nplt.xlabel('Pharse Character Length')\nplt.ylabel('Sample Count')\nplt.grid(axis='y')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:46:58.422264Z","iopub.execute_input":"2023-07-11T08:46:58.422681Z","iopub.status.idle":"2023-07-11T08:46:59.0359Z","shell.execute_reply.started":"2023-07-11T08:46:58.422644Z","shell.execute_reply":"2023-07-11T08:46:59.034667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read First Parquet File\nexample_parquet_df = pd.read_parquet(train['file_path'][0])\n\n# Each parquet file contains 1000 unique recordings\nprint(f'Number of Unique Recording: {example_parquet_df.index.nunique()}')\n# Display DataFrame layout\nexample_parquet_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:46:59.037617Z","iopub.execute_input":"2023-07-11T08:46:59.038095Z","iopub.status.idle":"2023-07-11T08:47:18.390498Z","shell.execute_reply.started":"2023-07-11T08:46:59.038053Z","shell.execute_reply":"2023-07-11T08:47:18.389616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Define a variable `N` that indicates the number of Parquet file chunks to be analyzed. If the variable `IS_INTERACTIVE` is True, set `N` to 5; otherwise, set it to 25.\n- Define an empty list `N_UNIQUE_FRAMES` that will be used to store the number of unique frames in each recording.\n- Create a series `UNIQUE_FILE_PATHS` that contains the unique file paths from the `train` DataFrame in the \"file_path\" column.\n- Then, iterate over a random sample of `N` file paths obtained from `UNIQUE_FILE_PATHS` using `UNIQUE_FILE_PATHS.sample(N, random_state=SEED)`. This involves randomly selecting a subset of file paths to analyze.\n- For each selected file path, read the corresponding Parquet file using `pd.read_parquet(file_path)` and store it in the DataFrame `df`. Next, group the DataFrame `df` by the value in the \"sequence_id\" column using `groupby('sequence_id')`.\n- For each group (a group of sequences with the same ID), calculate the number of unique frames in the \"frame\" column using `group_df['frame'].nunique()`. The obtained number of unique frames is added to the `N_UNIQUE_FRAMES` list.","metadata":{}},{"cell_type":"code","source":"# Number of parquet chunks to analyse\nN = 5 if IS_INTERACTIVE else 25\n# Number of Unique Frames in Recording\nN_UNIQUE_FRAMES = []\n\nUNIQUE_FILE_PATHS = pd.Series(train['file_path'].unique()) # total file paths\n\nfor idx, file_path in enumerate(tqdm(UNIQUE_FILE_PATHS.sample(N, random_state=SEED))):\n    df = pd.read_parquet(file_path) # read the parquet \n    for group, group_df in df.groupby('sequence_id'):\n        N_UNIQUE_FRAMES.append(group_df['frame'].nunique())\n\n# Convert to Numpy Array\nN_UNIQUE_FRAMES = np.array(N_UNIQUE_FRAMES)","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:47:18.391575Z","iopub.execute_input":"2023-07-11T08:47:18.392692Z","iopub.status.idle":"2023-07-11T08:49:07.714953Z","shell.execute_reply.started":"2023-07-11T08:47:18.39266Z","shell.execute_reply":"2023-07-11T08:49:07.713536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,8))\nplt.title('Number of Unique Frames', size=24)\npd.Series(N_UNIQUE_FRAMES).plot(kind='hist', bins=128)\nplt.grid()\nxlim = math.ceil(plt.xlim()[1])\nplt.xlim(0, xlim)\nplt.xticks(np.arange(0, xlim+50, 50))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:07.719706Z","iopub.execute_input":"2023-07-11T08:49:07.720102Z","iopub.status.idle":"2023-07-11T08:49:08.35224Z","shell.execute_reply.started":"2023-07-11T08:49:07.720069Z","shell.execute_reply":"2023-07-11T08:49:08.350998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Define the function `get_idxs`, which takes as input a DataFrame `df`, a list of positive keyword `words_pos`, another list of negative keywords `words_neg` (optional), a boolean indicator `ret_names` (indicating whether column names should be returned), and a list of positive column indices `idxs_pos` (optional).\n\nThe function iterates over the DataFrame `df` and searches for columns whose names meet certain conditions. The goal is to find columns that contain all positive keywords and none of the negative keywords.\n\nFor each positive keyword in `words_pos`, the function iterates over the columns of the DataFrame. If a column is not a landmark column (such as the \"frame\" column), it is skipped and the next column is processed.\n\nThen, the function extracts the column index of the current column and checks if the column name contains the positive keyword and none of the negative keywords. If it meets these conditions and, optionally, if the column index is in the `idxs_pos` list, the column index is added to the `idxs` list and the column name is added to the `names` list.\n\nFinally, the function converts the `idxs` and `names` lists into NumPy arrays and returns them. Depending on the value of `ret_names`, the function may return only the column indices or both the column indices and the column names.","metadata":{}},{"cell_type":"code","source":"def get_idxs(df, words_pos, words_neg=['z'], ret_names=True, idxs_pos=None):\n    # words_pos is the first element we want to select : ['face','right_hand','left_hand','pose']\n    # words_neg are the components we want to exclude: ['x','y','z']\n    idxs = []\n    names = []\n    for w in words_pos:\n        for col_idx, col in enumerate(example_parquet_df.columns):\n            # Exclude Non Landmark Columns\n            if col in ['frame']:\n                continue\n                \n            col_idx = int(col.split('_')[-1])\n            # Check if column name contains all words\n            if (w in col) and (idxs_pos is None or col_idx in idxs_pos) and all([w not in col for w in words_neg]):\n                idxs.append(col_idx)\n                names.append(col)\n    # Convert to Numpy arrays\n    idxs = np.array(idxs)\n    names = np.array(names)\n    # Returns either both column indices and names\n    if ret_names:\n        return idxs, names\n    # Or only columns indices\n    else:\n        return idxs","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:08.353428Z","iopub.execute_input":"2023-07-11T08:49:08.353739Z","iopub.status.idle":"2023-07-11T08:49:08.363682Z","shell.execute_reply.started":"2023-07-11T08:49:08.353712Z","shell.execute_reply":"2023-07-11T08:49:08.36243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we extract the different indices from the Parquet files that are relevant to solve the problem. We will focus on the indices of the left and right hands and the corresponding indices of the lips, as they can provide us with the most information. The indices of the lips are obtained from Google ([lips landmarks](https://github.com/tensorflow/tfjs-models/blob/838611c02f51159afdd77469ce67f0e26b7bbb23/face-landmarks-detection/src/mediapipe-facemesh/keypoints.ts)).","metadata":{}},{"cell_type":"code","source":"# Lips Landmark Face Ids\nLIPS_LANDMARK_IDXS = np.array([\n        61, 185, 40, 39, 37, 0, 267, 269, 270, 409,\n        291, 146, 91, 181, 84, 17, 314, 405, 321, 375,\n        78, 191, 80, 81, 82, 13, 312, 311, 310, 415,\n        95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n    ])\n\n# Landmark Indices for Left/Right hand without z axis in raw data\nLEFT_HAND_IDXS0, LEFT_HAND_NAMES0 = get_idxs(example_parquet_df, ['left_hand'], ['z'])\nRIGHT_HAND_IDXS0, RIGHT_HAND_NAMES0 = get_idxs(example_parquet_df, ['right_hand'], ['z'])\nLIPS_IDXS0, LIPS_NAMES0 = get_idxs(example_parquet_df, ['face'], ['z'], idxs_pos=LIPS_LANDMARK_IDXS)\nCOLUMNS0 = np.concatenate((LEFT_HAND_NAMES0, RIGHT_HAND_NAMES0, LIPS_NAMES0))\nN_COLS0 = len(COLUMNS0)\n# Only X/Y axes are used\nN_DIMS0 = 2\n\nprint(f'N_COLS0: {N_COLS0}')","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:08.365647Z","iopub.execute_input":"2023-07-11T08:49:08.366103Z","iopub.status.idle":"2023-07-11T08:49:08.394454Z","shell.execute_reply.started":"2023-07-11T08:49:08.366069Z","shell.execute_reply":"2023-07-11T08:49:08.393129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f' Indexes of the left hand: {LEFT_HAND_IDXS0.tolist()}')\nprint(f'\\n Indexes of the lips: {LIPS_IDXS0.tolist()}')","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:08.395864Z","iopub.execute_input":"2023-07-11T08:49:08.39631Z","iopub.status.idle":"2023-07-11T08:49:08.411235Z","shell.execute_reply.started":"2023-07-11T08:49:08.396279Z","shell.execute_reply":"2023-07-11T08:49:08.410157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we select the columns from the Parquet that are of interest to us, namely, those that contain the indices of the left and right hands and the lips:\n\n- The `np.isin()` function is used to check which columns from the `COLUMNS0` array are present in the `LEFT_HAND_NAMES0` list. The `np.argwhere()` function is used to obtain the indices of the matching columns. The result is assigned to the `LEFT_HAND_IDXS` variable. This process is repeated for the `RIGHT_HAND_NAMES0` and `LIPS_NAMES0` lists, and the corresponding column indices are obtained in `RIGHT_HAND_IDXS` and `LIPS_IDXS`, respectively. These indices represent the locations of the columns in the `COLUMNS0` array.\n\n- The `N_COLS` variable is assigned the value of `N_COLS0`. This indicates the total number of columns present in `COLUMNS0`.\n\n- The `N_DIMS` variable is set to 2. This indicates that only the X and Y axes of the landmark coordinates will be used.","metadata":{}},{"cell_type":"code","source":"# Landmark Indices in subset of dataframe with only COLUMNS selected\nLEFT_HAND_IDXS = np.argwhere(np.isin(COLUMNS0, LEFT_HAND_NAMES0)).squeeze()\n# obtain indices of the LEFT_HAND_NAMES0 in COLUMNS0\nRIGHT_HAND_IDXS = np.argwhere(np.isin(COLUMNS0, RIGHT_HAND_NAMES0)).squeeze()\nLIPS_IDXS = np.argwhere(np.isin(COLUMNS0, LIPS_NAMES0)).squeeze()\nN_COLS = N_COLS0\n# Only x and y axes are used\nN_DIMS = 2\n\nprint(f'N_COLS: {N_COLS}')","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:08.412601Z","iopub.execute_input":"2023-07-11T08:49:08.41348Z","iopub.status.idle":"2023-07-11T08:49:08.428582Z","shell.execute_reply.started":"2023-07-11T08:49:08.413442Z","shell.execute_reply":"2023-07-11T08:49:08.427251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We create arrays of indices and column names specific to the X and Y coordinates of the left hand in a processed dataset. This can be useful for accessing and manipulating the data of these coordinates more conveniently.","metadata":{}},{"cell_type":"code","source":"# Indices in processed data by axes with only dominant hand\nHAND_X_IDXS = np.array(\n        [idx for idx, name in enumerate(LEFT_HAND_NAMES0) if 'x' in name]\n    ).squeeze()\nHAND_Y_IDXS = np.array(\n        [idx for idx, name in enumerate(LEFT_HAND_NAMES0) if 'y' in name]\n    ).squeeze()\n# Names in processed data by axes\nHAND_X_NAMES = LEFT_HAND_NAMES0[HAND_X_IDXS]\nHAND_Y_NAMES = LEFT_HAND_NAMES0[HAND_Y_IDXS]","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:08.430445Z","iopub.execute_input":"2023-07-11T08:49:08.431214Z","iopub.status.idle":"2023-07-11T08:49:08.439097Z","shell.execute_reply.started":"2023-07-11T08:49:08.431145Z","shell.execute_reply":"2023-07-11T08:49:08.437691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preprocessing\n\nNow, we begin with the preprocessing and define a custom preprocessing layer called `PreprocessLayerNonNaN`. The class inherits from `tf.keras.layers.Layer`, making it a Keras layer. In the `__init__` constructor, no additional operations are performed.\n\nThe `call` method is the core of the layer and is executed when the layer is called with input data. It is decorated with `@tf.function` and specifies an `input_signature` that indicates the shape and type of the input data. `@tf.function` is a decorator in TensorFlow that allows Python functions to be compiled and optimized for execution in computational graphs. By decorating a function with `@tf.function`, TensorFlow analyzes the code of the function and creates an optimized computational graph that can be executed efficiently.\n\n`input_signature` is an argument of tf.function that is used to specify the input signature of the function. In this case, `input_signature=(tf.TensorSpec(shape=[None,N_COLS0], dtype=tf.float32),)` is used to indicate that the function expects a single input argument that should be a float32 tensor with a shape of `[None, N_COLS0]`. The None value in the tensor shape indicates that it can have any size in that dimension.\n\nInside the `call` method, the following preprocessing is performed:\n\n- NaN values in the input data `data0` are filled with zeros using the `tf.where` function and the condition `tf.math.is_nan(data0)`. This ensures that there are no NaN values in the data.\n- Manipulation of the shape of the data is performed to add an additional dimension at the beginning. This is achieved by using `data[None]`, which adds a dimension of size 1 at the beginning of the data.\n- Next, a filtering called \"Empty Hand Frame Filtering\" is applied to the data. Hand frames are extracted from the data using `tf.slice`. The line `hands = tf.slice(data, [0,0,0], [-1, -1, 84])` selects a portion of the data from the beginning to the end in the first two dimensions and the first 84 elements in the third dimension.\n- The absolute value of the frames is then calculated using `tf.abs(hands)`. This is done to ensure that all values are positive. Then, a mask is created by summing the values along the hand frames axis using `tf.reduce_sum(hands, axis=2)`. The mask is then checked if the sum is not equal to zero using `tf.not_equal(mask, 0)`.\n- The mask is used to select only the data frames that contain hand information, and they are assigned to `data` using `data = data[mask][None]`. This filters out the data frames that do not contain hand information.\n- Finally, additional manipulation is performed to remove the previously added additional dimension using `data = tf.squeeze(data, axis=[0])`. `data`, which represents the preprocessed input data, is returned.","metadata":{}},{"cell_type":"code","source":"class PreprocessLayerNonNaN(tf.keras.layers.Layer):\n    def __init__(self):\n        super(PreprocessLayerNonNaN, self).__init__()\n    \n    @tf.function(\n        input_signature=(tf.TensorSpec(shape=[None,N_COLS0], dtype=tf.float32),),\n    )\n    def call(self, data0):\n        # Fill NaN Values With 0\n        data = tf.where(tf.math.is_nan(data0), 0.0, data0)\n        \n        # Add another dimension\n        data = data[None]\n        \n        # Empty Hand Frame Filtering\n        hands = tf.slice(data, [0,0,0], [-1, -1, 84]) # 84 its 21 x 4, left_hand, right_hand and x,y\n        hands = tf.abs(hands)\n        mask = tf.reduce_sum(hands, axis=2)\n        mask = tf.not_equal(mask, 0)\n        data = data[mask][None]\n        data = tf.squeeze(data, axis=[0])\n        \n        return data\n    \npreprocess_layer_non_nan = PreprocessLayerNonNaN()","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:08.440647Z","iopub.execute_input":"2023-07-11T08:49:08.441212Z","iopub.status.idle":"2023-07-11T08:49:08.475002Z","shell.execute_reply.started":"2023-07-11T08:49:08.441167Z","shell.execute_reply":"2023-07-11T08:49:08.474145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.read_parquet(file_path)\nfor group, group_df in df.groupby('sequence_id'):\n    # extract sequence_id = group and the rest of the dataframe = group_df\n    datagroup = group\n    datagroupdf = group_df\n\ndisplay(datagroupdf[COLUMNS0].iloc[:,:84].head())\nprint('\\n')\ndisplay(datagroup)","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:08.476608Z","iopub.execute_input":"2023-07-11T08:49:08.477234Z","iopub.status.idle":"2023-07-11T08:49:11.863733Z","shell.execute_reply.started":"2023-07-11T08:49:08.477204Z","shell.execute_reply":"2023-07-11T08:49:11.862515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Unique Parquet Files\nUNIQUE_FILE_PATHS = pd.Series(train['file_path'].unique())\n# Number of parquet chunks to analyse\nN = 5 if (IS_INTERACTIVE or FAST) else len(UNIQUE_FILE_PATHS)\n# Number of Non Nan Frames in Recording\nN_NON_NAN_FRAMES = []\n\nfor idx, file_path in enumerate(tqdm(UNIQUE_FILE_PATHS.sample(N, random_state=SEED))):\n    df = pd.read_parquet(file_path)\n    for group, group_df in df.groupby('sequence_id'):\n        # group = sequence_id \n        # group_df = the rest of the dataframe (left_hand, right_hand,...)\n        # preprocess the left hand and right hand values for that sequence id\n        frames = preprocess_layer_non_nan(group_df[COLUMNS0].values).numpy()\n        # append the number of frames that are not NaN\n        N_NON_NAN_FRAMES.append(len(frames))\n\n# Convert to Numpy Array\nN_NON_NAN_FRAMES = pd.Series(N_NON_NAN_FRAMES).to_frame('# Frames')","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:11.865251Z","iopub.execute_input":"2023-07-11T08:49:11.865579Z","iopub.status.idle":"2023-07-11T08:49:41.186001Z","shell.execute_reply.started":"2023-07-11T08:49:11.865552Z","shell.execute_reply":"2023-07-11T08:49:41.184834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"N_NON_NAN_FRAMES.plot(kind='hist', bins=128, figsize=(15,8))\nplt.title('Number of Non NaN Frames', size=24)\nplt.grid()\nxlim = np.percentile(N_NON_NAN_FRAMES, 99)\nplt.xlim(0, xlim)\nplt.xticks(np.arange(0, xlim+32, 32))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:41.187751Z","iopub.execute_input":"2023-07-11T08:49:41.188259Z","iopub.status.idle":"2023-07-11T08:49:41.836652Z","shell.execute_reply.started":"2023-07-11T08:49:41.188216Z","shell.execute_reply":"2023-07-11T08:49:41.835472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We redefine the `PreprocessLayerNonNaN` class, adding the padding phase:\n\n- Calculate the current length of the input data `N_FRAMES` using `len(data[0])`.\n- Compare the current length of the data `N_FRAMES` with a target length `N_TARGET_FRAMES`. If the current length is less than the target length, padding with zeros is performed.\n- Use `tf.concat` to concatenate additional zeros at the end of the data to reach the target length. Create a zeros tensor using `tf.zeros` with shape `[1, N_TARGET_FRAMES-N_FRAMES, N_COLS]`, where `N_COLS` represents the number of columns in the input data.\n- Concatenate the zeros to the data using `tf.concat((data, zeros), axis=1)`, where `axis=1` indicates that the concatenation is done along the column axis.\n\nAfter padding, we move to the next stage:\n\n- The data tensor will be resized using the `tf.image.resize()` function with a target size of `[1, N_TARGET_FRAMES]`, which in this case is `[1, 128]`. The resizing function will use the bilinear method to adjust the tensor size. We will end up with a tensor of shape `[1, 128, ...]`.\n- Next, apply the `tf.squeeze()` function to the data tensor with the argument `axis=[0]`. This will remove the batch dimension (1), resulting in a tensor of size `[128, ...]`.","metadata":{}},{"cell_type":"code","source":"class PreprocessLayer(tf.keras.layers.Layer):\n    def __init__(self):\n        super(PreprocessLayer, self).__init__()\n    \n    @tf.function(\n        input_signature=(tf.TensorSpec(shape=[None,N_COLS0], dtype=tf.float32),),\n    )\n    def call(self, data0, resize=True):\n        # Fill NaN Values With 0\n        data = tf.where(tf.math.is_nan(data0), 0.0, data0)\n        \n        # Add another dimension\n        data = data[None]\n        \n        # Empty Hand Frame Filtering\n        hands = tf.slice(data, [0,0,0], [-1, -1, 84])\n        hands = tf.abs(hands)\n        mask = tf.reduce_sum(hands, axis=2)\n        mask = tf.not_equal(mask, 0)\n        data = data[mask][None]\n        \n        # Padding with Zeros\n        N_FRAMES = len(data[0]) # set N_FRAMES as the original length of the data\n        if N_FRAMES < N_TARGET_FRAMES: \n            # N_TARGET_FRAMES = 128\n            # then we use padding \n            data = tf.concat((\n                data,    \n                tf.zeros([1,N_TARGET_FRAMES-N_FRAMES,N_COLS], dtype=tf.float32)\n            ), axis=1)\n        # Downsample\n        data = tf.image.resize(\n            data,\n            [1, N_TARGET_FRAMES], # [1,128]\n            method=tf.image.ResizeMethod.BILINEAR,)\n        \n        # Squeeze Batch Dimension\n        data = tf.squeeze(data, axis=[0])\n        \n        return data\n    \npreprocess_layer = PreprocessLayer()\n\ninputs = group_df[COLUMNS0].values # 164 columns (left_hand,right_hand,face) \n# dimension -> [n_rows, n_columns = 164]\nprint(inputs.shape) \ninputs = inputs[:1] # select the first row for the 164 columns\n# dimension -> [1, 164]\nprint(inputs.shape)\n\nframes = preprocess_layer(inputs) # Use PreprocessLayer in the example parquet\n\nprint(f'inputs shape: {inputs.shape}, NaN count: {np.isnan(inputs).sum()}')\nprint(f'frames shape: {frames.shape}, NaN count: {np.isnan(frames).sum()}')","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:41.838129Z","iopub.execute_input":"2023-07-11T08:49:41.838529Z","iopub.status.idle":"2023-07-11T08:49:42.163519Z","shell.execute_reply.started":"2023-07-11T08:49:41.838499Z","shell.execute_reply":"2023-07-11T08:49:42.162215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" ","metadata":{"execution":{"iopub.status.busy":"2023-07-06T07:40:50.080843Z","iopub.execute_input":"2023-07-06T07:40:50.082112Z","iopub.status.idle":"2023-07-06T07:40:50.08974Z","shell.execute_reply.started":"2023-07-06T07:40:50.082049Z","shell.execute_reply":"2023-07-06T07:40:50.088351Z"}}},{"cell_type":"markdown","source":"As we can see, we went from an array object with 164 values (1 value for each column) and 42 NaN values (approximately 25% null values) to a padded array.\n\nLet's try to perform the process manually to visualize it better:","metadata":{}},{"cell_type":"code","source":"''' define data to preprocess'''\ninputs = group_df[COLUMNS0].values # 164 columns (left_hand,right_hand,face) \n# dimension -> [n_rows, n_columns = 164]\nprint(inputs.shape) \ninputs = inputs[:1] # select the first row for the 164 columns\n# dimension -> [1, 164]\nprint(inputs.shape)","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:42.165137Z","iopub.execute_input":"2023-07-11T08:49:42.16552Z","iopub.status.idle":"2023-07-11T08:49:42.174531Z","shell.execute_reply.started":"2023-07-11T08:49:42.16549Z","shell.execute_reply":"2023-07-11T08:49:42.173158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Fill NaN Values With 0\ndata = tf.where(tf.math.is_nan(inputs), 0.0, inputs)    \nprint(f' elementos 40:45 del array \"inputs\" : {inputs[:,40:45]}')\nprint(f'\\n elementos 40:45 del array \"data\" : {data[:,40:45]}\\n')\n# Add another dimension\ndata = data[None]\nprint(f' shape del array \"data\" : {data.shape}')          ","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:42.176265Z","iopub.execute_input":"2023-07-11T08:49:42.176921Z","iopub.status.idle":"2023-07-11T08:49:42.20466Z","shell.execute_reply.started":"2023-07-11T08:49:42.176877Z","shell.execute_reply":"2023-07-11T08:49:42.203264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, we have changed the 0's to 1's and added a new dimension to the left of the array. Now, let's look at the empty hand filtering:","metadata":{}},{"cell_type":"code","source":"# Empty Hand Frame Filtering\nhands = tf.slice(data, [0,0,0], [-1, -1, 84])\nprint(hands) # from al the data (face and hands)\n             # we select the hands (last 84 elements in 3rd dimension)","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:42.206163Z","iopub.execute_input":"2023-07-11T08:49:42.206553Z","iopub.status.idle":"2023-07-11T08:49:42.216848Z","shell.execute_reply.started":"2023-07-11T08:49:42.206524Z","shell.execute_reply":"2023-07-11T08:49:42.215381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"hands = tf.abs(hands) # original values [-1,1] -> make then [0,1] to make the sum\nmask = tf.reduce_sum(hands, axis=2) # sum of all the values\nprint(f'\\n')\nmask = tf.not_equal(mask, 0) # is this is True, this frame contains hand information\nprint(f'\\n does the frame contain information? : {mask}') \nprint(f'\\n is bool value its True : {data[mask][None]}') # select or not the values\nprint(f'\\n is bool value its False : {data[tf.fill(data.shape, False)][None]}')","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:42.218644Z","iopub.execute_input":"2023-07-11T08:49:42.219009Z","iopub.status.idle":"2023-07-11T08:49:42.259373Z","shell.execute_reply.started":"2023-07-11T08:49:42.21898Z","shell.execute_reply":"2023-07-11T08:49:42.258116Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, this part of the code determines whether the video frame contains any hand gesture information (if at least one element is not null). If this is the case, we save the object; otherwise, we discard it (keeping an empty object).\n\nNow, let's move on to the padding part:","metadata":{}},{"cell_type":"code","source":"data = data[mask][None] # select or not the values\n\n# Padding with Zeros\nN_FRAMES = len(data[0]) # N_FRAMES as the original length of the data \nprint(f'number of frames of the video : {N_FRAMES}')\nprint('\\n')\nif N_FRAMES < N_TARGET_FRAMES: \n    # N_TARGET_FRAMES = 128\n    # then we use padding \n    print(f'Original shape of the data pre concat : {data.shape}')\n    newdata = tf.concat((\n                data,   # [1,1,164]  \n                        # [1,128-1,164] -> concat the two produces [1,128,164]\n                tf.zeros([1,N_TARGET_FRAMES-N_FRAMES,N_COLS], dtype=tf.float32)),\n                axis=1) # will add 127 in 2nd dimension full of zeros \n    print(f'\\n Resulting tensor post concat : {newdata.shape}')\n    print('\\n')\n    print(f'\\n First vector in 2nd dimension : {newdata[:,0,:]}')\n    print(f'\\n  Rest of the vectors (127) : {newdata[:,10,:]}')\n    \n    # Now we have a Tensor with shape [1,128,164]\n    \n    # Downsample\n    newdata = tf.image.resize(\n            newdata,\n            [1, N_TARGET_FRAMES], # [1,128,164]\n            method=tf.image.ResizeMethod.BILINEAR,)\n    # in this particular case it doesnt change nothing,\n    # because our tensor was already [1,128,164]\n    \n    # Delete First Dimension (Batch dimension)\n    newdata = tf.squeeze(newdata, axis=[0])\nprint(f'\\n Final result : {newdata.shape}')  ","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:42.260981Z","iopub.execute_input":"2023-07-11T08:49:42.26145Z","iopub.status.idle":"2023-07-11T08:49:42.298893Z","shell.execute_reply.started":"2023-07-11T08:49:42.26141Z","shell.execute_reply":"2023-07-11T08:49:42.296672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Converts any tensor to one with shape `[128,164]`.\n\n# Create X/y","metadata":{"execution":{"iopub.status.busy":"2023-07-06T09:23:23.629438Z","iopub.execute_input":"2023-07-06T09:23:23.629871Z","iopub.status.idle":"2023-07-06T09:23:23.638758Z","shell.execute_reply.started":"2023-07-06T09:23:23.629841Z","shell.execute_reply":"2023-07-06T09:23:23.637405Z"}}},{"cell_type":"code","source":"# Number Of Train Samples\nN_SAMPLES = len(train)\nprint(f'N_SAMPLES: {N_SAMPLES}')\n\n# Target Arrays Processed Input Videos\nX = np.zeros([N_SAMPLES, N_TARGET_FRAMES, N_COLS], dtype=np.float32)\nprint(f'\\nX Shape : {X.shape}')\n# Ordinally Encoded Target With value 59 for pad token\ny = np.full(shape=[N_SAMPLES, N_TARGET_FRAMES], fill_value=N_UNIQUE_CHARACTERS, dtype=np.int8)\nprint(f'\\ny Shape : {y.shape}')\n# Phrase Type\ny_phrase_type = np.empty(shape=[N_SAMPLES], dtype=object)\nprint(f'\\ny_phrase_type Shape : {y_phrase_type.shape}')","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:42.300325Z","iopub.execute_input":"2023-07-11T08:49:42.300704Z","iopub.status.idle":"2023-07-11T08:49:42.310821Z","shell.execute_reply.started":"2023-07-11T08:49:42.300673Z","shell.execute_reply":"2023-07-11T08:49:42.309433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:42.312524Z","iopub.execute_input":"2023-07-11T08:49:42.312975Z","iopub.status.idle":"2023-07-11T08:49:42.333208Z","shell.execute_reply.started":"2023-07-11T08:49:42.312928Z","shell.execute_reply":"2023-07-11T08:49:42.331944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We now proceed to create the `X` and `y` datasets, which we will work with later:\n\n- First, we define the `UNIQUE_FILE_PATHS` and `N_UNIQUE_FILE_PATHS` variables, which contain all the unique file paths and the number of existing paths, respectively.\n- We define `row` and `counter` as counters that we will use in the loop. `counter` represents the total number of iterations, while `row` will represent only those rows for which the number of frames (rows in the Parquet file) per character (`parquet_rows/characters`) is greater than the minimum number, `MIN_NUM_FRAMES_PER_CHARACTER=4`.\n- We define two empty lists, `N_FRAMES_PER_CHARACTER` to store the number of frames per character, and `VALID_IDXS` to store the different indices.","metadata":{}},{"cell_type":"code","source":"# All unique parquet files\nUNIQUE_FILE_PATHS = pd.Series(train['file_path'].unique())\nN_UNIQUE_FILE_PATHS = len(UNIQUE_FILE_PATHS)\n# Counter to keep track of sample\nrow = 0\ncount = 0\n# Compressed Parquet Files\nPath('train_landmark_subsets').mkdir(parents=True, exist_ok=True)\n# Number Of Frames Per Character\nN_FRAMES_PER_CHARACTER = []\n# Minimum Number Of Frames Per Character\nMIN_NUM_FRAMES_PER_CHARACTER = 4\nVALID_IDXS = []","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:42.334827Z","iopub.execute_input":"2023-07-11T08:49:42.335446Z","iopub.status.idle":"2023-07-11T08:49:42.358887Z","shell.execute_reply.started":"2023-07-11T08:49:42.335399Z","shell.execute_reply":"2023-07-11T08:49:42.357836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Once these variables are defined, we create a loop that iterates over the different indices of the `train` DataFrame and the corresponding Parquet file paths:\n\n- First, we read the Parquet file corresponding to the `file_path`.\n- We save the file name using `file_path.split('/')[-1]` (for example, `105143404.parquet`).\n- For the first 10 index values, we select the columns of interest from the DataFrame and convert it back to Parquet format to save it in the `train_landmark_subsets` folder.","metadata":{}},{"cell_type":"code","source":"i = 0\nfor idx, file_path in enumerate((UNIQUE_FILE_PATHS)):\n    print(f'Index of the dataframe : {idx}')\n    print(f'File path : {file_path}')\n    # read the parquet corresponding to that path\n    df = pd.read_parquet(file_path)\n    # Save COLUMN Subset of parquet files for TFLite Model verficiation\n    name = file_path.split('/')[-1] # parquet name (ex : 105143404.parquet)\n    print(f'File name : {name}')\n    if idx < 10: # first 10 values\n        display(df[COLUMNS0].head())\n        # make it parquet again and save it in /kaggle/working/train_landmark_subsets\n        df[COLUMNS0].to_parquet(f'train_landmark_subsets/{name}', engine='pyarrow', compression='zstd')\n    # this is just to stop at three iters\n    print('\\n\\n')\n    i += 1 \n    if i == 1:\n        break","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:42.360513Z","iopub.execute_input":"2023-07-11T08:49:42.361283Z","iopub.status.idle":"2023-07-11T08:49:48.387857Z","shell.execute_reply.started":"2023-07-11T08:49:42.361251Z","shell.execute_reply":"2023-07-11T08:49:48.386661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we add another loop that iterates over the DataFrames (Parquet files) grouped by `sequence_id`. We extract the `sequence_id` column and the remaining columns from the DataFrame. We calculate the number of rows (frames) per character length of the associated phrase, `n_frames_per_character`. If it is greater than a minimum length, in this case 4, we consider it valid and save it.","metadata":{}},{"cell_type":"code","source":"# All unique parquet files\nUNIQUE_FILE_PATHS = pd.Series(train['file_path'].unique())\nN_UNIQUE_FILE_PATHS = len(UNIQUE_FILE_PATHS)\n# Counter to keep track of sample\nrow = 0\ncount = 0\n# Compressed Parquet Files\nPath('train_landmark_subsets').mkdir(parents=True, exist_ok=True)\n# Number Of Frames Per Character\nN_FRAMES_PER_CHARACTER = []\n# Minimum Number Of Frames Per Character\nMIN_NUM_FRAMES_PER_CHARACTER = 4\nVALID_IDXS = []\n\nfor idx, file_path in enumerate((UNIQUE_FILE_PATHS)):\n    # read the parquet corresponding to that path\n    df = pd.read_parquet(file_path)\n    # save COLUMN Subset of parquet files for TFLite Model verification\n    name = file_path.split('/')[-1] # parquet name (ex: 105143404.parquet)\n    if idx < 10: # first 10 values\n        # make it parquet again and save it in /kaggle/working/train_landmark_subsets\n        df[COLUMNS0].to_parquet(f'train_landmark_subsets/{name}', engine='pyarrow', compression='zstd')\n    # iterate over samples\n    for group, group_df in df.groupby('sequence_id'):\n        # Number of Frames Per Character\n        print(f'count: {count}')\n        print(f'group (sequence_id): {group}')\n        print(f'group_df[:1,:5] : \\n')\n        print(f'----------------------------------------------------------')\n        print(f'{group_df.iloc[:1,:5]}\\n')\n        print(f'----------------------------------------------------------')\n        print(f'shape of the dataframe : {group_df.shape}')\n        print(f'len group_df[COLUMNS0].values : {len(group_df[COLUMNS0].values)}')\n        print(f'len char : {len(train_sequence_id.loc[group, \"phrase_char\"]) }')\n        n_frames_per_character =  len(group_df[COLUMNS0].values) / len(train_sequence_id.loc[group, 'phrase_char'])\n        print(f'frames per character: {n_frames_per_character}')\n        N_FRAMES_PER_CHARACTER.append(n_frames_per_character)\n        \n        if n_frames_per_character < MIN_NUM_FRAMES_PER_CHARACTER:\n            count = count + 1\n            continue\n        else:\n            # Add Valid Index\n            VALID_IDXS.append(count)\n            print(f'valid indexes: {VALID_IDXS}')\n            print('\\n\\n')\n            count = count + 1\n        \n        if count == 2:\n            break\n    if count == 2:\n        break","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:48.389438Z","iopub.execute_input":"2023-07-11T08:49:48.390349Z","iopub.status.idle":"2023-07-11T08:49:55.162499Z","shell.execute_reply.started":"2023-07-11T08:49:48.39031Z","shell.execute_reply":"2023-07-11T08:49:55.161243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, we process the DataFrame (Parquet) with the `PreprocessLayer` class to standardize the length of the elements and remove NAs by replacing them with 0s, and save it in the `frames` variable. The frames will have a shape of `[128, 164]`, where 128 is the target number of frames, `N_TARGET_FRAMES`, and 164 is the number of columns we kept. Then, we save the `frames` variable in the `row` row of the `X` variable, and in that same row of the `y` variable, we add the numbers or tokens corresponding to each character, with the `END_TOKEN=61` inserted at the end of the phrase. The rest of the values in that row (the rest of the columns) have the `PAD_TOKEN=59` (padding token) assigned. Additionally, we save the phrase type in the `y_phrase_type` variable.","metadata":{}},{"cell_type":"code","source":"# All unique parquet files\nUNIQUE_FILE_PATHS = pd.Series(train['file_path'].unique())\nN_UNIQUE_FILE_PATHS = len(UNIQUE_FILE_PATHS)\n# Counter to keep track of sample\nrow = 0\ncount = 0\n# Compressed Parquet Files\nPath('train_landmark_subsets').mkdir(parents=True, exist_ok=True)\n# Number Of Frames Per Character\nN_FRAMES_PER_CHARACTER = []\n# Minimum Number Of Frames Per Character\nMIN_NUM_FRAMES_PER_CHARACTER = 4\nVALID_IDXS = []\n\n# Fill Arrays\ni = 0\nstop = False\nfor idx, file_path in enumerate(tqdm(UNIQUE_FILE_PATHS)):\n    # Print the progress\n    print(f'Processed {idx:02d}/{N_UNIQUE_FILE_PATHS} parquet files')\n    # Read parquet file\n    df = pd.read_parquet(file_path)\n    # Save COLUMN Subset of parquet files for TFLite Model verification\n    name = file_path.split('/')[-1] # parquet name (ex : 105143404.parquet)\n    if idx < 10: # first 10 values\n        df[COLUMNS0].to_parquet(f'train_landmark_subsets/{name}', engine='pyarrow', compression='zstd')\n    # Iterate Over Samples\n    for group, group_df in df.groupby('sequence_id'):\n        # Number of Frames Per Character\n        n_frames_per_character =  len(group_df[COLUMNS0].values) / len(train_sequence_id.loc[group, 'phrase_char'])\n        N_FRAMES_PER_CHARACTER.append(n_frames_per_character)\n        if n_frames_per_character < MIN_NUM_FRAMES_PER_CHARACTER:\n            count = count + 1\n            continue\n        else:\n            # Add Valid Index\n            VALID_IDXS.append(count)\n            count = count + 1\n        \n        # Get Processed Frames and non empty frame indices\n        frames = preprocess_layer(group_df[COLUMNS0].values)\n        assert frames.ndim == 2\n        print(f'frames shape : {frames.shape}\\n')\n        # Assign\n        X[row] = frames # each row contains a 128,164 tensor\n        # Add Target By Ordinally Encoding Characters\n        phrase_char = train_sequence_id.loc[group, 'phrase_char']\n        for col, char in enumerate(phrase_char):\n            # for every column and character\n            y[row, col] = CHAR2ORD.get(char) # set in corresponding row \n            \n        # Add End of Sentence Token = 61\n        y[row, col+1] = END_TOKEN # 61\n        print(f'y[row,:] :')\n        print(f'----------------------------------------------------------')\n        print(f'{y[row,:]}\\n')\n        print(f'----------------------------------------------------------')\n        # Phrase Type\n        y_phrase_type[row] = train_sequence_id.loc[group, 'phrase_type']\n        print(f'y_phrase_type[row] : {y_phrase_type[row]}\\n')\n        # Row Count\n        row += 1\n        \n        # Check if 3 iterations of the second loop have been completed\n        i += 1\n        if i == 3:\n            stop = True\n            break\n    \n    if stop:\n        break\n    \n    print('\\n\\n')\n    # Clean up\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:49:55.169733Z","iopub.execute_input":"2023-07-11T08:49:55.170191Z","iopub.status.idle":"2023-07-11T08:50:02.198737Z","shell.execute_reply.started":"2023-07-11T08:49:55.170145Z","shell.execute_reply":"2023-07-11T08:50:02.19748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With every step explained, we can now create the variables `X`, `y`, `y_phrase_type`:","metadata":{}},{"cell_type":"code","source":"# All unique parquet files\nUNIQUE_FILE_PATHS = pd.Series(train['file_path'].unique())\nN_UNIQUE_FILE_PATHS = len(UNIQUE_FILE_PATHS)\n# Counter to keep track of sample\nrow = 0\ncount = 0\n# Compressed Parquet Files\nPath('train_landmark_subsets').mkdir(parents=True, exist_ok=True)\n# Number Of Frames Per Character\nN_FRAMES_PER_CHARACTER = []\n# Minimum Number Of Frames Per Character\nMIN_NUM_FRAMES_PER_CHARACTER = 4\nVALID_IDXS = []\n\n# Fill Arrays\nfor idx, file_path in enumerate(tqdm(UNIQUE_FILE_PATHS)):\n    # Print the progress\n    print(f'Processed {idx:02d}/{N_UNIQUE_FILE_PATHS} parquet files')\n    # Read parquet file\n    df = pd.read_parquet(file_path)\n    # Save COLUMN Subset of parquet files for TFLite Model verficiation\n    name = file_path.split('/')[-1] # parquet name (ex : 105143404.parquet)\n    if idx < 10: # first 10 values\n        df[COLUMNS0].to_parquet(f'train_landmark_subsets/{name}', engine='pyarrow', compression='zstd')\n    # Iterate Over Samples\n    for group, group_df in df.groupby('sequence_id'):\n        # Number of Frames Per Character\n        n_frames_per_character =  len(group_df[COLUMNS0].values) / len(train_sequence_id.loc[group, 'phrase_char'])\n        N_FRAMES_PER_CHARACTER.append(n_frames_per_character)\n        if n_frames_per_character < MIN_NUM_FRAMES_PER_CHARACTER:\n            count = count + 1\n            continue\n        else:\n            # Add Valid Index\n            VALID_IDXS.append(count)\n            count = count + 1\n        \n        # Get Processed Frames and non empty frame indices\n        frames = preprocess_layer(group_df[COLUMNS0].values)\n        assert frames.ndim == 2\n        # Assign\n        X[row] = frames\n        # Add Target By Ordinally Encoding Characters\n        phrase_char = train_sequence_id.loc[group, 'phrase_char']\n        for col, char in enumerate(phrase_char):\n            y[row, col] = CHAR2ORD.get(char)\n        # Add EOS Token\n        y[row, col+1] = END_TOKEN\n        # Phrase Type\n        y_phrase_type[row] = train_sequence_id.loc[group, 'phrase_type']\n        # Row Count\n        row += 1\n    # clean up\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-07-11T08:50:02.200439Z","iopub.execute_input":"2023-07-11T08:50:02.200778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# rows denotes the number of samples with frames/character above threshold\n# count is the total \nprint(f'row: {row}, count: {count}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filter X/y\nX = X[:row]\ny = y[:row]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we save the `X` and `y` data into *numpy* (`*.npy`) files using the `np.save()` function. Next, we use the `GroupShuffleSplit` class from *scikit-learn* to split the data into training and validation sets based on the `participant_id`. Then, we save the training and validation sets into separate *numpy* files. Finally, we perform a check to ensure that there is no overlap of participant IDs between the training and validation sets, and print out the size of the training and validation sets.","metadata":{}},{"cell_type":"code","source":"# Save X/y\nnp.save('X.npy', X)\nnp.save('y.npy', y)\n# Save Validation\nsplitter = GroupShuffleSplit(test_size=0.10, n_splits=2, random_state=SEED)\nPARTICIPANT_IDS = train['participant_id'].values[VALID_IDXS]\ntrain_idxs, val_idxs = next(splitter.split(X, y, groups=PARTICIPANT_IDS))\n\n# Save Train\nnp.save('X_train.npy', X[train_idxs])\nnp.save('y_train.npy', y[train_idxs])\n# Save Validation\nnp.save('X_val.npy', X[val_idxs])\nnp.save('y_val.npy', y[val_idxs])\n# Verify Train/Val is correctly split by participan id\nprint(f'Patient ID Intersection Train/Val: {set(PARTICIPANT_IDS[train_idxs]).intersection(PARTICIPANT_IDS[val_idxs])}')\n# Train/Val Sizes\nprint(f'# Train Samples: {len(train_idxs)}, # Val Samples: {len(val_idxs)}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Number of Frames per Phrase Character","metadata":{}},{"cell_type":"code","source":"N_FRAMES_PER_CHARACTER_S = pd.Series(N_FRAMES_PER_CHARACTER)\n\nplt.figure(figsize=(20,10))\nplt.title('Number Of Frames Per Phrase Character')\nN_FRAMES_PER_CHARACTER_S.plot(kind='hist', bins=128)\n# Plot till 99th percentile\np99 = math.ceil(np.percentile(N_FRAMES_PER_CHARACTER_S, 99))\nplt.xticks(np.arange(0, p99+1, 1))\nplt.xlim(0, p99)\nplt.xlabel('Number Of Frames Per Phrase Character')\nplt.ylabel('Sample Count')\nplt.grid()\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Dominant Hand Statistics\n\nWe calculate the mean and standard deviation of the dominant hand in the input data. We plot a boxplot to visualize the distribution of values in the left hand, right hand, and lips. Finally, we save the means and standard deviations into numpy files to normalize the input data for a neural network model.","metadata":{}},{"cell_type":"code","source":"def get_left_right_hand_mean_std():\n    # Dominant Hand Statistics\n    MEANS = np.zeros([N_COLS], dtype=np.float32)\n    STDS = np.zeros([N_COLS], dtype=np.float32)\n    \n    # Plot\n    fig, axes = plt.subplots(3, figsize=(20, 3*8))\n    \n    # Iterate over all landmarks\n    for col, v in enumerate(tqdm(X.reshape([-1, N_COLS]).T)):\n        v = v[np.nonzero(v)]\n        # Remove zero values as they are NaN values\n        MEANS[col] = v.astype(np.float32).mean()\n        STDS[col] = v.astype(np.float32).std()\n        if col in LEFT_HAND_IDXS:\n            axes[0].boxplot(v, notch=False, showfliers=False, positions=[col], whis=[5,95])\n        elif col in RIGHT_HAND_IDXS:\n            axes[1].boxplot(v, notch=False, showfliers=False, positions=[col], whis=[5,95])\n        else:\n            axes[2].boxplot(v, notch=False, showfliers=False, positions=[col], whis=[5,95])\n        \n    for ax, name in zip(axes, ['Left Hand', 'Right Hand', 'Lips']):\n        ax.set_title(f'{name}', size=24)\n        ax.tick_params(axis='x', labelsize=8, rotation=45)\n        ax.set_ylim(0.0, 1.0)\n        ax.grid(axis='y')\n\n    plt.show()\n    \n    return MEANS, STDS\n\n# Get Dominant Hand Mean/Standard Deviation\nMEANS, STDS = get_left_right_hand_mean_std()\n# Save Mean/STD to normalize input in neural network model\nnp.save('MEANS.npy', MEANS)\nnp.save('STDS.npy', STDS)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we have the parquet files processed, so we're ready to train with the them in the [Modelization Notebook](https://www.kaggle.com/code/m4nugnzl/aslfr-transformer-training-fully-explained).","metadata":{}}]}