{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<p style=\"background-color:#ffffff;font-family:candaralight;color:#C15D06;font-size:215%;text-align:center;border-radius:10px 10px;\"> Google ASL Fingerspelling</p>\n<p style=\"background-color:#ffffff;font-family:candaralight;color:#B0B0B0;font-size:150%;text-align:center;border-radius:10px 10px;\">✋ EDA and PREPROCESSING ✋</p>\n\n<div style=\"width:100%;text-align: center;\"> <img align=middle src=\"https://media1.giphy.com/media/Co5TKVg51CmFsxPpNP/giphy.webp\" alt=\"Heat beating\" > </div>\n\n\nThe Google ASL Finger Recognition dataset is a collection of hand landmarks captured from videos of people performing American Sign Language (ASL) finger spelling. The goal of this dataset is to develop a machine learning model that can accurately predict the corresponding ASL finger spelling from the hand landmarks.\n\nIn this Kaggle notebook, we will explore the data processing and exploratory data analysis (EDA) steps for this task. The dataset is provided in the form of parquet files containing the hand landmarks for each video sequence. Our objective is to process these files, extract meaningful features, and prepare the data for training a machine learning model\n        \n   <center><div class=\"alert alert-block alert-warning\" style=\"margin: 2em; line-height: 1.7em; font-family: candaralight;\">\n    <b style=\"font-size: 18px;\">👏 &nbsp; IF YOU FORK THIS OR FIND THIS HELPFUL &nbsp; 👏</b><br><br><b style=\"font-size: 22px; color: darkorange\">PLEASE UPVOTE!</b><br><br>This was a lot of work for me and while it may seem silly, it makes me feel appreciated when others like my work. 😅\n</div></center>\n    \n\n<a id=\"toc\"></a>\n\n<br><br>\n\n# <p style=\"font-family: candaralight; font-size: 28px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; color: #C15D06; background-color: #ffffff;\">TABLE OF CONTENTS</p>\n\n* [1. DATA OVERVIEW](#1)\n    \n    - [Import Libraries](#1.1)\n    \n    - [Loading dataset](#1.2)\n    \n    - [Data Description](#1.3)\n        \n* [2. Exploratory Data Analysis](#2)\n    \n    - [Inspect the PATH column](#2.1)\n\n    - [Inspect the PARTICIPANT_ID Column](#2.2)\n    \n    - [Inspect the SEQUENCE_ID Column](#2.3)\n    \n    -[Inspect the PHASES Column](#2.4)\n    \n    -[Inspect the PARQUET Column](#2.5)\n    \n* [3. PREPROCESSING DATA](#3)\n\n* [4. Go Further](#4)\n\n    - [Calculating Levenshtein Distance](#4.1)\n    \n    - [Data Preparation and Splitting](#4.2)\n    \n    - [Model Definition](#4.3)\n    \n    - [Optimization Setup](#4.4)\n    \n    - [Training Loop](#4.5)\n    \n    - [Evaluation](#4.6)   \n        \n<a id=\"1\"></a>\n# <p style=\"font-family: candaralight; font-size: 28px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; color: #C15D06; background-color: #ffffff;\">1. DATA OVERVIEW</p>\n\n\nIn this notebook, we will be exploring and analyzing the Apple stock prices dataset. We will start by importing important libraries, loading the data, and giving a brief description of the dataset.\n\n<a id=\"1.1\"></a>\n<br>\n\n<h3 style=\"font-family: candaralight; font-size: 20px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: #C15D06; background-color: #ffffff;\">1.1 <b>Import</b> Libraries</h3>\n\n---\n\n","metadata":{}},{"cell_type":"code","source":"import gc\n\nimport json\nfrom tqdm import tqdm\n\nimport torch\nimport torch.nn as nn\n\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import accuracy_score\n\nimport warnings\nwarnings.filterwarnings(action='ignore')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-06-20T17:39:58.51951Z","iopub.execute_input":"2023-06-20T17:39:58.519919Z","iopub.status.idle":"2023-06-20T17:40:00.506025Z","shell.execute_reply.started":"2023-06-20T17:39:58.519885Z","shell.execute_reply":"2023-06-20T17:40:00.505034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# import data processing and visualisation libraries\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport plotly.io as pio\npio.templates.default = \"simple_white\"\n\n# import tensorflow and keras\nimport tensorflow as tf\nimport os\n\nprint(\"Packages imported...\")","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:08.349229Z","iopub.execute_input":"2023-06-20T17:37:08.349575Z","iopub.status.idle":"2023-06-20T17:37:17.21298Z","shell.execute_reply.started":"2023-06-20T17:37:08.349545Z","shell.execute_reply":"2023-06-20T17:37:17.211937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nLet's load the training and supplemental_metadata dataframe\n<a id=\"1.2\"></a>\n\n\n<h3 style=\"font-family: candaralight; font-size: 20px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: #C15D06; background-color: #ffffff;\">1.2 <b>Loading</b> Dataset</h3>\n\n---\n","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(\"/kaggle/input/asl-fingerspelling/train.csv\")\nmetadata = pd.read_csv(\"/kaggle/input/asl-fingerspelling/supplemental_metadata.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:17.214992Z","iopub.execute_input":"2023-06-20T17:37:17.215823Z","iopub.status.idle":"2023-06-20T17:37:17.464142Z","shell.execute_reply.started":"2023-06-20T17:37:17.21576Z","shell.execute_reply":"2023-06-20T17:37:17.463213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"1.3\"></a>\n\n<h3 style=\"font-family: candaralight; font-size: 20px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: #C15D06; background-color: #ffffff;\">1.3 <b>Data</b> Description</h3>\n\n---","metadata":{}},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:19.734985Z","iopub.execute_input":"2023-06-20T17:37:19.735804Z","iopub.status.idle":"2023-06-20T17:37:19.755154Z","shell.execute_reply.started":"2023-06-20T17:37:19.735749Z","shell.execute_reply":"2023-06-20T17:37:19.754012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The phrases in the training set contains random websites/addresses/phone numbers.","metadata":{}},{"cell_type":"code","source":"print(f\"Total number of files : {df.shape[0]}\")\nprint(f\"Total number of Participant in the dataset : {df.participant_id.nunique()}\")\nprint(f\"Total number of unique phrases : {df.phrase.nunique()}\")","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:20.346704Z","iopub.execute_input":"2023-06-20T17:37:20.347476Z","iopub.status.idle":"2023-06-20T17:37:20.373806Z","shell.execute_reply.started":"2023-06-20T17:37:20.347442Z","shell.execute_reply":"2023-06-20T17:37:20.372713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"metadata.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:20.761679Z","iopub.execute_input":"2023-06-20T17:37:20.762381Z","iopub.status.idle":"2023-06-20T17:37:20.77518Z","shell.execute_reply.started":"2023-06-20T17:37:20.762347Z","shell.execute_reply":"2023-06-20T17:37:20.773927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" the metadata phrases are mostly normal sentences!\n","metadata":{}},{"cell_type":"code","source":"print(f\"Total number of files in metadata: {metadata.shape[0]}\")\nprint(f\"Total number of participants in metadata : {metadata.participant_id.nunique()}\")\nprint(f\"Total number of unique phrases in metadata : {metadata.phrase.nunique()}\")","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:21.239079Z","iopub.execute_input":"2023-06-20T17:37:21.240035Z","iopub.status.idle":"2023-06-20T17:37:21.253141Z","shell.execute_reply.started":"2023-06-20T17:37:21.239992Z","shell.execute_reply":"2023-06-20T17:37:21.252099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The data consist of:    \n\n    path - The path to the landmark file.\n    file_id - A unique identifier for the data file.\n    participant_id - A unique identifier for the data contributor.\n    sequence_id - A unique identifier for the landmark sequence. Each data file may contain many sequences.\n    phrase - The labels for the landmark sequence. The train and test datasets contain randomly generated addresses, phone numbers, and urls derived from components of real addresses/phone numbers/urls.\n    \n**Lets dig deeper into the dataframe info**","metadata":{}},{"cell_type":"code","source":"# Check the dimensions of the dataset\ndf.shape\n\n# Check the data types of columns\ndf.info()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:21.688925Z","iopub.execute_input":"2023-06-20T17:37:21.690129Z","iopub.status.idle":"2023-06-20T17:37:21.743582Z","shell.execute_reply.started":"2023-06-20T17:37:21.69009Z","shell.execute_reply":"2023-06-20T17:37:21.742514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.describe()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:21.910255Z","iopub.execute_input":"2023-06-20T17:37:21.910603Z","iopub.status.idle":"2023-06-20T17:37:21.939614Z","shell.execute_reply.started":"2023-06-20T17:37:21.910573Z","shell.execute_reply":"2023-06-20T17:37:21.938539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2\"></a>\n\n# <p style=\"font-family: candaralight; font-size: 28px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; color: #C15D06; background-color: #ffffff;\">2. EXPLORATORY DATA ANALYSIS</p>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"2.1\"></a>\n\n<h3 style=\"font-family: candaralight; font-size: 20px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: #C15D06; background-color: #ffffff;\">2.1 Inspect the <b>'PATH'</b> Column</h3>\n\n---\n","metadata":{}},{"cell_type":"code","source":"np.array(list(df[\"path\"].value_counts().to_dict().values())).min()\ndf[\"path\"].describe().to_frame().T","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:22.809514Z","iopub.execute_input":"2023-06-20T17:37:22.809932Z","iopub.status.idle":"2023-06-20T17:37:22.84591Z","shell.execute_reply.started":"2023-06-20T17:37:22.8099Z","shell.execute_reply":"2023-06-20T17:37:22.844912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The path column is simply the path to the landmark file (parquet).\n<ul>\n    <li><b>Number unique paths</b>: 68</li>\n    <li><b>Minimum Number of repeated path is</b>: 287</li>\n    <li><b>Maximum Number of repeated path is</b>: 1000</li>\n</ul>\n\n<a id=\"2.2\"></a>\n\n<h3 style=\"font-family: candaralight; font-size: 20px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: #C15D06; background-color: #ffffff;\">2.2 Inspect the <b>`PARTICIPANT_ID`</b> Column</h3>\n\n---\n\n","metadata":{}},{"cell_type":"code","source":"df[\"participant_id\"].astype(str).describe().to_frame().T","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:23.302604Z","iopub.execute_input":"2023-06-20T17:37:23.302965Z","iopub.status.idle":"2023-06-20T17:37:23.373736Z","shell.execute_reply.started":"2023-06-20T17:37:23.302935Z","shell.execute_reply":"2023-06-20T17:37:23.372823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The participant_id statistics indicate a varied distribution of data contributions among participants, with some participants contributing more examples than others.\n\n<ul>\n    <li><b>Number of Unique Participants</b>: 94</li>\n    <li><b>Average Number of Rows Per Participant</b>: 715.82</li>\n    <li><b>Standard Deviation in Counts Per Participant</b>: 230.86</li>\n    <li><b>Minimum Number of Examples For One Participant</b>: 1</li>\n    <li><b>Maximum Number of Examples For One Participant</b>: 1537</li>\n</ul>","metadata":{}},{"cell_type":"code","source":"# Set the custom color scheme\ncolor_scheme = [\"#4f000b\", \"#720026\", \"#ce4257\", \"#ff7f51\", \"#ff9b54\"]\n\n#The column is set to strings as it is an ID\ndf[\"participant_id\"] = df[\"participant_id\"].astype(str)\n\n# Calculate the counts for each participant_id\ncounts = df[\"participant_id\"].value_counts()\n\n# Set up the figure and axes\nfig, ax = plt.subplots(figsize=(15, 6))\n\n# Plot the histogram\nbars = ax.bar(counts.index, counts.values, color=color_scheme[1])\n\n# Set the labels and title\nax.set_xlabel(\"Participant ID\")\nax.set_ylabel(\"Total Row Count\")\nax.set_title(\"Row Counts by Participant ID\")\n\n# Rotate the x-axis labels if needed\nplt.xticks(rotation=90, ha='center')\nplt.xlim(-1, 94)\n\n# Show the plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:23.800344Z","iopub.execute_input":"2023-06-20T17:37:23.800694Z","iopub.status.idle":"2023-06-20T17:37:24.683609Z","shell.execute_reply.started":"2023-06-20T17:37:23.800665Z","shell.execute_reply":"2023-06-20T17:37:24.682721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2.3\"></a>\n\n<h3 style=\"font-family: candaralight; font-size: 20px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: #C15D06; background-color: #ffffff;\">2.3 Inspect the <b>`SEQUENCE_ID`</b> Column</h3>\n\n---\n","metadata":{}},{"cell_type":"code","source":"df[\"sequence_id\"].astype(str).describe().to_frame().T","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:24.685483Z","iopub.execute_input":"2023-06-20T17:37:24.686067Z","iopub.status.idle":"2023-06-20T17:37:24.783709Z","shell.execute_reply.started":"2023-06-20T17:37:24.686034Z","shell.execute_reply":"2023-06-20T17:37:24.78276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A unique identifier for the landmark sequence. Each data file may contain many sequences. Every value is unique for every row\n\n<a id=\"2.4\"></a>\n\n<h3 style=\"font-family: candaralight; font-size: 20px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: #C15D06; background-color: #ffffff;\">2.4 Inspect the <b>`PHASES`</b> Column</h3>\n\n---\n\nHow long are the phrases?\n","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\n\ndf['phrase_len'] = df.phrase.str.len()\nmetadata['phrase_len'] = metadata.phrase.str.len()\n\nfor param in ['text.color', 'axes.labelcolor', 'xtick.color', 'ytick.color']:\n    plt.rcParams[param] = '#000000'  # very light grey\n\nfor param in ['figure.facecolor', 'axes.facecolor', 'savefig.facecolor']:\n    plt.rcParams[param] = '#ffffff'  # bluish dark grey\n\nfig, axs = plt.subplots(1, 1, figsize=(10, 7), tight_layout=True)\n\n# Remove axes splines\nfor s in ['top', 'bottom', 'left', 'right']:\n    axs.spines[s].set_visible(False)\n\n# Remove x, y ticks\naxs.xaxis.set_ticks_position('none')\naxs.yaxis.set_ticks_position('none')\n\n# Add padding between axes and labels\naxs.xaxis.set_tick_params(pad=5)\naxs.yaxis.set_tick_params(pad=10)\n\n# Set the custom color scheme\ncolor_scheme = [\"#4f000b\", \"#720026\", \"#ce4257\", \"#ff7f51\", \"#ff9b54\"]\n\n# Add x, y gridlines\naxs.grid(visible=True, color='grey', linestyle='-.', linewidth=0.5, alpha=0.6)\n\n\nplt.subplot(1, 2, 1)\nsns.histplot(df.phrase_len, kde=True, binwidth = 2, color=color_scheme[0])\nplt.title('Character occurences in each phrase in training set')\nplt.xlabel('Phrase length')\nplt.ylabel('Sample Count')\n\nplt.subplot(1, 2, 2)\nsns.histplot(metadata.phrase_len, kde=True, binwidth = 2, color=color_scheme[0])\nplt.title('    and Supplementary metadata')\nplt.xlabel('Unique characters')\nplt.ylabel('Sample Count')\nplt.grid(axis='y')\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:24.916883Z","iopub.execute_input":"2023-06-20T17:37:24.917236Z","iopub.status.idle":"2023-06-20T17:37:26.358311Z","shell.execute_reply.started":"2023-06-20T17:37:24.917206Z","shell.execute_reply":"2023-06-20T17:37:26.357387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2.5\"></a>\n\n<h3 style=\"font-family: candaralight; font-size: 20px; font-style: normal; font-weight: normal; text-decoration: none; text-transform: none; letter-spacing: 2px; color: #C15D06; background-color: #ffffff;\">2.5 Inspect the <b>`PARQUET`</b> Column</h3>\n\n---\n\n\n\n","metadata":{}},{"cell_type":"code","source":"sample = pd.read_parquet(\"/kaggle/input/asl-fingerspelling/train_landmarks/1019715464.parquet\")\n","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:32.089363Z","iopub.execute_input":"2023-06-20T17:37:32.089721Z","iopub.status.idle":"2023-06-20T17:37:47.460464Z","shell.execute_reply.started":"2023-06-20T17:37:32.089692Z","shell.execute_reply":"2023-06-20T17:37:47.459506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"Sample shape = {sample.shape}\")\nsample.sample(10)","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:47.46283Z","iopub.execute_input":"2023-06-20T17:37:47.463161Z","iopub.status.idle":"2023-06-20T17:37:47.501596Z","shell.execute_reply.started":"2023-06-20T17:37:47.46313Z","shell.execute_reply":"2023-06-20T17:37:47.500615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample.describe()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:47.502998Z","iopub.execute_input":"2023-06-20T17:37:47.50402Z","iopub.status.idle":"2023-06-20T17:37:58.518351Z","shell.execute_reply.started":"2023-06-20T17:37:47.503988Z","shell.execute_reply":"2023-06-20T17:37:58.517376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"MULTI_HAND_LANDMARKS Collection of detected/tracked hands, where each hand is represented as a list of 21 hand landmarks and each landmark is composed of x, y and z. x and y are normalized to [0.0, 1.0] by the image width and height respectively. z represents the landmark depth with the depth at the wrist being the origin, and the smaller the value the closer the landmark is to the camera. The magnitude of z uses roughly the same scale as x.\n\nInaddition, there are few negative coordinates above. negative values are not expected for x and y perhaps. It is also important to note that for a good chunk of the video either or both hands will not be visible or in other words, will not have any landmark data. \n\nLets extract the columns and check the nans in either hands.\n","metadata":{}},{"cell_type":"code","source":"def get_cols(df, words_pos, words_neg=[], ret_names=True):\n    cols = []\n    names = []\n    for col in df.columns:\n        # Check if column name contains all words\n        if all([w in col for w in words_pos]) and all([w not in col for w in words_neg]):\n            cols.append(df[col])  # Append the entire column to the list\n            names.append(col)\n\n    if ret_names:\n        return cols, names\n    else:\n        return cols\n","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:58.521804Z","iopub.execute_input":"2023-06-20T17:37:58.524298Z","iopub.status.idle":"2023-06-20T17:37:58.531571Z","shell.execute_reply.started":"2023-06-20T17:37:58.524263Z","shell.execute_reply":"2023-06-20T17:37:58.530445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Landmark Indices for Left/Right hand without z axis in raw data\nLH_Index, LEFT_HAND_NAME = get_cols(sample, ['left_hand'], ['z'])\nRH_Index,RIGHT_HAND_NAME = get_cols(sample, ['right_hand'], ['z'])\n#RIGHT_HAND_NAMES0.insert(0, \"frame\")\nLEFT_HAND_NAME.insert(0, \"frame\")\nCOLUMNS = np.concatenate((LEFT_HAND_NAME, RIGHT_HAND_NAME))\n\nN_COLS0 = len(COLUMNS)\n\nprint(f'Total number of columns: {N_COLS0}')","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:58.532943Z","iopub.execute_input":"2023-06-20T17:37:58.533987Z","iopub.status.idle":"2023-06-20T17:37:58.554084Z","shell.execute_reply.started":"2023-06-20T17:37:58.533945Z","shell.execute_reply":"2023-06-20T17:37:58.551929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"RIGHT_HAND = sample.loc[:, RIGHT_HAND_NAME] \nLEFT_HAND = sample.loc[:, LEFT_HAND_NAME] \nRIGHT_HAND","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:58.556688Z","iopub.execute_input":"2023-06-20T17:37:58.557088Z","iopub.status.idle":"2023-06-20T17:37:58.638783Z","shell.execute_reply.started":"2023-06-20T17:37:58.557055Z","shell.execute_reply":"2023-06-20T17:37:58.63781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"Percentage of nulls in Left Hand data = {100*np.mean(LEFT_HAND['x_left_hand_0'].isnull()):.02f} %\")\nprint(f\"Percentage of nulls in Right Hand data = {100*np.mean(RIGHT_HAND['x_right_hand_0'].isnull()):.02f} %\")","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:37:58.642328Z","iopub.execute_input":"2023-06-20T17:37:58.645107Z","iopub.status.idle":"2023-06-20T17:37:58.655064Z","shell.execute_reply.started":"2023-06-20T17:37:58.645069Z","shell.execute_reply":"2023-06-20T17:37:58.652088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3\"></a>\n\n# <p style=\"font-family: candaralight; font-size: 28px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; color: #C15D06; background-color: #ffffff;\">3. PREPROCESSING DATA</p>\n\nIn this section the datais processed to extract relevant features, and prepared the dataset for further analysis and model training. The processed data can be used for building machine learning models, such as neural networks, to predict the ASL finger spelling from the hand landmarks.","metadata":{}},{"cell_type":"code","source":"sample = pd.read_parquet('/kaggle/input/asl-fingerspelling/train_landmarks/1021040628.parquet')\nLANDMARK_FILES_DIR = \"/kaggle/input/asl-fingerspelling/train_landmarks\"\nTRAIN_FILE = \"/kaggle/input/asl-fingerspelling/train.csv\"\nlabel_map = json.load(open(\"/kaggle/input/asl-fingerspelling/character_to_prediction_index.json\", \"r\"))","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:40:07.558411Z","iopub.execute_input":"2023-06-20T17:40:07.559244Z","iopub.status.idle":"2023-06-20T17:40:12.2954Z","shell.execute_reply.started":"2023-06-20T17:40:07.559211Z","shell.execute_reply":"2023-06-20T17:40:12.294482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(label_map)","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:40:14.825384Z","iopub.execute_input":"2023-06-20T17:40:14.825982Z","iopub.status.idle":"2023-06-20T17:40:14.832456Z","shell.execute_reply.started":"2023-06-20T17:40:14.825943Z","shell.execute_reply":"2023-06-20T17:40:14.83158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So there are 59 characters in total in the vocabulary. :\n\n    alphabets a-z : total 26 characters\n    digits : 0-9 , 10 in total\n    The rest are special characters. Let's look at them","metadata":{}},{"cell_type":"code","source":"# Memory saving function credit to https://www.kaggle.com/gemartin/load-data-reduce-memory-usage\ndef reduce_mem_usage(df):\n    \"\"\" iterate through all the columns of a dataframe and modify the data type\n        to reduce memory usage.        \n    \"\"\"\n    #start_mem = df.memory_usage().sum() / 1024**2\n    #print('Memory usage of dataframe is {:.2f} MB'.format(start_mem))\n\n    for col in df.columns:\n        col_type = df[col].dtype\n\n        if col_type != object:\n            c_min = df[col].min()\n            c_max = df[col].max()\n            if str(col_type)[:3] == 'int':\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)  \n            else:\n                if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:40:17.260511Z","iopub.execute_input":"2023-06-20T17:40:17.260926Z","iopub.status.idle":"2023-06-20T17:40:17.273145Z","shell.execute_reply.started":"2023-06-20T17:40:17.260894Z","shell.execute_reply":"2023-06-20T17:40:17.272146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import multiprocessing as mp\nimport pandas as pd\nimport torch\nimport numpy as np\n\n# Function to process a single parquet file\ndef process_parquet(row):\n    path = os.path.join(\"/kaggle/input/asl-fingerspelling\", row[1].path)\n    data_columns = COLUMNS\n    landmark_df = pd.read_parquet(path, columns=data_columns)\n\n    # Group the landmarks by sequence_id\n    grouped_landmarks = landmark_df.groupby('sequence_id')\n\n    # Initialize empty lists to store features and labels\n    features = []\n    labels = []\n\n    # Iterate over each sequence_id\n    for sequence_id, group in grouped_landmarks:\n        # Get the label for the sequence\n        phrase = df.loc[df['sequence_id'] == sequence_id, 'phrase'].iloc[0]\n\n        # Map each letter in the phrase using label_map\n        mapped_phrase = [letter for letter in phrase]\n        \n        # Create a new Series with sequence and mapped_phrase\n        result_series = pd.DataFrame({'sequence_id': sequence_id, 'mapped_phrase': mapped_phrase}) \n        result_series['label'] = result_series['mapped_phrase'].map(label_map).astype(np.int8)  \n\n        # Convert the label Series to a list\n        #label_list = result_series['label'].tolist()\n\n        # Initialize an empty feature vector for the sequence\n        sequence_features = []\n        \n        # Iterate over each landmark index\n        for landmark_index in range(20):\n            # Generate feature names for x, y, z coordinates\n            x_feature = f'x_right_hand_{landmark_index}'\n            y_feature = f'y_right_hand_{landmark_index}'\n\n            # Get the x, y, z coordinates for the landmark\n            x = group[x_feature].values.astype(np.float16)\n            y = group[y_feature].values.astype(np.float16)\n\n            # Replace NaN values with 0\n            x[np.isnan(x)] = 0.0\n            y[np.isnan(y)] = 0.0\n                        \n            # Perform feature transformations or calculations\n            x = torch.tensor(x).contiguous().view(-1, x.shape[0])\n            y = torch.tensor(y).contiguous().view(-1, y.shape[0])\n            \n            x_mean = torch.mean(x, 0) \n            y_mean = torch.mean(y, 0) \n\n            x_std = torch.std(x, 1) \n            y_std = torch.std(y, 1) \n\n            # Add the calculated features to the sequence feature vector\n            sequence_features = torch.cat([x_mean, y_mean,x_std,y_std], axis=0)\n            #sequence_features = torch.where(torch.isnan(sequence_features), torch.tensor(0.0, dtype=torch.float32), sequence_features)\n\n            diff = 3258 - sequence_features.shape[0]\n            if diff > 0:\n                padding = torch.zeros(diff)\n                sequence_features = torch.cat((sequence_features, padding))\n            \n            features = sequence_features[:3258].cpu().numpy()\n            \n            \n            return features, result_series['label']\n\nif __name__ == '__main__':\n    df = pd.read_csv(TRAIN_FILE)\n    df = reduce_mem_usage(df) \n    df2 = df.head(30)\n\n    max_label_length = 30\n    all_features = np.zeros((df.shape[0], 3258))\n    labels = np.zeros((df.shape[0], max_label_length))\n    #all_features = []\n    #labels = []\n\n    # Process parquet files in parallel\n    with mp.Pool() as pool:\n        results = pool.imap(process_parquet, df.iterrows(), chunksize=250)\n        for i, (x, y) in tqdm(enumerate(results), total=df.shape[0]):\n            all_features[i, :] = x\n            labels[i, :len(y)] = y.values.reshape(1, -1)\n\n    # Print the shapes of the tensors\n    np.save(\"feature_data.npy\", all_features)\n    np.save(\"feature_labels.npy\", labels)\n\n    print(\"Features tensor shape:\", all_features.shape)\n    print(\"Labels tensor shape:\", labels.shape)\n","metadata":{"execution":{"iopub.status.busy":"2023-06-20T17:40:28.96055Z","iopub.execute_input":"2023-06-20T17:40:28.960937Z","iopub.status.idle":"2023-06-20T21:07:14.096742Z","shell.execute_reply.started":"2023-06-20T17:40:28.960905Z","shell.execute_reply":"2023-06-20T21:07:14.09567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The process_parquet function processes a single parquet file. It reads the file, groups the landmarks by sequence_id, retrieves the corresponding label for each sequence, maps the label using label_map, performs feature transformations on the landmarks, and constructs a feature vector for the sequence. The function returns the features and labels for each sequence.\n","metadata":{}},{"cell_type":"markdown","source":"<a id=\"4\"></a>\n\n# <p style=\"font-family: candaralight; font-size: 28px; font-style: normal; font-weight: bold; text-decoration: none; text-transform: none; letter-spacing: 3px; color: #C15D06; background-color: #ffffff;\">4. GO FURTHER</p>\n\n[ASLFR training & Evaluation Levenshtein dist.](https://www.kaggle.com/code/hebasaleh00/aslfr-training-and-evaluation-with-levenshtein-dis)\n\nThis notebook is a continuation for the Google ASL Finger Recognition competition. In the this notebook, I processed the data and extracted features for the task. In the next notebook, I will train and evaluate a model using the extracted features. Additionally, I will use the Levenshtein distance as a metric to measure the similarity between predicted labels and ground truth labels.\n\n **<span style=\"color:darkorange;\"> If you liked this Notebook, please do forget to upvote, and GooD LucK.</span>**\n\n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}