{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<!-- Codes by HTMLcodes.ws -->\n<h1 style = \"background-color:MediumSpringGreen;font-family:newtimeroman;font-size:250%;text-align:center;border-radius:15px 50px;\">\"ASL Fingerspelling Accuracy with Levenshtein Distance\"</h1>\n","metadata":{}},{"cell_type":"markdown","source":"# Introduction\n\nWelcome to the ASL Fingerspelling Translation competition, where AI empowers the Deaf and Hard of Hearing community to enhance communication. By training a specialized model, participants have the opportunity to revolutionize sign language recognition technology.\n\nWhile voice-enabled assistants and AI solutions have revolutionized modern devices, they often overlook the 70+ million Deaf individuals worldwide and the 1.5+ billion people affected by hearing loss. Fingerspelling, a key aspect of ASL, uses hand shapes to represent letters and is frequently used for text input on mobile devices. Deaf smartphone users can fingerspell words faster than they can type on virtual keyboards. However, sign language recognition AI for text entry has been limited due to the lack of comprehensive datasets.\n\nThis competition aligns with Google's mission of universal accessibility and AI principles by exploring scalable solutions for sign language recognition. In collaboration with the Deaf Professional Arts Network, the competition aims to address individual user needs and expand to other sign languages.\n\nParticipating in this competition empowers Deaf and Hard of Hearing users to use fingerspelling instead of traditional keyboards. Beyond convenient text entry, there is potential for an app that translates fingerspelling into spoken words, facilitating smoother communication between the Deaf and non-signing individuals.\n\nJoin this competition to contribute to the advancement of sign language technology, bridging the gap between sign language and mainstream AI applications. Together, we can build a more inclusive and accessible future.","metadata":{}},{"cell_type":"markdown","source":"## What is ASL Fingerspelling Translation?\n\nASL Fingerspelling Translation refers to the process of translating American Sign Language (ASL) fingerspelling into written or spoken language using artificial intelligence (AI) technology. Fingerspelling is a fundamental aspect of ASL where hand shapes are used to represent letters of the alphabet. It is commonly used for spelling out words, names, or other specific terms in sign language.\n\nThe ASL Fingerspelling Translation competition harnesses the power of AI to improve the recognition and interpretation of fingerspelling gestures. Participants in the competition train AI models on specialized datasets to enhance the accuracy and efficiency of translating fingerspelling into written or spoken words. The goal is to develop scalable AI solutions that can benefit the Deaf and Hard of Hearing community by improving communication and accessibility through sign language recognition technology.","metadata":{}},{"cell_type":"markdown","source":"# **Install Dependencies**","metadata":{}},{"cell_type":"code","source":"%%capture\n!pip install python-Levenshtein==0.12.0","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:04:07.778105Z","iopub.execute_input":"2023-05-28T10:04:07.77849Z","iopub.status.idle":"2023-05-28T10:04:18.036366Z","shell.execute_reply.started":"2023-05-28T10:04:07.77846Z","shell.execute_reply":"2023-05-28T10:04:18.034189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ref: [python-Levenshtein](https://pypi.org/project/python-Levenshtein/0.12.0/)","metadata":{}},{"cell_type":"markdown","source":"# **Import Modules**","metadata":{}},{"cell_type":"code","source":"import cv2\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\n%matplotlib inline\nimport json\nimport tensorflow as tf\nfrom tensorflow.keras import layers, optimizers, constraints, regularizers\nimport plotly.graph_objects as go\nimport plotly.io as pio\nimport os\nfrom Levenshtein import distance\nfrom datetime import datetime\n\nplt.rcParams['figure.figsize'] = (12,6)\nplt.style.use('fivethirtyeight')\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:04:35.50783Z","iopub.execute_input":"2023-05-28T10:04:35.508548Z","iopub.status.idle":"2023-05-28T10:04:35.517767Z","shell.execute_reply.started":"2023-05-28T10:04:35.508523Z","shell.execute_reply":"2023-05-28T10:04:35.516715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Load the Dataset**","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/asl-fingerspelling/train.csv')\ntrain_df.head(4).style.set_properties(**{'background-color':'royalblue','color':'black','border-color':'#8b8c8c'})","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:04:42.122001Z","iopub.execute_input":"2023-05-28T10:04:42.122412Z","iopub.status.idle":"2023-05-28T10:04:42.29664Z","shell.execute_reply.started":"2023-05-28T10:04:42.122385Z","shell.execute_reply":"2023-05-28T10:04:42.29537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check the dimensions of the dataset\nprint(train_df.shape)\n\n# Check the data types of columns\nprint(train_df.info())","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:04:46.458412Z","iopub.execute_input":"2023-05-28T10:04:46.458901Z","iopub.status.idle":"2023-05-28T10:04:46.493369Z","shell.execute_reply.started":"2023-05-28T10:04:46.458869Z","shell.execute_reply":"2023-05-28T10:04:46.492221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the statistical of the dataset\nstyled_data = train_df.describe().style\\\n.background_gradient(cmap='coolwarm')\\\n.set_properties(**{'text-align':'center','border':'1px solid black'})\n\n# display styled data\ndisplay(styled_data)","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:04:50.43842Z","iopub.execute_input":"2023-05-28T10:04:50.439308Z","iopub.status.idle":"2023-05-28T10:04:50.49027Z","shell.execute_reply.started":"2023-05-28T10:04:50.439237Z","shell.execute_reply":"2023-05-28T10:04:50.487469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sequence_id = 1817362238\nfile_id = 5414471","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:05:04.850118Z","iopub.execute_input":"2023-05-28T10:05:04.850539Z","iopub.status.idle":"2023-05-28T10:05:04.855845Z","shell.execute_reply.started":"2023-05-28T10:05:04.850503Z","shell.execute_reply":"2023-05-28T10:05:04.854657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sign_path = f\"/kaggle/input/asl-fingerspelling/train_landmarks/{file_id}.parquet\"\nsign = pd.read_parquet(sign_path)\n","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:05:08.944916Z","iopub.execute_input":"2023-05-28T10:05:08.945301Z","iopub.status.idle":"2023-05-28T10:05:16.414658Z","shell.execute_reply.started":"2023-05-28T10:05:08.945271Z","shell.execute_reply":"2023-05-28T10:05:16.413896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(np.unique(sign.index))","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:05:35.776053Z","iopub.execute_input":"2023-05-28T10:05:35.77749Z","iopub.status.idle":"2023-05-28T10:05:35.788738Z","shell.execute_reply.started":"2023-05-28T10:05:35.777422Z","shell.execute_reply":"2023-05-28T10:05:35.787831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sequence = sign[sign.index == sequence_id]","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:05:40.652332Z","iopub.execute_input":"2023-05-28T10:05:40.652735Z","iopub.status.idle":"2023-05-28T10:05:40.663596Z","shell.execute_reply.started":"2023-05-28T10:05:40.652705Z","shell.execute_reply":"2023-05-28T10:05:40.662549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"suppl_df = pd.read_csv('/kaggle/input/asl-fingerspelling/supplemental_metadata.csv')\nsuppl_df.head(4).style.set_properties(**{'background-color':'lightgreen','color':'black','border-color':'#8b8c8c'})","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:05:46.10661Z","iopub.execute_input":"2023-05-28T10:05:46.106977Z","iopub.status.idle":"2023-05-28T10:05:46.225659Z","shell.execute_reply.started":"2023-05-28T10:05:46.106952Z","shell.execute_reply":"2023-05-28T10:05:46.224878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check the dimensions of the dataset\nprint(suppl_df.shape)\n\n# Check the data types of columns\nprint(suppl_df.info())","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:05:50.295006Z","iopub.execute_input":"2023-05-28T10:05:50.29538Z","iopub.status.idle":"2023-05-28T10:05:50.325321Z","shell.execute_reply.started":"2023-05-28T10:05:50.295352Z","shell.execute_reply":"2023-05-28T10:05:50.32359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the statistical of the dataset\nstyled_data = suppl_df.describe().style\\\n.background_gradient(cmap='coolwarm')\\\n.set_properties(**{'text-align':'center','border':'1px solid black'})\n\n# display styled data\ndisplay(styled_data)","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:05:54.525432Z","iopub.execute_input":"2023-05-28T10:05:54.52587Z","iopub.status.idle":"2023-05-28T10:05:54.557094Z","shell.execute_reply.started":"2023-05-28T10:05:54.525838Z","shell.execute_reply":"2023-05-28T10:05:54.555559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sequence_id = 1535585216\nfile_id = 33432165","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:05:58.525725Z","iopub.execute_input":"2023-05-28T10:05:58.526129Z","iopub.status.idle":"2023-05-28T10:05:58.531844Z","shell.execute_reply.started":"2023-05-28T10:05:58.5261Z","shell.execute_reply":"2023-05-28T10:05:58.530361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sign_path = f\"/kaggle/input/asl-fingerspelling/supplemental_landmarks/{file_id}.parquet\"\nsign = pd.read_parquet(sign_path)\n","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:06:01.866641Z","iopub.execute_input":"2023-05-28T10:06:01.868153Z","iopub.status.idle":"2023-05-28T10:06:08.856588Z","shell.execute_reply.started":"2023-05-28T10:06:01.868099Z","shell.execute_reply":"2023-05-28T10:06:08.855892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# The Evaluation metric - Levenshtein distance:","metadata":{}},{"cell_type":"markdown","source":"## What is Levenshtein distance?\n\nThe Levenshtein distance, or edit distance, is a metric that quantifies the dissimilarity between two strings. It calculates the minimum number of single-character operations (insertions, deletions, or substitutions) needed to transform one string into another.\n\nNamed after Vladimir Levenshtein, a Soviet mathematician, the concept was introduced in 1965. The Levenshtein distance finds applications in various fields like spell checking, DNA sequence analysis, natural language processing, and computational linguistics.\n\nTo compute the Levenshtein distance, an algorithm constructs a matrix where each cell represents the cost of transforming one substring to another. Starting from the top-left cell and moving towards the bottom-right, the algorithm compares characters and determines the minimum cost using insertions, deletions, or substitutions. The value in the bottom-right cell represents the Levenshtein distance between the two strings.\n\nThe Levenshtein distance serves as a similarity measure between strings and is used in tasks such as string matching, string clustering, and fuzzy string searching. It also forms the basis for other string distance metrics, like the Damerau-Levenshtein distance, which incorporates transpositions as an additional operation.\n\nIn the context of the provided code, the Levenshtein distance is utilized to compute a distance matrix between two sequences. This matrix is then used for model training and performance evaluation.","metadata":{}},{"cell_type":"markdown","source":"* **Expression Metric**\n\nLevenshtein distance, the expression metric = (N - D) / N represents a way to measure the similarity or dissimilarity between two strings based on their Levenshtein distance.\n\nLet's break down the components of the expression:\n\n* N: N represents the length of the longer string between the two compared strings. It is the maximum possible number of character positions that need to be considered.\n\n* D: D corresponds to the Levenshtein distance between the two strings. It is the actual number of single-character edits required to transform one string into the other.\n\nThe expression (N - D) represents the number of character positions that are unchanged or require no edit operations to transform one string into the other. By subtracting D from N, we obtain the number of common characters or positions between the two strings.\n\nDividing (N - D) by N normalizes this value by the maximum possible number of character positions, N. This normalization results in a similarity metric ranging between 0 and 1, where 0 represents no similarity and 1 indicates an exact match.\n\nTherefore, the expression metric = (N - D) / N provides a measure of similarity between two strings based on the Levenshtein distance. A value close to 1 suggests a high degree of similarity, while a value closer to 0 indicates a larger difference between the strings.\n\nRef:[Wikipedia - Levenshtein distance](https://en.wikipedia.org/wiki/Levenshtein_distance)","metadata":{}},{"cell_type":"markdown","source":"For example, the Levenshtein distance between \"kitten\" and \"sitting\" is 3, since the following 3 edits change one into the other, and there is no way to do it with fewer than 3 edits:\n\n* kitten → sitten (substitution of \"s\" for \"k\"),\n* sitten → sittin (substitution of \"i\" for \"e\"),\n* sittin → sitting (insertion of \"g\" at the end).\n\n![image](https://upload.wikimedia.org/wikipedia/commons/d/d1/Levenshtein_distance_animation.gif)","metadata":{}},{"cell_type":"code","source":"from Levenshtein import distance as lev","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:06:16.844019Z","iopub.execute_input":"2023-05-28T10:06:16.844402Z","iopub.status.idle":"2023-05-28T10:06:16.849337Z","shell.execute_reply.started":"2023-05-28T10:06:16.844375Z","shell.execute_reply":"2023-05-28T10:06:16.848325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def levenshtein(seq1, seq2):\n    size_x = len(seq1) + 1\n    size_y = len(seq2) + 1\n    matrix = np.zeros((size_x, size_y))\n\n    for x in range(size_x):\n        matrix[x, 0] = x\n\n    for y in range(size_y):\n        matrix[0, y] = y\n\n    for x in range(1, size_x):\n        for y in range(1, size_y):\n            if seq1[x-1] == seq2[y-1]:\n                matrix[x, y] = min(\n                    matrix[x-1, y] + 1,\n                    matrix[x-1, y-1],\n                    matrix[x, y-1] + 1\n                )\n            else:\n                matrix[x, y] = min(\n                    matrix[x-1, y] + 1,\n                    matrix[x-1, y-1] + 1,\n                    matrix[x, y-1] + 1\n                )\n\n    return matrix[size_x - 1, size_y - 1]\n","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:06:22.813631Z","iopub.execute_input":"2023-05-28T10:06:22.814682Z","iopub.status.idle":"2023-05-28T10:06:22.82269Z","shell.execute_reply.started":"2023-05-28T10:06:22.814643Z","shell.execute_reply":"2023-05-28T10:06:22.821407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import plotly.graph_objects as go\n\ndef plot_levenshtein_matrix(matrix):\n    fig = go.Figure(data=go.Heatmap(z=matrix, colorscale='Viridis'))\n    fig.update_layout(\n        title='Levenshtein Distance Matrix',\n        xaxis_title='Sequence 2',\n        yaxis_title='Sequence 1'\n    )\n    fig.show()\n\n# Example usage\nseq1 = '3 creekhouse'\nseq2 = 'scales/kuhaylah'\nmatrix = np.zeros((len(seq1) + 1, len(seq2) + 1))\n\ndistance = levenshtein(seq1, seq2)\nprint(\"Levenshtein Distance:\", distance)\n\nplot_levenshtein_matrix(matrix)\n","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:06:30.614045Z","iopub.execute_input":"2023-05-28T10:06:30.614746Z","iopub.status.idle":"2023-05-28T10:06:30.670308Z","shell.execute_reply.started":"2023-05-28T10:06:30.614713Z","shell.execute_reply":"2023-05-28T10:06:30.669244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Inference Model","metadata":{}},{"cell_type":"code","source":"basedir = \"/kaggle/working/\"\nNUM_CHARACTERS = 59\n\nSEL_FEATURES = ['x_right_hand_0', 'y_right_hand_0', 'z_right_hand_0',\n                # ... rest of the features ...\n                'x_left_hand_20', 'y_left_hand_20', 'z_left_hand_20'\n                ]\nNUM_FEATURES = len(SEL_FEATURES)\n\nd = {\"selected_columns\": SEL_FEATURES}\n\nwith open(f\"{basedir}/inference_args.json\", \"w\") as f:\n    json.dump(d, f)\n\n\ndef get_dummy_model():\n    inputs = tf.keras.Input(shape=(NUM_FEATURES), dtype=tf.float32, name=\"inputs\")\n    x = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n    x = tf.keras.layers.Dense(NUM_CHARACTERS)(x)\n    out = tf.keras.layers.Activation(\"linear\", name=\"outputs\")(x)\n    inference_model = tf.keras.Model(inputs=inputs, outputs=out)\n    inference_model.compile(loss=\"sparse_categorical_crossentropy\",\n                            metrics=\"accuracy\")\n    return inference_model\n\n\ndummy_model_test = get_dummy_model()\ndummy_model_test.summary()","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:06:48.578781Z","iopub.execute_input":"2023-05-28T10:06:48.579171Z","iopub.status.idle":"2023-05-28T10:06:48.99061Z","shell.execute_reply.started":"2023-05-28T10:06:48.579144Z","shell.execute_reply":"2023-05-28T10:06:48.989294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"converter = tf.lite.TFLiteConverter.from_keras_model(dummy_model_test)\n\ntflite_model = converter.convert()\nmodel_path = 'model.tflite'\n\nwith open(model_path, 'wb') as f:\n    f.write(tflite_model)\n\n!zip submission.zip  '/kaggle/working/model.tflite' '/kaggle/working/inference_args.json'","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:07:09.624411Z","iopub.execute_input":"2023-05-28T10:07:09.624897Z","iopub.status.idle":"2023-05-28T10:07:10.638679Z","shell.execute_reply.started":"2023-05-28T10:07:09.624862Z","shell.execute_reply":"2023-05-28T10:07:10.637622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"CHECKING = False\n\nif CHECKING:\n    !pip install tflite-runtime==2.9.1\n    import tflite_runtime.interpreter as tflite\n\n    def load_relevant_data_subset(pq_path):\n        return pd.read_parquet(pq_path, columns=SEL_FEATURES) #selected_columns)\n    \n    data_path = \"/kaggle/input/asl-fingerspelling/train_landmarks/1021040628.parquet\"\n    frames = load_relevant_data_subset(data_path).values\n    \n    interpreter = tflite.Interpreter(model_path)\n    found_signatures = list(interpreter.get_signature_list().keys())\n    prediction_fn = interpreter.get_signature_runner(\"serving_default\")\n    \n    with open (\"/kaggle/input/asl-fingerspelling/character_to_prediction_index.json\", \"r\") as f:\n        character_map = json.load(f)\n    rev_character_map = {j:i for i,j in character_map.items()}","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:07:22.055642Z","iopub.execute_input":"2023-05-28T10:07:22.056078Z","iopub.status.idle":"2023-05-28T10:07:22.066457Z","shell.execute_reply.started":"2023-05-28T10:07:22.056045Z","shell.execute_reply":"2023-05-28T10:07:22.065559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if CHECKING:\n    output = prediction_fn(inputs=frames)\n    prediction_str = \"\".join([rev_character_map.gets(s,\"\")for s in np.argmax(output['outputs'], axis=1)])\n    print(\"\\n\\n\",prediction_str[:100])","metadata":{"execution":{"iopub.status.busy":"2023-05-28T10:07:27.503958Z","iopub.execute_input":"2023-05-28T10:07:27.504378Z","iopub.status.idle":"2023-05-28T10:07:27.510457Z","shell.execute_reply.started":"2023-05-28T10:07:27.504348Z","shell.execute_reply":"2023-05-28T10:07:27.509206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![image](https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcTDAD-Z6e0Lad_MWmVJw-crpHqq-SFh9aBOdA&usqp=CAU)","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#5642C5;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n\n<p style=\"padding: 10px;\n              color:white;\">\nYour upvote is a great way to show your support and help others discover this valuable resource.","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\"> 📌 Note: If you forks my notebook, please don't forget to upvote it. </div>\n","metadata":{}}]}