{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Goal of the Competition\n\n\nThe goal of this competition is to detect and translate American Sign Language (ASL) fingerspelling into text.\n\nFingerspelling uses hand shapes that represent individual letters to convey words. While fingerspelling is only a part of ASL, it is often used for communicating names, addresses, phone numbers, and other information commonly entered on a mobile phone. Many Deaf smartphone users can fingerspell words faster than they can type on mobile keyboards. In fact, ASL fingerspelling can be substantially faster than typing on a smartphone’s virtual keyboard (57 words/minute average versus 36 words/minute US average). But sign language recognition AI for text entry lags far behind voice-to-text or even gesture-based typing, as robust datasets didn't previously exist.\n\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# Importing Libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport plotly.graph_objects as go\nimport seaborn as sns\n\n# import os\n# for dirname, _, filenames in os.walk('/kaggle/input'):\n#     for filename in filenames:\n#         print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2023-06-20T12:26:04.387009Z","iopub.execute_input":"2023-06-20T12:26:04.387387Z","iopub.status.idle":"2023-06-20T12:26:04.406523Z","shell.execute_reply.started":"2023-06-20T12:26:04.387358Z","shell.execute_reply":"2023-06-20T12:26:04.405043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"BASE_DIR = \"/kaggle/input/asl-fingerspelling/\"","metadata":{"execution":{"iopub.status.busy":"2023-06-20T11:18:56.706568Z","iopub.execute_input":"2023-06-20T11:18:56.707599Z","iopub.status.idle":"2023-06-20T11:18:56.712601Z","shell.execute_reply.started":"2023-06-20T11:18:56.707535Z","shell.execute_reply":"2023-06-20T11:18:56.711322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA","metadata":{}},{"cell_type":"markdown","source":"## Importing Data\n\nThe landmarks were extracted from raw videos with the MediaPipe holistic model. Not all of the frames necessarily had visible hands or hands that could be detected by the model.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/asl-fingerspelling/train.csv')\nprint(\"Size of training data: \",train_df.shape)\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T11:18:56.714016Z","iopub.execute_input":"2023-06-20T11:18:56.714689Z","iopub.status.idle":"2023-06-20T11:18:56.895834Z","shell.execute_reply.started":"2023-06-20T11:18:56.71464Z","shell.execute_reply":"2023-06-20T11:18:56.894769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Target Variable","metadata":{}},{"cell_type":"code","source":"print(\"Total Unique Phrases: \",train_df['phrase'].nunique())\n\nimport matplotlib.pyplot as plt\nfig, ax = plt.subplots(figsize=(8, 8))\ntrain_df[\"phrase\"].value_counts().head(50).sort_values(ascending=True).plot(\n    kind=\"barh\", ax=ax, title=\"Top 100 Signs in Training Dataset\"\n)\nax.set_xlabel(\"Number of Training Examples\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T11:18:56.898523Z","iopub.execute_input":"2023-06-20T11:18:56.898813Z","iopub.status.idle":"2023-06-20T11:18:57.660686Z","shell.execute_reply.started":"2023-06-20T11:18:56.898792Z","shell.execute_reply":"2023-06-20T11:18:57.659262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"(Min count, Max count) of phrase:\")\ntrain_df[\"phrase\"].value_counts().min(), train_df[\"phrase\"].value_counts().max()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T11:18:57.662106Z","iopub.execute_input":"2023-06-20T11:18:57.662625Z","iopub.status.idle":"2023-06-20T11:18:57.727524Z","shell.execute_reply.started":"2023-06-20T11:18:57.662598Z","shell.execute_reply":"2023-06-20T11:18:57.726491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,8))\nplt.hist(train_df['phrase'].value_counts(), bins=np.arange(0,train_df[\"phrase\"].value_counts().max()+1,1))\nplt.title(\"Distribution of value counts of unique phrases\")\nplt.ylabel(\"Number of phrases\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T11:18:57.729078Z","iopub.execute_input":"2023-06-20T11:18:57.729387Z","iopub.status.idle":"2023-06-20T11:18:58.040097Z","shell.execute_reply.started":"2023-06-20T11:18:57.729359Z","shell.execute_reply":"2023-06-20T11:18:58.039404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['phrase'].value_counts().value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T11:18:58.041214Z","iopub.execute_input":"2023-06-20T11:18:58.041678Z","iopub.status.idle":"2023-06-20T11:18:58.096654Z","shell.execute_reply.started":"2023-06-20T11:18:58.041653Z","shell.execute_reply":"2023-06-20T11:18:58.095902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training Data - Landmark\n\n* Each Parquet file is in the path:\n    *  train_landmark_files/[train/supplemental].parquet\n    * The parquet's associated phrase can be found in train.csv","metadata":{}},{"cell_type":"code","source":"id = 0\nprint(\"Phrase: \",train_df.loc[id,'phrase'])\nprint(\"Landmark data in parquet file:\\n\")\nparq_example_df = pd.read_parquet(BASE_DIR+train_df.loc[id,'path'])\nparq_example_df","metadata":{"execution":{"iopub.status.busy":"2023-06-20T11:23:10.188129Z","iopub.execute_input":"2023-06-20T11:23:10.188488Z","iopub.status.idle":"2023-06-20T11:23:11.64799Z","shell.execute_reply.started":"2023-06-20T11:23:10.188459Z","shell.execute_reply":"2023-06-20T11:23:11.647053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* sequence_id - A unique identifier for the landmark sequence. landmark files contain approximately 1,000 sequences. The sequence ID is used as the dataframe index.\n* frame - The frame number within a landmark sequence.\n* [x/y/z]\\_[type]\\_[landmark_index] - There are now 1,629 spatial coordinate columns for the x, y and z coordinates for each of the 543 landmarks. The type of landmark is one of ['face', 'left_hand', 'pose', 'right_hand']. Details of the hand landmark locations can be found here. The spatial coordinates have already been normalized by MediaPipe. \n\nNote that the MediaPipe model is not fully trained to predict depth so you may wish to ignore the z values. The landmarks have been converted to float32.","metadata":{}},{"cell_type":"code","source":"# Number of unique sequence Id\nprint(\"Number of unique seq id: \",parq_example_df.index.nunique())\nprint(\"Number of unique frame: \",parq_example_df.frame.nunique())","metadata":{"execution":{"iopub.status.busy":"2023-06-20T11:23:11.649681Z","iopub.execute_input":"2023-06-20T11:23:11.650034Z","iopub.status.idle":"2023-06-20T11:23:11.660385Z","shell.execute_reply.started":"2023-06-20T11:23:11.650007Z","shell.execute_reply":"2023-06-20T11:23:11.659204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"parq_example_df[['frame']].groupby('sequence_id').count()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T11:19:11.915401Z","iopub.execute_input":"2023-06-20T11:19:11.915676Z","iopub.status.idle":"2023-06-20T11:19:11.931818Z","shell.execute_reply.started":"2023-06-20T11:19:11.915646Z","shell.execute_reply":"2023-06-20T11:19:11.930149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"parq_example_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T11:19:11.933817Z","iopub.execute_input":"2023-06-20T11:19:11.934176Z","iopub.status.idle":"2023-06-20T11:19:12.152895Z","shell.execute_reply.started":"2023-06-20T11:19:11.934148Z","shell.execute_reply":"2023-06-20T11:19:12.152166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of rows with no co-ordinates at all i.e all NAN \n(parq_example_df.isna().sum(axis=1)>0).value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T11:19:12.154315Z","iopub.execute_input":"2023-06-20T11:19:12.154769Z","iopub.status.idle":"2023-06-20T11:19:12.515394Z","shell.execute_reply.started":"2023-06-20T11:19:12.154733Z","shell.execute_reply":"2023-06-20T11:19:12.51459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"parq_example_df","metadata":{"execution":{"iopub.status.busy":"2023-06-20T11:20:18.963295Z","iopub.execute_input":"2023-06-20T11:20:18.963634Z","iopub.status.idle":"2023-06-20T11:20:18.998099Z","shell.execute_reply.started":"2023-06-20T11:20:18.963612Z","shell.execute_reply":"2023-06-20T11:20:18.997113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualization of landmark data","metadata":{}},{"cell_type":"code","source":"def wide_to_long_format(sequence):\n    # extracting types and landmark_indexes from col names\n    types = []\n    landmark_indexes = []\n    for column in list(sequence.columns)[1:544]:\n        parts = column.split(\"_\")\n        # 'face', 'left_hand', 'pose', 'right_hand'\n        if len(parts) == 4:\n            types.append(parts[1] + \"_\" + parts[2])\n        else:\n            types.append(parts[1])\n\n        landmark_indexes.append(int(parts[-1]))\n\n\n    data = {\n        \"frame\": [],\n        \"type\": [],\n        \"landmark_index\": [],\n        \"x\": [],\n        \"y\": [],\n        \"z\": []\n    }\n\n    for index, row in sequence.iterrows():\n        data[\"frame\"] += [int(row.frame)]*543\n        data[\"type\"] += types\n        data[\"landmark_index\"] += landmark_indexes\n\n        for _type, landmark_index in zip(types, landmark_indexes):\n            data[\"x\"].append(row[f\"x_{_type}_{landmark_index}\"])\n            data[\"y\"].append(row[f\"y_{_type}_{landmark_index}\"])\n            data[\"z\"].append(row[f\"z_{_type}_{landmark_index}\"])\n\n    return pd.DataFrame.from_dict(data)\n\n# assign desired colors to landmarks\ndef assign_color(row):\n    if row == 'face':\n        return 'red'\n    elif 'hand' in row:\n        return 'dodgerblue'\n    else:\n        return 'green'\n\n# specifies the plotting order \ndef assign_order(row):\n    if row.type == 'face':\n        return row.landmark_index + 101\n    elif row.type == 'pose':\n        return row.landmark_index + 30\n    elif row.type == 'left_hand':\n        return row.landmark_index + 80\n    else:\n        return row.landmark_index\n\n\ndef visualize2d_landmarks(parquet_df, title=\"\"):\n    connections = [  \n        [0, 1, 2, 3, 4,],\n        [0, 5, 6, 7, 8],\n        [0, 9, 10, 11, 12],\n        [0, 13, 14, 15, 16],\n        [0, 17, 18, 19, 20],\n\n        \n        [38, 36, 35, 34, 30, 31, 32, 33, 37],\n        [40, 39],\n        [52, 46, 50, 48, 46, 44, 42, 41, 43, 45, 47, 49, 45, 51],\n        [42, 54, 56, 58, 60, 62, 58],\n        [41, 53, 55, 57, 59, 61, 57],\n        [54, 53],\n\n        \n        [80, 81, 82, 83, 84, ],\n        [80, 85, 86, 87, 88],\n        [80, 89, 90, 91, 92],\n        [80, 93, 94, 95, 96],\n        [80, 97, 98, 99, 100], ]\n    \n    parquet_df = wide_to_long_format(parquet_df)\n    frames = sorted(set(parquet_df.frame))\n    first_frame = min(frames)\n    parquet_df['color'] = parquet_df.type.apply(lambda row: assign_color(row))\n    parquet_df['plot_order'] = parquet_df.apply(lambda row: assign_order(row), axis=1)\n    first_frame_df = parquet_df[parquet_df.frame == first_frame].copy()\n    first_frame_df = first_frame_df.sort_values([\"plot_order\"]).set_index('plot_order')\n\n\n    frames_l = []\n    for frame in frames:\n        filtered_df = parquet_df[parquet_df.frame == frame].copy()\n        filtered_df = filtered_df.sort_values([\"plot_order\"]).set_index(\"plot_order\")\n        \n        #scatter plot\n        traces = [go.Scatter(\n            x=filtered_df['x'],\n            y=filtered_df['y'],\n            mode='markers',\n            marker=dict(\n                color=filtered_df.color,\n                size=9))]\n\n        # drawing lines \n        for i, seg in enumerate(connections):\n            trace = go.Scatter(\n                    x=filtered_df.loc[seg]['x'],\n                    y=filtered_df.loc[seg]['y'],\n                    mode='lines',\n            )\n            traces.append(trace)\n        frame_data = go.Frame(data=traces, traces = [i for i in range(17)])\n        frames_l.append(frame_data)\n\n    # First frame -  scatter plot\n    traces = [go.Scatter(\n        x=first_frame_df['x'],\n        y=first_frame_df['y'],\n        mode='markers',\n        marker=dict(\n            color=first_frame_df.color,\n            size=9\n        )\n    )]\n    # Drawing lines in First Frame\n    for i, seg in enumerate(connections):\n        trace = go.Scatter(\n            x=first_frame_df.loc[seg]['x'],\n            y=first_frame_df.loc[seg]['y'],\n            mode='lines',\n            line=dict(\n                color='black',\n                width=2\n            )\n        )\n        traces.append(trace)\n    \n    fig = go.Figure(\n        data=traces,\n        frames=frames_l\n    )\n\n    fig.update_layout(\n        width=500,\n        height=800,\n        scene={\n            'aspectmode': 'data',\n        },\n        title_text=title,\n        title_x=0.5,\n        updatemenus=[\n            {\n                \"buttons\": [\n                    {\n                        \"args\": [None, {\"frame\": {\"duration\": 100,\n                                                  \"redraw\": True},\n                                        \"fromcurrent\": True,\n                                        \"transition\": {\"duration\": 0}}],\n                        \"label\": \"&#9654;\",\n                        \"method\": \"animate\",\n                    },\n\n                ],\n                \"direction\": \"left\",\n                \"pad\": {\"r\": 100, \"t\": 100},\n                \"font\": {\"size\":30},\n                \"type\": \"buttons\",\n                \"x\": 0.1,\n                \"y\": 0,\n            }\n        ],\n    )\n    # set up camera position \n    camera = dict(\n        up=dict(x=0, y=-1, z=0),\n        eye=dict(x=0, y=0, z=0)\n    )\n    fig.update_layout(scene_camera=camera, showlegend=False)\n    fig.update_layout(xaxis = dict(visible=False),\n            yaxis = dict(visible=False),\n    )\n    fig.update_yaxes(autorange=\"reversed\")\n\n    fig.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T14:10:13.545746Z","iopub.execute_input":"2023-06-20T14:10:13.546133Z","iopub.status.idle":"2023-06-20T14:10:13.574648Z","shell.execute_reply.started":"2023-06-20T14:10:13.546109Z","shell.execute_reply":"2023-06-20T14:10:13.572438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"id = 10\nprint(\"Phrase: \",train_df.loc[id,'phrase'])\nprint(\"Landmark data in parquet file:\\n\")\nphrase = train_df.loc[id,'phrase']\nparq_example_df = pd.read_parquet(BASE_DIR+train_df.loc[id,'path'])\nparq_example_df","metadata":{"execution":{"iopub.status.busy":"2023-06-20T13:23:41.930598Z","iopub.execute_input":"2023-06-20T13:23:41.931024Z","iopub.status.idle":"2023-06-20T13:23:45.010899Z","shell.execute_reply.started":"2023-06-20T13:23:41.930991Z","shell.execute_reply":"2023-06-20T13:23:45.009893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"parq_example_df.index.nunique()","metadata":{"execution":{"iopub.status.busy":"2023-06-20T11:26:54.095834Z","iopub.execute_input":"2023-06-20T11:26:54.096166Z","iopub.status.idle":"2023-06-20T11:26:54.103605Z","shell.execute_reply.started":"2023-06-20T11:26:54.09614Z","shell.execute_reply":"2023-06-20T11:26:54.102762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sequence_id = 1816796431 #parq_example_df.index[0]\nseq_df = parq_example_df[parq_example_df.index==sequence_id]\nseq_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-06-20T11:28:56.664642Z","iopub.execute_input":"2023-06-20T11:28:56.665002Z","iopub.status.idle":"2023-06-20T11:28:56.674625Z","shell.execute_reply.started":"2023-06-20T11:28:56.664975Z","shell.execute_reply":"2023-06-20T11:28:56.673623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of z co-ordinates\ncount = 0\nfor col in list(seq_df.columns):\n    if 'z' in col:\n        count+=1\nprint(count)","metadata":{"execution":{"iopub.status.busy":"2023-06-20T12:12:49.757024Z","iopub.execute_input":"2023-06-20T12:12:49.760646Z","iopub.status.idle":"2023-06-20T12:12:49.776327Z","shell.execute_reply.started":"2023-06-20T12:12:49.760557Z","shell.execute_reply":"2023-06-20T12:12:49.774147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"visualize2d_landmarks(seq_df, phrase)","metadata":{"execution":{"iopub.status.busy":"2023-06-20T14:10:16.382138Z","iopub.execute_input":"2023-06-20T14:10:16.382533Z","iopub.status.idle":"2023-06-20T14:10:29.072446Z","shell.execute_reply.started":"2023-06-20T14:10:16.382505Z","shell.execute_reply":"2023-06-20T14:10:29.070795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"execution":{"iopub.status.busy":"2023-06-20T14:06:16.592609Z","iopub.execute_input":"2023-06-20T14:06:16.593024Z","iopub.status.idle":"2023-06-20T14:06:16.609814Z","shell.execute_reply.started":"2023-06-20T14:06:16.592998Z","shell.execute_reply":"2023-06-20T14:06:16.609013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}