{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# ASL- Fingerspelling ","metadata":{}},{"cell_type":"markdown","source":"## What is American Sign Language Fingerspelling Recognition ?\n\nAmerican Sign Language Fingerspelling Recognition is a technology that uses computer vision and machine learning algorithms to recognize and interpret the hand gestures used in American Sign Language (ASL) fingerspelling. It can be used to create tools and applications that help people with hearing impairments to communicate more effectively with others. The technology involves comparing the input image of the hand gesture to a pre-defined set of templates, extracting relevant features from the input image, or training a neural network on a large dataset of ASL fingerspelling images to learn the patterns and features that are most important for recognition. Despite some challenges, ASL Fingerspelling Recognition has the potential to greatly improve the lives of people with hearing impairments.","metadata":{}},{"cell_type":"markdown","source":"## Data Overview\n\n### Files\n#### [train/supplemental_metadata].csv\n\n\n* path - The path to the landmark file.\n* file_id - A unique identifier for the data file.\n* participant_id - A unique identifier for the data contributor.\n* sequence_id - A unique identifier for the landmark sequence. Each data file may contain many sequences.\n* phrase - The labels for the landmark sequence. The train and test datasets contain randomly generated addresses, phone numbers, and urls derived from components of real addresses/phone numbers/urls. Any overlap with real addresses, phone numbers, or urls is purely accidental. The supplemental dataset consists of fingerspelled sentences. Note that some of the urls include adult content. The intent of this competition is to support the Deaf and Hard of Hearing community in engaging with technology on an equal footing with other adults.\n\n### character_to_prediction_index.json\n\n#### [train/supplemental]_landmarks/ \nThe landmark data. The landmarks were extracted from raw videos with the MediaPipe holistic model. Not all of the frames necessarily had visible hands or hands that could be detected by the model.\nThe landmark files contain the same data as in the ASL Signs competition (minus the row ID column) but reshaped into a wide format. This allows you to take advantage of the Parquet format to entirely skip loading landmarks that you aren't using.\n\n* sequence_id - A unique identifier for the landmark sequence. Most landmark files contain 1,000 sequences. The sequence ID is used as the dataframe index.\n* frame - The frame number within a landmark sequence.\n* [x/y/z]_[type]_[landmark_index] - There are now 1,629 spatial coordinate columns for the x, y and z coordinates for each of the 543 landmarks. The type of landmark is one of ['face', 'left_hand', 'pose', 'right_hand']. Details of the hand landmark locations can be found here. The spatial coordinates have already been normalized by MediaPipe. Note that the MediaPipe model is not fully trained to predict depth so you may wish to ignore the z values. The landmarks have been converted to float32.","metadata":{}},{"cell_type":"code","source":"### import libraries\nimport pandas as pd,numpy as np,os\nimport json\nimport plotly.graph_objects as go\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nimport plotly.io as pio\nfrom pathlib import Path\npio.templates.default = \"simple_white\"\nprint(\"importing..\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-05-14T14:10:57.789472Z","iopub.execute_input":"2023-05-14T14:10:57.789884Z","iopub.status.idle":"2023-05-14T14:11:00.830554Z","shell.execute_reply.started":"2023-05-14T14:10:57.78985Z","shell.execute_reply":"2023-05-14T14:11:00.829453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def map_new_to_old_style(sequence):\n    types = []\n    landmark_indexes = []\n    for column in list(sequence.columns)[1:544]:\n        parts = column.split(\"_\")\n        if len(parts) == 4:\n            types.append(parts[1] + \"_\" + parts[2])\n        else:\n            types.append(parts[1])\n\n        landmark_indexes.append(int(parts[-1]))\n\n    data = {\n        \"frame\": [],\n        \"type\": [],\n        \"landmark_index\": [],\n        \"x\": [],\n        \"y\": [],\n        \"z\": []\n    }\n\n    for index, row in sequence.iterrows():\n        data[\"frame\"] += [int(row.frame)]*543\n        data[\"type\"] += types\n        data[\"landmark_index\"] += landmark_indexes\n\n        for _type, landmark_index in zip(types, landmark_indexes):\n            data[\"x\"].append(row[f\"x_{_type}_{landmark_index}\"])\n            data[\"y\"].append(row[f\"y_{_type}_{landmark_index}\"])\n            data[\"z\"].append(row[f\"z_{_type}_{landmark_index}\"])\n\n    return pd.DataFrame.from_dict(data)\n\n# assign desired colors to landmarks\ndef assign_color(row):\n    if row == 'face':\n        return 'red'\n    elif 'hand' in row:\n        return 'dodgerblue'\n    else:\n        return 'green'\n\n# specifies the plotting order\ndef assign_order(row):\n    if row.type == 'face':\n        return row.landmark_index + 101\n    elif row.type == 'pose':\n        return row.landmark_index + 30\n    elif row.type == 'left_hand':\n        return row.landmark_index + 80\n    else:\n        return row.landmark_index\n    \ndef visualise2d_landmarks(parquet_df, title=\"\"):\n    connections = [  \n        [0, 1, 2, 3, 4,],\n        [0, 5, 6, 7, 8],\n        [0, 9, 10, 11, 12],\n        [0, 13, 14, 15, 16],\n        [0, 17, 18, 19, 20],\n\n        \n        [38, 36, 35, 34, 30, 31, 32, 33, 37],\n        [40, 39],\n        [52, 46, 50, 48, 46, 44, 42, 41, 43, 45, 47, 49, 45, 51],\n        [42, 54, 56, 58, 60, 62, 58],\n        [41, 53, 55, 57, 59, 61, 57],\n        [54, 53],\n\n        \n        [80, 81, 82, 83, 84, ],\n        [80, 85, 86, 87, 88],\n        [80, 89, 90, 91, 92],\n        [80, 93, 94, 95, 96],\n        [80, 97, 98, 99, 100], ]\n\n    parquet_df = map_new_to_old_style(parquet_df)\n    frames = sorted(set(parquet_df.frame))\n    first_frame = min(frames)\n    parquet_df['color'] = parquet_df.type.apply(lambda row: assign_color(row))\n    parquet_df['plot_order'] = parquet_df.apply(lambda row: assign_order(row), axis=1)\n    first_frame_df = parquet_df[parquet_df.frame == first_frame].copy()\n    first_frame_df = first_frame_df.sort_values([\"plot_order\"]).set_index('plot_order')\n\n\n    frames_l = []\n    for frame in frames:\n        filtered_df = parquet_df[parquet_df.frame == frame].copy()\n        filtered_df = filtered_df.sort_values([\"plot_order\"]).set_index(\"plot_order\")\n        traces = [go.Scatter(\n            x=filtered_df['x'],\n            y=filtered_df['y'],\n            mode='markers',\n            marker=dict(\n                color=filtered_df.color,\n                size=9))]\n\n        for i, seg in enumerate(connections):\n            trace = go.Scatter(\n                    x=filtered_df.loc[seg]['x'],\n                    y=filtered_df.loc[seg]['y'],\n                    mode='lines',\n            )\n            traces.append(trace)\n        frame_data = go.Frame(data=traces, traces = [i for i in range(17)])\n        frames_l.append(frame_data)\n\n    traces = [go.Scatter(\n        x=first_frame_df['x'],\n        y=first_frame_df['y'],\n        mode='markers',\n        marker=dict(\n            color=first_frame_df.color,\n            size=9\n        )\n    )]\n    for i, seg in enumerate(connections):\n        trace = go.Scatter(\n            x=first_frame_df.loc[seg]['x'],\n            y=first_frame_df.loc[seg]['y'],\n            mode='lines',\n            line=dict(\n                color='black',\n                width=2\n            )\n        )\n        traces.append(trace)\n    fig = go.Figure(\n        data=traces,\n        frames=frames_l\n    )\n\n\n    fig.update_layout(\n        width=500,\n        height=800,\n        scene={\n            'aspectmode': 'data',\n        },\n        updatemenus=[\n            {\n                \"buttons\": [\n                    {\n                        \"args\": [None, {\"frame\": {\"duration\": 100,\n                                                  \"redraw\": True},\n                                        \"fromcurrent\": True,\n                                        \"transition\": {\"duration\": 0}}],\n                        \"label\": \"&#9654;\",\n                        \"method\": \"animate\",\n                    },\n                    {\n                        \"args\": [[None], {\"frame\": {\"duration\": 0, \"redraw\": False},\n                                          \"mode\": \"immediate\",\n                                          \"transition\": {\"duration\": 0}}],\n                        \"label\": \"&#9612;&#9612;\",\n                        \"method\": \"animate\",\n                    },\n                ],\n                \"direction\": \"left\",\n                \"pad\": {\"r\": 100, \"t\": 100},\n                \"font\": {\"size\":20},\n                \"type\": \"buttons\",\n                \"x\": 0.1,\n                \"y\": 0,\n            }\n        ],\n    )\n    camera = dict(\n        up=dict(x=0, y=-1, z=0),\n        eye=dict(x=0, y=0, z=2.5)\n    )\n    fig.update_layout(title_text=title, title_x=0.5)\n    fig.update_layout(scene_camera=camera, showlegend=False)\n    fig.update_layout(xaxis = dict(visible=False),\n            yaxis = dict(visible=False),\n    )\n    fig.update_yaxes(autorange=\"reversed\")\n\n    fig.show()\n    \ndef get_phrase(df, file_id, sequence_id):\n    return df[\n        np.logical_and(\n            df.file_id == file_id, \n            df.sequence_id == sequence_id\n        )\n    ].phrase.iloc[0]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-05-14T11:23:22.210937Z","iopub.execute_input":"2023-05-14T11:23:22.212522Z","iopub.status.idle":"2023-05-14T11:23:22.249877Z","shell.execute_reply.started":"2023-05-14T11:23:22.212471Z","shell.execute_reply":"2023-05-14T11:23:22.24848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explore Metadata","metadata":{}},{"cell_type":"markdown","source":"### explore supplemental_metadata","metadata":{}},{"cell_type":"code","source":"# Load the supplemental_metadata.csv file into memory\nsupplemental_df = pd.read_csv(\"/kaggle/input/asl-fingerspelling/supplemental_metadata.csv\")\npd.set_option('display.max_columns', None)\nsupplemental_df.head(3)","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:13.720783Z","iopub.execute_input":"2023-05-14T14:11:13.721188Z","iopub.status.idle":"2023-05-14T14:11:13.912242Z","shell.execute_reply.started":"2023-05-14T14:11:13.72116Z","shell.execute_reply":"2023-05-14T14:11:13.911026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## get count phrases\nphrase_count = supplemental_df[\"phrase\"]","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:13.986964Z","iopub.execute_input":"2023-05-14T14:11:13.98747Z","iopub.status.idle":"2023-05-14T14:11:13.993096Z","shell.execute_reply.started":"2023-05-14T14:11:13.987429Z","shell.execute_reply":"2023-05-14T14:11:13.991951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## get count of unique phrases\nunique_phrase = supplemental_df[\"phrase\"].unique()","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:14.227736Z","iopub.execute_input":"2023-05-14T14:11:14.228499Z","iopub.status.idle":"2023-05-14T14:11:14.244582Z","shell.execute_reply.started":"2023-05-14T14:11:14.228459Z","shell.execute_reply":"2023-05-14T14:11:14.243259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"number of phrase is : {} and number of unique phrase is : {}\".format(len(phrase_count), len(unique_phrase)))","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:14.35293Z","iopub.execute_input":"2023-05-14T14:11:14.354Z","iopub.status.idle":"2023-05-14T14:11:14.361062Z","shell.execute_reply.started":"2023-05-14T14:11:14.353953Z","shell.execute_reply":"2023-05-14T14:11:14.359931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# type(phrase_count),type(unique_phrase)","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:14.363957Z","iopub.execute_input":"2023-05-14T14:11:14.364894Z","iopub.status.idle":"2023-05-14T14:11:14.370304Z","shell.execute_reply.started":"2023-05-14T14:11:14.364837Z","shell.execute_reply":"2023-05-14T14:11:14.369143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### create separete dataframe to store phrases and their value counts","metadata":{}},{"cell_type":"code","source":"# Get the value counts for the 'phrase' columns\nphrase_counts = supplemental_df['phrase'].value_counts()\n\n# Create a new DatFrame with 'phrase' and 'count' columns \nphrase_data = pd.DataFrame({'phrases': phrase_counts.index, 'phrase_count': phrase_counts.values})","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:14.372192Z","iopub.execute_input":"2023-05-14T14:11:14.37259Z","iopub.status.idle":"2023-05-14T14:11:14.394033Z","shell.execute_reply.started":"2023-05-14T14:11:14.372548Z","shell.execute_reply":"2023-05-14T14:11:14.392735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"phrase_data.head(10)\n","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:14.397146Z","iopub.execute_input":"2023-05-14T14:11:14.398698Z","iopub.status.idle":"2023-05-14T14:11:14.409478Z","shell.execute_reply.started":"2023-05-14T14:11:14.398657Z","shell.execute_reply":"2023-05-14T14:11:14.408427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### visualize data for 5 most frequent and least frequent phrases","metadata":{}},{"cell_type":"code","source":"fig = px.bar(phrase_data.iloc[:5,:], x='phrase_count', y='phrases', color='phrases', orientation='h')\nfig.update_layout(\n    title={\n        'text': \"count of top 5 most frequent phrases\",\n        'y':0.96,\n        'x':0.4,\n        'xanchor': 'center',\n        'yanchor': 'top'\n    },\n    legend_title_text='Aspect:'\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:14.412926Z","iopub.execute_input":"2023-05-14T14:11:14.413849Z","iopub.status.idle":"2023-05-14T14:11:14.911677Z","shell.execute_reply.started":"2023-05-14T14:11:14.413811Z","shell.execute_reply":"2023-05-14T14:11:14.910324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(phrase_data.iloc[504:508,:], x='phrase_count', y='phrases', color='phrases', orientation='h')\nfig.update_layout(\n    title={\n        'text': \"count of 5 least phrases\",\n        'y':0.96,\n        'x':0.4,\n        'xanchor': 'center',\n        'yanchor': 'top'\n    },\n    legend_title_text='Aspect:'\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:14.914836Z","iopub.execute_input":"2023-05-14T14:11:14.915251Z","iopub.status.idle":"2023-05-14T14:11:15.018853Z","shell.execute_reply.started":"2023-05-14T14:11:14.915199Z","shell.execute_reply":"2023-05-14T14:11:15.017766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explore Landmark","metadata":{}},{"cell_type":"markdown","source":"## loading parquest file of ","metadata":{}},{"cell_type":"code","source":"##create subset of dataset where phrase is \"coming up with killer sound bites\"\ntop_phrase = supplemental_df[supplemental_df[\"phrase\"]==\"coming up with killer sound bites\"]['path'].values[0]\ntop_phrase","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:15.020397Z","iopub.execute_input":"2023-05-14T14:11:15.021233Z","iopub.status.idle":"2023-05-14T14:11:15.041871Z","shell.execute_reply.started":"2023-05-14T14:11:15.02117Z","shell.execute_reply":"2023-05-14T14:11:15.040429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base_dir=Path(\"/kaggle/input/asl-fingerspelling\")","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:15.043651Z","iopub.execute_input":"2023-05-14T14:11:15.044112Z","iopub.status.idle":"2023-05-14T14:11:15.050918Z","shell.execute_reply.started":"2023-05-14T14:11:15.044072Z","shell.execute_reply":"2023-05-14T14:11:15.049822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### explore landmark file of top_phrase","metadata":{}},{"cell_type":"code","source":"landmark_file = pd.read_parquet(base_dir/top_phrase)\nlandmark_file.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:15.052493Z","iopub.execute_input":"2023-05-14T14:11:15.053006Z","iopub.status.idle":"2023-05-14T14:11:34.948233Z","shell.execute_reply.started":"2023-05-14T14:11:15.052962Z","shell.execute_reply":"2023-05-14T14:11:34.947295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark_file=landmark_file.reset_index(inplace=False)","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:34.949509Z","iopub.execute_input":"2023-05-14T14:11:34.950072Z","iopub.status.idle":"2023-05-14T14:11:35.500465Z","shell.execute_reply.started":"2023-05-14T14:11:34.950043Z","shell.execute_reply":"2023-05-14T14:11:35.499265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark_file.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:35.50177Z","iopub.execute_input":"2023-05-14T14:11:35.502271Z","iopub.status.idle":"2023-05-14T14:11:36.694743Z","shell.execute_reply.started":"2023-05-14T14:11:35.502242Z","shell.execute_reply":"2023-05-14T14:11:36.693691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# len(landmark_file.columns)  ## 1630","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:36.696168Z","iopub.execute_input":"2023-05-14T14:11:36.696588Z","iopub.status.idle":"2023-05-14T14:11:36.701672Z","shell.execute_reply.started":"2023-05-14T14:11:36.696551Z","shell.execute_reply":"2023-05-14T14:11:36.700764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark_file.shape","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:36.703139Z","iopub.execute_input":"2023-05-14T14:11:36.704215Z","iopub.status.idle":"2023-05-14T14:11:36.716179Z","shell.execute_reply.started":"2023-05-14T14:11:36.704157Z","shell.execute_reply":"2023-05-14T14:11:36.715107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# view number of unique sequence_ids in dataset \n# landmark_file[\"sequence_id\"].nunique() # 1000","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:36.72285Z","iopub.execute_input":"2023-05-14T14:11:36.723563Z","iopub.status.idle":"2023-05-14T14:11:36.728183Z","shell.execute_reply.started":"2023-05-14T14:11:36.723522Z","shell.execute_reply":"2023-05-14T14:11:36.727054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# return 1st two sequence_ids\nlandmark_file[\"sequence_id\"].unique()[:2]","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:36.730017Z","iopub.execute_input":"2023-05-14T14:11:36.73043Z","iopub.status.idle":"2023-05-14T14:11:36.745136Z","shell.execute_reply.started":"2023-05-14T14:11:36.730394Z","shell.execute_reply":"2023-05-14T14:11:36.744076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# landmark_file[\"frame\"].nunique() # 507","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:36.746698Z","iopub.execute_input":"2023-05-14T14:11:36.747062Z","iopub.status.idle":"2023-05-14T14:11:36.752102Z","shell.execute_reply.started":"2023-05-14T14:11:36.747033Z","shell.execute_reply":"2023-05-14T14:11:36.750961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#fetch landmark data for sequence id=1535467051\nlandmark_1st_id=landmark_file[landmark_file[\"sequence_id\"]==1535467051]","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:36.753802Z","iopub.execute_input":"2023-05-14T14:11:36.754098Z","iopub.status.idle":"2023-05-14T14:11:36.765827Z","shell.execute_reply.started":"2023-05-14T14:11:36.754072Z","shell.execute_reply":"2023-05-14T14:11:36.764662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmark_1st_id","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:36.767276Z","iopub.execute_input":"2023-05-14T14:11:36.767626Z","iopub.status.idle":"2023-05-14T14:11:38.310824Z","shell.execute_reply.started":"2023-05-14T14:11:36.767596Z","shell.execute_reply":"2023-05-14T14:11:38.309517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# explore train file","metadata":{}},{"cell_type":"code","source":"train_data=pd.read_csv(\"/kaggle/input/asl-fingerspelling/train.csv\")\ntrain_data.shape","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:38.312446Z","iopub.execute_input":"2023-05-14T14:11:38.312875Z","iopub.status.idle":"2023-05-14T14:11:38.474016Z","shell.execute_reply.started":"2023-05-14T14:11:38.312829Z","shell.execute_reply":"2023-05-14T14:11:38.472682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:38.477577Z","iopub.execute_input":"2023-05-14T14:11:38.47792Z","iopub.status.idle":"2023-05-14T14:11:38.490629Z","shell.execute_reply.started":"2023-05-14T14:11:38.47789Z","shell.execute_reply":"2023-05-14T14:11:38.489416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### explore file ->/kaggle/input/asl-fingerspelling/character_to_prediction_index.json","metadata":{}},{"cell_type":"code","source":"char_to_pred=\"/kaggle/input/asl-fingerspelling/character_to_prediction_index.json\"\n# Python program to read\n# json file\nchar=[]\nvalues=[]\nimport json\n\n# Opening JSON file\nf = open(char_to_pred)\n\n# returns JSON object as\n# a dictionary\ndata = json.load(f)\n\n# Iterating through the json\n# list\nfor i,j in data.items():\n    char.append(i)\n    values.append(j)\n#   print(\"key:\"+str(i),\"values:\"+str(j))\n\n# Closing file\nf.close()\n\n# print(\"\\n characters list:\",char)\n# print(\"\\n values list:\",values)","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:38.492117Z","iopub.execute_input":"2023-05-14T14:11:38.492586Z","iopub.status.idle":"2023-05-14T14:11:38.505499Z","shell.execute_reply.started":"2023-05-14T14:11:38.492549Z","shell.execute_reply":"2023-05-14T14:11:38.504192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"char_to_pred_index=pd.DataFrame({\"char\":char,\"values\":values})\nchar_to_pred_index.head(20)","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:38.506933Z","iopub.execute_input":"2023-05-14T14:11:38.50732Z","iopub.status.idle":"2023-05-14T14:11:38.519474Z","shell.execute_reply.started":"2023-05-14T14:11:38.507289Z","shell.execute_reply":"2023-05-14T14:11:38.518635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualizing the Sequence","metadata":{}},{"cell_type":"markdown","source":"### Let's get random file ID and sequence to Visualize the sequence","metadata":{}},{"cell_type":"code","source":"# Get number of unique file_ids in train folder\nunique_file_ids = len(np.unique(supplemental_df['file_id']))\n\n# Generate a random integer between 0 and 53 (unique_file_ids)\nrandom_id = np.random.randint(0, unique_file_ids)\n\n# Getting the random file_id \nrandom_file_id = np.unique(supplemental_df['file_id'])[random_id]","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:38.520389Z","iopub.execute_input":"2023-05-14T14:11:38.520705Z","iopub.status.idle":"2023-05-14T14:11:38.533554Z","shell.execute_reply.started":"2023-05-14T14:11:38.520679Z","shell.execute_reply":"2023-05-14T14:11:38.532208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get all different sequences in random file\nsigns = supplemental_df[supplemental_df['file_id'] == random_file_id ]","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:38.534915Z","iopub.execute_input":"2023-05-14T14:11:38.53537Z","iopub.status.idle":"2023-05-14T14:11:38.543851Z","shell.execute_reply.started":"2023-05-14T14:11:38.535332Z","shell.execute_reply":"2023-05-14T14:11:38.542728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"signs.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:38.545251Z","iopub.execute_input":"2023-05-14T14:11:38.545688Z","iopub.status.idle":"2023-05-14T14:11:38.561139Z","shell.execute_reply.started":"2023-05-14T14:11:38.545653Z","shell.execute_reply":"2023-05-14T14:11:38.560256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The number of unique sequences in each .parquet file is 1000.","metadata":{}},{"cell_type":"code","source":"len(np.unique(signs.index))","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:38.562421Z","iopub.execute_input":"2023-05-14T14:11:38.56347Z","iopub.status.idle":"2023-05-14T14:11:38.573403Z","shell.execute_reply.started":"2023-05-14T14:11:38.563439Z","shell.execute_reply":"2023-05-14T14:11:38.572509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get a random Sequence id\nrandom_squence_id = signs.sample()['sequence_id'].item()","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:38.574725Z","iopub.execute_input":"2023-05-14T14:11:38.575345Z","iopub.status.idle":"2023-05-14T14:11:38.58372Z","shell.execute_reply.started":"2023-05-14T14:11:38.57531Z","shell.execute_reply":"2023-05-14T14:11:38.582505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Let's Load random file id to visualize the sequence ","metadata":{}},{"cell_type":"code","source":"path_to_sign = f\"/kaggle/input/asl-fingerspelling/supplemental_landmarks/{random_file_id}.parquet\"\nparquet = pd.read_parquet(path_to_sign)","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:11:38.585511Z","iopub.execute_input":"2023-05-14T14:11:38.586085Z","iopub.status.idle":"2023-05-14T14:12:02.102696Z","shell.execute_reply.started":"2023-05-14T14:11:38.586055Z","shell.execute_reply":"2023-05-14T14:12:02.101764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sequence = parquet[parquet.index == random_squence_id]\nsequence","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:12:02.104243Z","iopub.execute_input":"2023-05-14T14:12:02.104654Z","iopub.status.idle":"2023-05-14T14:12:03.760259Z","shell.execute_reply.started":"2023-05-14T14:12:02.104616Z","shell.execute_reply":"2023-05-14T14:12:03.759411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sequence_phrase = get_phrase(supplemental_df, random_file_id, random_squence_id)\nvisualise2d_landmarks(sequence, f\"Phrase: {sequence_phrase}\")","metadata":{"execution":{"iopub.status.busy":"2023-05-14T14:12:03.76224Z","iopub.execute_input":"2023-05-14T14:12:03.762618Z","iopub.status.idle":"2023-05-14T14:12:04.04092Z","shell.execute_reply.started":"2023-05-14T14:12:03.762587Z","shell.execute_reply":"2023-05-14T14:12:04.039053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## _________________________THANK YOU !!! ___________________","metadata":{}}]}