{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1><strong>American Sign Language Fingerspelling Recognition</strong></h1>","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"# **Overview**","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://studybreaks.com/wp-content/uploads/2016/06/asl.jpg\" alt=\"ASL Sign\"/>\n\nSigned languages use the visual-gestural modality to convey meaning through manual articulations in combination with non-manual elements like the face and body. They serve as the primary means of communication for numerous deaf and hard-of-hearing individuals.\n\nSign Language Processing is an emerging field of artificial intelligence concerned with the automatic processing and analysis of sign language content. The primary objective of this competition is to recognize and interpret American Sign Language fingerspelling into text. Fingerspelling uses hand shapes that represent individual letters to convey words.\n\nIn this notebook, I analyze and provide insights into the competition data. Let's first [import](#import-deps) some dependencies.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"import-deps\"></a>\n# **Import dependencies**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport plotly.io as pio\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\n\npd.options.plotting.backend = \"plotly\"","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-06-05T17:13:30.368259Z","iopub.execute_input":"2023-06-05T17:13:30.368725Z","iopub.status.idle":"2023-06-05T17:13:31.446971Z","shell.execute_reply.started":"2023-06-05T17:13:30.36869Z","shell.execute_reply":"2023-06-05T17:13:31.445837Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import sequence visualization code from Leonid Kulyk's notebook. Thanks Leonid!\n# Link: https://www.kaggle.com/code/leonidkulyk/eda-aslfr-animated-visualization\n\ndef map_new_to_old_style(sequence):\n    types = []\n    landmark_indexes = []\n    for column in list(sequence.columns)[1:544]:\n        parts = column.split(\"_\")\n        if len(parts) == 4:\n            types.append(parts[1] + \"_\" + parts[2])\n        else:\n            types.append(parts[1])\n\n        landmark_indexes.append(int(parts[-1]))\n\n    data = {\n        \"frame\": [],\n        \"type\": [],\n        \"landmark_index\": [],\n        \"x\": [],\n        \"y\": [],\n        \"z\": []\n    }\n\n    for index, row in sequence.iterrows():\n        data[\"frame\"] += [int(row.frame)]*543\n        data[\"type\"] += types\n        data[\"landmark_index\"] += landmark_indexes\n\n        for _type, landmark_index in zip(types, landmark_indexes):\n            data[\"x\"].append(row[f\"x_{_type}_{landmark_index}\"])\n            data[\"y\"].append(row[f\"y_{_type}_{landmark_index}\"])\n            data[\"z\"].append(row[f\"z_{_type}_{landmark_index}\"])\n\n    return pd.DataFrame.from_dict(data)\n\n# assign desired colors to landmarks\ndef assign_color(row):\n    if row == 'face':\n        return 'red'\n    elif 'hand' in row:\n        return 'dodgerblue'\n    else:\n        return 'green'\n\n# specifies the plotting order\ndef assign_order(row):\n    if row.type == 'face':\n        return row.landmark_index + 101\n    elif row.type == 'pose':\n        return row.landmark_index + 30\n    elif row.type == 'left_hand':\n        return row.landmark_index + 80\n    else:\n        return row.landmark_index\n\ndef visualise2d_landmarks(parquet_df, title=\"\"):\n    connections = [  \n        [0, 1, 2, 3, 4,],\n        [0, 5, 6, 7, 8],\n        [0, 9, 10, 11, 12],\n        [0, 13, 14, 15, 16],\n        [0, 17, 18, 19, 20],\n\n        \n        [38, 36, 35, 34, 30, 31, 32, 33, 37],\n        [40, 39],\n        [52, 46, 50, 48, 46, 44, 42, 41, 43, 45, 47, 49, 45, 51],\n        [42, 54, 56, 58, 60, 62, 58],\n        [41, 53, 55, 57, 59, 61, 57],\n        [54, 53],\n\n        \n        [80, 81, 82, 83, 84, ],\n        [80, 85, 86, 87, 88],\n        [80, 89, 90, 91, 92],\n        [80, 93, 94, 95, 96],\n        [80, 97, 98, 99, 100], ]\n\n    parquet_df = map_new_to_old_style(parquet_df)\n    frames = sorted(set(parquet_df.frame))\n    first_frame = min(frames)\n    parquet_df['color'] = parquet_df.type.apply(lambda row: assign_color(row))\n    parquet_df['plot_order'] = parquet_df.apply(lambda row: assign_order(row), axis=1)\n    first_frame_df = parquet_df[parquet_df.frame == first_frame].copy()\n    first_frame_df = first_frame_df.sort_values([\"plot_order\"]).set_index('plot_order')\n\n\n    frames_l = []\n    for frame in frames:\n        filtered_df = parquet_df[parquet_df.frame == frame].copy()\n        filtered_df = filtered_df.sort_values([\"plot_order\"]).set_index(\"plot_order\")\n        traces = [go.Scatter(\n            x=filtered_df['x'],\n            y=filtered_df['y'],\n            mode='markers',\n            marker=dict(\n                color=filtered_df.color,\n                size=9))]\n\n        for i, seg in enumerate(connections):\n            trace = go.Scatter(\n                    x=filtered_df.loc[seg]['x'],\n                    y=filtered_df.loc[seg]['y'],\n                    mode='lines',\n            )\n            traces.append(trace)\n        frame_data = go.Frame(data=traces, traces = [i for i in range(17)])\n        frames_l.append(frame_data)\n\n    traces = [go.Scatter(\n        x=first_frame_df['x'],\n        y=first_frame_df['y'],\n        mode='markers',\n        marker=dict(\n            color=first_frame_df.color,\n            size=9\n        )\n    )]\n    for i, seg in enumerate(connections):\n        trace = go.Scatter(\n            x=first_frame_df.loc[seg]['x'],\n            y=first_frame_df.loc[seg]['y'],\n            mode='lines',\n            line=dict(\n                color='black',\n                width=2\n            )\n        )\n        traces.append(trace)\n        \n    fig = go.Figure(\n        data=traces,\n        frames=frames_l\n    )\n\n\n    fig.update_layout(\n        width=500,\n        height=1000,\n        scene={\n            'aspectmode': 'data',\n        },\n        updatemenus=[\n            {\n                \"buttons\": [\n                    {\n                        \"args\": [None, {\"frame\": {\"duration\": 100,\n                                                  \"redraw\": True},\n                                        \"fromcurrent\": True,\n                                        \"transition\": {\"duration\": 0}}],\n                        \"label\": \"&#9654; Play\",\n                        \"method\": \"animate\",\n                    },\n\n                ],\n                \"direction\": \"left\",\n                \"font\": {\"size\": 18},\n                \"type\": \"buttons\",\n                \"xanchor\": \"left\",\n                \"yanchor\": \"top\",\n                \"x\": 0.4,\n                \"y\": 0,\n            }\n        ],\n    )\n    camera = dict(\n        up=dict(x=0, y=-1, z=0),\n        eye=dict(x=0, y=0, z=2.5)\n    )\n    \n    fig.update_layout(\n        title = {\n            'text': f'<b>{title}</b>',\n            'y': 0.97,\n            'x': 0.5,\n            'xanchor': 'center',\n            'yanchor': 'top',\n            'font_size': 24\n        },\n        scene_camera=camera,\n        showlegend=False,\n        margin = dict(t = 20, b = 20, l = 20, r = 20),\n        xaxis = dict(visible=False),\n        yaxis = dict(visible=False),\n        template = \"plotly_white\",\n    )\n    \n    fig.update_yaxes(autorange=\"reversed\")\n    fig.show()\n\n\ndef get_phrase(df, file_id, sequence_id):\n    return df[\n        np.logical_and(\n            df.file_id == file_id, \n            df.sequence_id == sequence_id\n        )\n    ].phrase.iloc[0]","metadata":{"execution":{"iopub.status.busy":"2023-06-05T17:13:31.449532Z","iopub.execute_input":"2023-06-05T17:13:31.450026Z","iopub.status.idle":"2023-06-05T17:13:31.492007Z","shell.execute_reply.started":"2023-06-05T17:13:31.449985Z","shell.execute_reply":"2023-06-05T17:13:31.490533Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Data Overview**","metadata":{}},{"cell_type":"markdown","source":"This is the dataset structure.\n```\nasl-fingerspelling\n├── character_to_prediction_index.json\n├── supplemental_landmarks\n│   └── 1032110484.parquet\n│   └── 1047404576.parquet\n│   └── ... \n├── supplemental_metadata.csv\n├── train.csv\n├── train_landmarks\n│   └── 1019715464.parquet\n│   └── 1021040628.parquet\n│   └── ...\n```\nThe dataset has two major parts, the supplimental landmarks and the training landmarks set. The landmarks seem to be distributed into several `.parquet` files. It would seem the metadata and data definitions for the supplemental and training landmarks are available in the respective csv files. Let's look into that in the [EDA](#eda) section.","metadata":{"execution":{"iopub.status.busy":"2023-06-04T10:29:34.448584Z","iopub.execute_input":"2023-06-04T10:29:34.448932Z","iopub.status.idle":"2023-06-04T10:29:34.457102Z","shell.execute_reply.started":"2023-06-04T10:29:34.448903Z","shell.execute_reply":"2023-06-04T10:29:34.454566Z"}}},{"cell_type":"markdown","source":"<a id=\"eda\"></a>\n# **EDA**","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(\"../input/asl-fingerspelling/train.csv\")","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that the **train_df** has the following columns:\n- **path**: Path to the landmarks file. We will come to this later when we visualize one of these sequences.\n- **file_id**: The file id which stores this sequence.\n- **sequence_id**: The sequence which the participant `participant_id` has signed which represents the `phrase`.\n- **participant_id**: The participant ID who has signed this sequence\n- **phrase**: The phrase which has been signed into this sequence.\n\nLet's look into this in much more detail.  \nEach row of this dataset represents a single **Mediapipe holistic** landmark sequence which is signed by a participant. The actual text is  stored in the `phrase` column. This means that the number of unique sequences should be equal to the number of datapoints, since each row represents a single sequence.","metadata":{}},{"cell_type":"code","source":"# We are trying to assert that the count of unique sequences is equal to the number of datapoints \nassert train_df[\"sequence_id\"].nunique() == train_df.shape[0] # True\n\n# We are trying to assert that each unique file id should have a unique path associated to it\nassert train_df[\"path\"].nunique() == train_df[\"file_id\"].nunique() # True","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"There are total {train_df['path'].nunique()} files that store these {train_df['sequence_id'].nunique()} sequences.\")","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let us understand the distribution of these sequences into these files.","metadata":{}},{"cell_type":"code","source":"file_seq_count = train_df.groupby([\"file_id\"])[\"sequence_id\"].count().value_counts()\ncolors = ['gainsboro', 'powderblue']\nfig = make_subplots(\n    rows=1,\n    cols=2,\n    column_widths=[0.4, 0.6],\n    specs=[[{\"type\": \"domain\"}, {\"type\": \"xy\"}]],\n    horizontal_spacing=0.2\n)\n\nfig.add_trace(\n    go.Pie(\n        labels=file_seq_count.index,\n        values=file_seq_count.values,\n        texttemplate = \"<b>%{label}</b> sequences in <b>%{value}</b> files\",\n    ),\n        row=1,\n    col=1\n)\n\nfig.update_traces(\n    hoverinfo='percent',\n    textinfo='label',\n    textfont_size=16,\n    marker=dict(\n        colors=colors,\n        line=dict(color='#000000', width=2)\n    )\n)\n\nfig.add_trace(\n    go.Bar(\n        x=file_seq_count.values,\n        y=[str(x) for x in file_seq_count.index],\n        orientation='h',\n        text=file_seq_count.values,\n        marker_color=colors,\n        hoverinfo=\"none\"\n    )\n)\n\nfig.update_layout(\n    title = {\n        'text': f'<b>Distibution of sequences vs file counts</b>',\n        'y': 0.95,\n        'x': 0.5,\n        'xanchor': 'center',\n        'yanchor': 'top',\n        'font_size': 22\n    },\n    showlegend=False,\n    title_x = 0.5,\n    hoverlabel = dict(\n        font_size = 14,\n    ),\n    xaxis_title=\"Number of files\",\n    yaxis_title=\"Count of sequences\",\n    template = \"plotly_white\",\n    margin = dict(t = 50, b = 20, l = 50, r = 20)\n)\nfig.show()","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that we have 67 files (98.5 %) that contain 1000 sequences each and 1 of the files (1.5%) containing the rest 287 samples. ","metadata":{}},{"cell_type":"code","source":"# Loop over all parquet files and collect data\nframes = {} # Save total frames as a dict\nframes_list = {} # Save a dict of frames, I know, it seems redundant\nfile_list = train_df[\"path\"].unique() # List of files\nfor index, path in enumerate(file_list):\n    df = pd.read_parquet(\"../input/asl-fingerspelling/\" + path) # read the corresponding parquet file into a dataframe\n    frame_count = df.groupby(df.index)[\"frame\"].count().tolist() # count the number of frames and store into an array\n    frames[path.split(\"/\")[1].split(\".\")[0]] = frame_count # parse the label for the file\n    frames_list |= df.groupby(df.index)[\"frame\"].count().to_dict() # add sequence_id and frame counts to the frames_list","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.merge(train_df, pd.DataFrame(frames_list.items(), columns=[\"sequence_id\", \"frame_count\"]), on=\"sequence_id\")\ntrain_df[\"phrase_len\"] = train_df[\"phrase\"].str.len()\ntrain_df[\"word_count\"] = train_df[\"phrase\"].str.split(\" \").apply(len)","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_frame_distribution(frames):\n    fig=make_subplots(\n        rows=1,\n        cols=2,\n        shared_yaxes=True,\n        subplot_titles=[\"Count\", \"Distribution\"],\n        column_widths=[0.85, 0.15],\n        horizontal_spacing=0.01\n    )\n\n    for key in frames.keys():\n        trace1 = go.Scatter(\n            y=frames[key],\n            name=key,\n            hovertemplate='Sequence Number: <b>%{x}</b><br>Number of frames: <b>%{y}</b><extra></extra>',\n            marker_color='#484bfa',\n            opacity=0.75,\n        )\n        fig.append_trace(trace1, 1, 1)\n        trace2 = go.Violin(\n            y=frames[key],\n            name=key,\n            fillcolor='aliceblue',\n            line_color='#484bfa',\n            hovertemplate='Number of frames: <b>%{y}</b><extra></extra>',\n            box_visible=True,\n            opacity=0.75,\n            meanline_visible=True,\n            x0=key\n        )\n        fig.append_trace(trace2, 1, 2)\n\n    length_of_data = len(fig.data)\n    length_of_frames = len(frames.keys())\n    for data in range(2, length_of_data):\n        fig.update_traces(visible=False, selector=data)\n\n    def create_layout_button(index, frame):\n        visibility=[False] * 2 * length_of_frames\n        for k in [2 * index, 2 * index + 1]:\n            visibility[k] =True\n        return dict(\n            label = frame,\n            method = 'restyle',\n            args = [\n                {\n                    'visible': visibility,\n                    'title': frame,\n                    'showlegend': False\n                }\n            ]\n        )    \n\n    fig.update_layout(\n        title={\n            'text': f'<b>Sequence frame count distribution by file</b>',\n            'y': 0.98,\n            'x': 0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'\n        },\n        annotations=[\n            {\n                \"text\": \"File ID:\",\n                \"x\": 0.42,\n                \"y\": 1.01,\n                \"align\": \"left\",\n                \"showarrow\": False\n            }\n        ],\n        updatemenus=[\n            go.layout.Updatemenu(\n                active=0,\n                buttons=[create_layout_button(index, frame) for index, frame in enumerate(frames)],\n                direction=\"down\",\n                x=0.45,\n                xanchor='left',\n                y=1,\n                yanchor='bottom',\n            )\n        ],\n        hoverlabel=dict(\n            bgcolor=\"white\",\n            font_size=14,\n        ),\n        title_font_size=24,\n        template=\"plotly_white\",\n        margin=dict(t=80, b=20, l=20, r=20),\n        yaxis_title=\"Number of Frames\",\n        showlegend=False\n    )\n    \n    fig.show()\n\nplot_frame_distribution(frames)","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Using the above visualization we can interact with the file data in a much better manner. But wait, we can already notice something strange, several of these sequences have very low frames in them. Well we know that Mediapipe uses image/video data and signs sequences using machine learning. If you have played around with Mediapipe you know that a lot of times the system is not able to capture landmarks from the data. This low sequence count begs the question, whether something useful was actually signed in those sequences. Let us look into that.","metadata":{}},{"cell_type":"code","source":"fig = go.Figure(\n    data=go.Box(\n        x=train_df[\"phrase_len\"],\n        y=train_df[\"frame_count\"],\n        boxpoints=\"outliers\",\n        boxmean=True,\n    )\n)\n\nfig.update_traces(\n    marker_color='rgb(158,202,250)',\n    marker_line_color='rgb(8,48,107)',\n)\n\nfig.update_layout(\n    title={\n        'text':f'<b>Distribution of frame length vs number of frames</b>',\n        'y':0.95,\n        'x':0.5,\n        'xanchor':'center',\n        'yanchor':'top'\n    },\n    xaxis_title=\"Phrase length\",\n    yaxis_title=\"Number of frames\",\n    height=500,\n    title_x=0.5,\n    hoverlabel=dict(\n        font_size=14,\n    ),\n    title_font_size=24,\n    template=\"plotly_white\",\n    margin=dict(t=60,b=20,l=20,r=20),\n    yaxis_tickformat='digits',\n    margin_pad=5\n)\n\nfig.show()","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = go.Figure(\n    data=[\n        go.Histogram(\n            x=train_df[\"frame_count\"],\n            xbins=dict(size=10),\n            hovertemplate='Number of frames: <b>%{x}</b><br>Number of sequences: <b>%{y}</b><extra></extra>'\n        )\n    ]\n)\n\nfig.update_traces(\n    marker_color='rgb(158,202,225)',\n    marker_line_color='rgb(8,48,107)',\n)\n\nfig.update_layout(\n    title={\n        'text':f'<b>Number of frames in landmark sequences</b>',\n        'y':0.95,\n        'x':0.5,\n        'xanchor':'center',\n        'yanchor':'top'\n    },\n    xaxis_title=\"Number of frames\",\n    yaxis_title=\"Number of sequences\",\n    height=500,\n    title_x=0.5,\n    hoverlabel=dict(\n        font_size=14,\n    ),\n    title_font_size=24,\n    template=\"plotly_white\",\n    margin=dict(t=60,b=20,l=20,r=20),\n    yaxis_tickformat='digits',\n    margin_pad=5\n)\n\nfig.show()","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Aha! Just like we suspected. If you observe the first few bins (0-10, 10-20), There are a lot of sequences with very low number of frames. Now this may be due to several reasons, it could be \"good\" data that simply comprises of small sequences, or it could be \"bad\" data that crept in without the participants and organizers/maintainers noticing. To provide a conclusive reason, we need to dig into this more. Off the top of my head, I can think of two analyses that could be done:  \n1. Since this is a histogram data, the next logical thing we can do is to look more closely into the first few bins to get more information on the potential \"bad\" data problem.\n2. The only thing that can give a better insight into this is to check the actual phrases that they are signing. The small frame counts could very well be explained if they are actually signing let's say one or two character lengths, although certainly not one single frame.","metadata":{}},{"cell_type":"code","source":"fig = go.Figure(\n    data=[\n        go.Histogram(\n            name=\"Decent frames\",\n            x=train_df[train_df[\"frame_count\"] > 20][\"frame_count\"],\n            xbins=dict(start=20, size=10),\n            marker_color='rgb(158,202,225)',\n            marker_line_color='rgb(8,48,107)',\n            hovertemplate='Number of frames: <b>%{x}</b><br>Number of sequences: <b>%{y}</b><extra></extra>'\n        ),\n        go.Histogram(\n            name=\"Low frames\",\n            x=train_df[train_df[\"frame_count\"] <= 20][\"frame_count\"],\n            xbins=dict(start=0, end=20, size=10),\n            marker_color='rgb(250, 180, 180)',\n            marker_line_color='rgb(250, 180, 180)',\n            hovertemplate='Number of frames: <b>%{x}</b><br>Number of sequences: <b>%{y}</b><extra></extra>'\n        )\n    ],\n    layout=go.Layout(barmode='overlay')\n)\n\nfig.update_layout(\n    title={\n        'text':f'<b>Number of frames in landmark sequences</b>',\n        'y':0.95,\n        'x':0.5,\n        'xanchor':'center',\n        'yanchor':'top'\n    },\n    xaxis_title=\"Number of frames\",\n    yaxis_title=\"Number of sequences\",\n    height=500,\n    title_x=0.5,\n    hoverlabel=dict(\n        font_size=14,\n    ),\n    title_font_size=24,\n    template=\"plotly_white\",\n    margin=dict(t=60,b=20,l=20,r=20),\n    yaxis_tickformat='digits',\n    margin_pad=5,\n    legend=dict(\n        x=0.75,\n        y=1,\n        traceorder=\"normal\",\n        font=dict(\n            family=\"sans-serif\",\n            size=12,\n            color=\"black\"\n        ),\n    )\n)\n\nfig.show()","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"low_frame_count = train_df.loc[train_df[\"frame_count\"] <= 20, [\"sequence_id\", \"frame_count\", \"phrase\", \"phrase_len\"]].reset_index(drop=True)","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"low_frame_count = low_frame_count.head(50)","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = make_subplots(\n    rows=1,\n    cols=2,\n    specs=[[{}, {}]],\n    shared_xaxes=True,\n    shared_yaxes=True,\n    horizontal_spacing=0,\n    subplot_titles=[\n        '<-- Frame Count',\n        'Phrase Length -->'\n    ]\n)\n\nfig.append_trace(\n    go.Bar(\n        name='Frame count',\n        y=low_frame_count['sequence_id'].astype(str),\n        x=low_frame_count['frame_count'],\n        orientation='h',\n        width=0.5,\n        marker_color='#80bff3',\n    ),\n    1,\n    1\n)\n\nfig.append_trace(\n    go.Bar(\n        name='Phrase length',\n        y=low_frame_count['sequence_id'].astype(str),\n        x=low_frame_count['phrase_len'],\n        orientation='h',\n        width=0.5,\n        marker_color='#0e668b',\n    ),\n    1,\n    2\n)\n\nfig.update_traces(\n    textposition='outside',\n    hovertemplate=\"<br>\".join(\n        [\"ID: <b>%{y}</b>\", \"Count: <b>%{x}</b><extra></extra>\"]\n    )\n)\n\nfig.update_layout(\n    template=\"plotly_white\",\n    title = {\n        'text': f'<b></b>',\n        'y': 0.99,\n        'x': 0.5,\n        'xanchor': 'center',\n        'yanchor': 'top',\n        'font_size': 22\n    },\n    legend=dict(\n        orientation=\"h\",\n        yanchor=\"bottom\",\n        y=1.01,\n        xanchor=\"right\",\n        x=0.58\n    ),\n    height=1000,\n    hoverlabel = dict(\n        font_size = 14,\n    ),\n    margin = dict(\n        t=100,\n        b=20,\n        l=50,\n        r=20\n    )\n)\nfig.update_yaxes(tickmode='linear')\nfig['layout']['xaxis1']['autorange'] = \"reversed\"\nfig['layout']['yaxis']['title']='Sequence ID'\n\nfig.show()","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = go.Figure(\n    data=[\n        go.Histogram(\n            x=train_df[\"phrase_len\"],\n            hovertemplate='Phrase length: <b>%{x}</b><br>Number of sequences: <b>%{y}</b><extra></extra>'\n        )\n    ]\n)\n\nfig.update_traces(\n    marker_color='rgb(158,202,225)',\n    marker_line_color='rgb(8,48,107)',\n)\n\nfig.update_layout(\n    title={\n        'text':f'<b>Phrase length distribution of landmark sequences</b>',\n        'y':0.95,\n        'x':0.5,\n        'xanchor':'center',\n        'yanchor':'top'\n    },\n    xaxis_title=\"Number of frames\",\n    yaxis_title=\"Number of sequences\",\n    height=500,\n    title_x=0.5,\n    hoverlabel=dict(\n        font_size=14,\n    ),\n    title_font_size=24,\n    template=\"plotly_white\",\n    margin=dict(t=60,b=20,l=20,r=20),\n    yaxis_tickformat='digits',\n    margin_pad=5\n)\n\nfig.show()","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"character_count = {}\n\nfor phrase in train_df[\"phrase\"]:\n    for char in phrase:\n        try:\n            character_count[char] += 1\n        except KeyError:\n            character_count[char] = 1\n            \ncharacter_count = dict(sorted(character_count.items(), key=lambda k: k[1], reverse=True))\n\nfig = go.Figure(\n    data=[\n        go.Bar(\n            x=list(character_count.keys()),\n            y=list(character_count.values()),\n            hovertemplate='Character: <b>%{x}</b><br>Count: <b>%{y}</b><extra></extra>'\n        )\n    ]\n)\n\nfig.update_traces(\n    marker_color='rgb(158,202,225)',\n    marker_line_color='rgb(8,48,107)',\n    marker_line_width=1.5,\n)\n\nfig.update_layout(\n    title={\n        'text':f'<b>Count of each character in training set</b>',\n        'y':0.95,\n        'x':0.5,\n        'xanchor':'center',\n        'yanchor':'top'\n    },\n    xaxis_title=\"Character\",\n    yaxis_title=\"Count\",\n    height=500,\n    title_x=0.5,\n    hoverlabel=dict(\n        font_size=14,\n    ),\n    title_font_size=24,\n    template=\"plotly_white\",\n    margin=dict(t=40,b=20,l=20,r=20),\n    yaxis_tickformat='digits',\n    margin_pad=5\n)\n\nfig.show()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_15_phrase = train_df[\"phrase\"].value_counts().head(15)\n\nfig = go.Figure(\n    data=[\n        go.Bar(\n            x=list(top_15_phrase.index),\n            y=list(top_15_phrase.values),\n            hovertemplate='Phrase: <b>%{x}</b><br>Count: <b>%{y}</b><extra></extra>'\n        )\n    ]\n)\n\nfig.update_traces(\n    marker_color='rgb(158,202,225)',\n    marker_line_color='rgb(8,48,107)',\n    marker_line_width=1.5,\n)\n\nfig.update_layout(\n    title={\n        'text':f'<b>Top 15 signs in training set</b>',\n        'y':0.95,\n        'x':0.5,\n        'xanchor':'center',\n        'yanchor':'top'\n    },\n    yaxis_title=\"Count\",\n    height=500,\n    title_x=0.5,\n    hoverlabel=dict(\n        font_size=14,\n    ),\n    title_font_size=24,\n    template=\"plotly_white\",\n    margin=dict(t=60,b=20,l=20,r=20),\n    margin_pad=5\n)\n\nfig.show()","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"participant_count = train_df[\"participant_id\"].value_counts().head(20)\n\nfig = go.Figure(\n    data=[\n        go.Bar(\n            y=list(map(str, participant_count.index)),\n            x=list(participant_count.values),\n            hovertemplate='Participant ID: <b>%{y}</b><br>Signed sequences: <b>%{x}</b><extra></extra>',\n            orientation=\"h\"\n        )\n    ]\n)\n\nfig.update_traces(\n    marker_color='rgb(158,202,225)',\n    marker_line_color='rgb(8,48,107)',\n    marker_line_width=1.5,\n)\n\nfig.update_layout(\n    title={\n        'text':f'<b>Participants with most sequences signed in training set</b>',\n        'y':0.95,\n        'x':0.5,\n        'xanchor':'center',\n        'yanchor':'top'\n    },\n    yaxis_title=\"Participant ID\",\n    xaxis_title=\"Count\",\n    yaxis_autorange=\"reversed\",\n    height=500,\n    title_x=0.5,\n    hoverlabel=dict(\n        font_size=14,\n    ),\n    title_font_size=24,\n    template=\"plotly_white\",\n    margin=dict(t=80,b=20,l=20,r=20),\n    margin_pad=5,\n)\n\nfig.show()","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Animation","metadata":{}},{"cell_type":"markdown","source":"Mediapipe Holistic model can be understood as a machine learning system that can detect human pose, face landmarks and hand tracking. The landmark dataset can be understood as the integration of above models\n- 33 pose\n- 468 face\n- 21 left_hand\n- 21 right_hand  \nEach landmark has 3 coordinates, i.e. **x**, **y**, **z**","metadata":{}},{"cell_type":"code","source":"file_id = 5414471\nsequence_id = 1816796431\nsign = pd.read_parquet(f\"/kaggle/input/asl-fingerspelling/train_landmarks/{file_id}.parquet\")\nsequence = sign[sign.index == sequence_id]\nsequence_phrase = get_phrase(train_df, file_id, sequence_id)\nvisualise2d_landmarks(sequence, f\"Phrase: {sequence_phrase}\")","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# WIP\n### TODO:\n- better insights on train data\n- better viz on landmarks\n- analyze supplimental landmarks","metadata":{}}]}