{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Files\n[train/supplemental_metadata].csv\n\npath - The path to the landmark file.\nfile_id - A unique identifier for the data file.\nparticipant_id - A unique identifier for the data contributor.\nsequence_id - A unique identifier for the landmark sequence. Each data file may contain many sequences.\nphrase - The labels for the landmark sequence. The train and test datasets contain randomly generated addresses, phone numbers, and urls derived from components of real addresses/phone numbers/urls. Any overlap with real addresses, phone numbers, or urls is purely accidental. The supplemental dataset consists of fingerspelled sentences. Note that some of the urls include adult content. The intent of this competition is to support the Deaf and Hard of Hearing community in engaging with technology on an equal footing with other adults.\ncharacter_to_prediction_index.json\n\n[train/supplemental]_landmarks/ The landmark data. The landmarks were extracted from raw videos with the MediaPipe holistic model. Not all of the frames necessarily had visible hands or hands that could be detected by the model.\nThe landmark files contain the same data as in the ASL Signs competition (minus the row ID column) but reshaped into a wide format. This allows you to take advantage of the Parquet format to entirely skip loading landmarks that you aren't using.\n\nsequence_id - A unique identifier for the landmark sequence. Most landmark filestrain/supplemental_metadata contain 1,000 sequences. The sequence ID is used as the dataframe index.\nframe - The frame number within a landmark sequence.\n[x/y/z]_[type]_[landmark_index] - There are now 1,629 spatial coordinate columns for the x, y and z coordinates for each of the 543 landmarks. The type of landmark is one of ['face', 'left_hand', 'pose', 'right_hand']. Details of the hand landmark locations can be found here. The spatial coordinates have already been normalized by MediaPipe. Note that the MediaPipe model is not fully trained to predict depth so you may wish to ignore the z values. The landmarks have been converted to float32.\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# Import all dependencies","metadata":{}},{"cell_type":"code","source":"\nimport pandas as pd,numpy as np,os\nimport json\nimport plotly.graph_objects as go\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nimport plotly.io as pio\nfrom pathlib import Path","metadata":{"execution":{"iopub.status.busy":"2023-05-12T20:29:29.130385Z","iopub.execute_input":"2023-05-12T20:29:29.130864Z","iopub.status.idle":"2023-05-12T20:29:30.725115Z","shell.execute_reply.started":"2023-05-12T20:29:29.130819Z","shell.execute_reply":"2023-05-12T20:29:30.723907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA","metadata":{}},{"cell_type":"code","source":"\n\n# Load dataset metadata\nmetadata = pd.read_csv('/kaggle/input/asl-fingerspelling/supplemental_metadata.csv')\n\n# Print the dimensions of the metadata dataset\nprint(\"Metadata dimensions: \", metadata.shape)\n\n","metadata":{"execution":{"iopub.status.busy":"2023-05-12T20:28:22.308609Z","iopub.execute_input":"2023-05-12T20:28:22.308982Z","iopub.status.idle":"2023-05-12T20:28:22.449784Z","shell.execute_reply.started":"2023-05-12T20:28:22.308953Z","shell.execute_reply":"2023-05-12T20:28:22.448651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\n\n# Load dataset landmark\nlandmarks = pd.read_parquet('/kaggle/input/asl-fingerspelling/supplemental_landmarks/33432165.parquet')\n\n# Print the first 5 rows of the dataset\nprint(landmarks.head())\n","metadata":{"execution":{"iopub.status.busy":"2023-05-12T20:27:57.66494Z","iopub.execute_input":"2023-05-12T20:27:57.665985Z","iopub.status.idle":"2023-05-12T20:28:18.293455Z","shell.execute_reply.started":"2023-05-12T20:27:57.66594Z","shell.execute_reply":"2023-05-12T20:28:18.292589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot a histogram of the x coordinates of the first hand landmark\nplt.hist(landmarks['x_left_hand_0'], bins=50)\nplt.xlabel('x coordinate')\nplt.ylabel('count')\nplt.title('Histogram of x coordinate of first left hand landmark')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-12T20:29:39.883064Z","iopub.execute_input":"2023-05-12T20:29:39.883515Z","iopub.status.idle":"2023-05-12T20:29:40.41708Z","shell.execute_reply.started":"2023-05-12T20:29:39.88348Z","shell.execute_reply":"2023-05-12T20:29:40.415939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport pyarrow.parquet as pq\nimport matplotlib.pyplot as plt\n\n# Load metadata\n# Load landmark data in parquet format\nbase_dir=Path(\"/kaggle/input/asl-fingerspelling\")\nlandmarks = pq.read_table(base_dir).to_pandas()\n\n# Merge metadata and landmark data based on sequence_id\ndata = pd.merge(metadata, landmarks, on=\"sequence_id\")\n\n# Plot x and y coordinates for left_hand landmarks\nleft_hand_x = data.filter(regex=\"^left_hand_x\")\nleft_hand_y = data.filter(regex=\"^left_hand_y\")\nplt.plot(left_hand_x.values.T, -left_hand_y.values.T, alpha=0.1)\n\n# Add title and labels\nplt.title(\"Left hand landmarks\")\nplt.xlabel(\"X coordinate\")\nplt.ylabel(\"Y coordinate\")\n\n# Show plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-05-12T20:10:00.797386Z","iopub.execute_input":"2023-05-12T20:10:00.798336Z","iopub.status.idle":"2023-05-12T20:10:01.166656Z","shell.execute_reply.started":"2023-05-12T20:10:00.798293Z","shell.execute_reply":"2023-05-12T20:10:01.165163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# TRAIN","metadata":{}}]}