{
  "id": 410214,
  "title": "The ratio of frames where the coordinates of both hands are NaN out of the total frames.",
  "url": "/competitions/asl-fingerspelling/discussion/410214",
  "author_name": "canlion",
  "post_date": "2023-05-14T11:37:43.614000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I plan to use only the landmarks of both hands. First, I checked the hand landmark coordinates of all parquet files. Out of a total of 19,865,410 frames, there are 7,992,066 frames where both hands are NaN, which is about 40% of the total. I'm doubtful if I calculated correctly.</p>\n<p>my code:</p>\n<ol>\n<li>parquet to numpy array: load parquet -&gt; extract sequence -&gt; assign numpy array</li>\n</ol>\n<pre><code>columns = [, , , ...]\nseq_arr = np.zeros((, ), dtype=np.float32)\n\n pq_path  uniq_pq_path:  \n    pq = pd.read_parquet(os.path.join(INPUT_DIR, pq_path), columns=columns)\n    pq_train_csv = train_csv[train_csv.path == pq_path]  \n\n     row  pq_train_csv[[, ]].itertuples():\n        seq_id, phrase = row.sequence_id, row.phrase\n        seq = pq[pq.index == seq_id].values\n        sorted_idx = np.argsort(seq[:, ])  \n        sorted_seq = seq[sorted_idx]\n        seq_len = (sorted_seq)\n        seq_arr[start_idx: start_idx + seq_len] = sorted_seq[:, :]  \n</code></pre>\n<ol>\n<li>all nan frame ratio</li>\n</ol>\n<pre><code>all_nan_rows = np.(np.isnan(seq_arr), axis=-).()\n\nall_nan_row_ratio = all_nan_rows / (seq_arr)\n()\n</code></pre>\n<p><strong>all nan row ratio: 40.23106495159174 % (7992066/19865410)</strong></p>",
  "messages": [
    {
      "id": 2258648,
      "postDate": "2023-05-14T11:37:43.613Z",
      "content": "<p>I plan to use only the landmarks of both hands. First, I checked the hand landmark coordinates of all parquet files. Out of a total of 19,865,410 frames, there are 7,992,066 frames where both hands are NaN, which is about 40% of the total. I'm doubtful if I calculated correctly.</p>\n<p>my code:</p>\n<ol>\n<li>parquet to numpy array: load parquet -&gt; extract sequence -&gt; assign numpy array</li>\n</ol>\n<pre><code>columns = [, , , ...]\nseq_arr = np.zeros((, ), dtype=np.float32)\n\n pq_path  uniq_pq_path:  \n    pq = pd.read_parquet(os.path.join(INPUT_DIR, pq_path), columns=columns)\n    pq_train_csv = train_csv[train_csv.path == pq_path]  \n\n     row  pq_train_csv[[, ]].itertuples():\n        seq_id, phrase = row.sequence_id, row.phrase\n        seq = pq[pq.index == seq_id].values\n        sorted_idx = np.argsort(seq[:, ])  \n        sorted_seq = seq[sorted_idx]\n        seq_len = (sorted_seq)\n        seq_arr[start_idx: start_idx + seq_len] = sorted_seq[:, :]  \n</code></pre>\n<ol>\n<li>all nan frame ratio</li>\n</ol>\n<pre><code>all_nan_rows = np.(np.isnan(seq_arr), axis=-).()\n\nall_nan_row_ratio = all_nan_rows / (seq_arr)\n()\n</code></pre>\n<p><strong>all nan row ratio: 40.23106495159174 % (7992066/19865410)</strong></p>",
      "rawMarkdown": "I plan to use only the landmarks of both hands. First, I checked the hand landmark coordinates of all parquet files. Out of a total of 19,865,410 frames, there are 7,992,066 frames where both hands are NaN, which is about 40% of the total. I'm doubtful if I calculated correctly.\n\n\nmy code:\n\n0. parquet to numpy array: load parquet -> extract sequence -> assign numpy array\n\n\n```python\ncolumns = ['sequence_id', 'frame', 'x_left_hand_0', ...]\nseq_arr = np.zeros((19865410, 126), dtype=np.float32)\n\nfor pq_path in uniq_pq_path:  # iter unique parquet paths\n    pq = pd.read_parquet(os.path.join(INPUT_DIR, pq_path), columns=columns)\n    pq_train_csv = train_csv[train_csv.path == pq_path]  # extract corresponding part\n\n    for row in pq_train_csv[['sequence_id', 'phrase']].itertuples():\n        seq_id, phrase = row.sequence_id, row.phrase\n        seq = pq[pq.index == seq_id].values\n        sorted_idx = np.argsort(seq[:, 1])  # sort by frame\n        sorted_seq = seq[sorted_idx]\n        seq_len = len(sorted_seq)\n        seq_arr[start_idx: start_idx + seq_len] = sorted_seq[:, 1:]  # assign sequence to numpy array\n```\n\n\n1. all nan frame ratio\n\n```python\nall_nan_rows = np.all(np.isnan(seq_arr), axis=-1).sum()\n\nall_nan_row_ratio = all_nan_rows / len(seq_arr)\nprint(f'all nan row ratio: {all_nan_row_ratio * 100} % ({all_nan_rows}/{len(seq_arr)})')\n```\n\n**all nan row ratio: 40.23106495159174 % (7992066/19865410)**",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2258648": "I plan to use only the landmarks of both hands. First, I checked the hand landmark coordinates of all parquet files. Out of a total of 19,865,410 frames, there are 7,992,066 frames where both hands are NaN, which is about 40% of the total. I'm doubtful if I calculated correctly.\n\n\nmy code:\n\n0. parquet to numpy array: load parquet -> extract sequence -> assign numpy array\n\n\n```python\ncolumns = ['sequence_id', 'frame', 'x_left_hand_0', ...]\nseq_arr = np.zeros((19865410, 126), dtype=np.float32)\n\nfor pq_path in uniq_pq_path:  # iter unique parquet paths\n    pq = pd.read_parquet(os.path.join(INPUT_DIR, pq_path), columns=columns)\n    pq_train_csv = train_csv[train_csv.path == pq_path]  # extract corresponding part\n\n    for row in pq_train_csv[['sequence_id', 'phrase']].itertuples():\n        seq_id, phrase = row.sequence_id, row.phrase\n        seq = pq[pq.index == seq_id].values\n        sorted_idx = np.argsort(seq[:, 1])  # sort by frame\n        sorted_seq = seq[sorted_idx]\n        seq_len = len(sorted_seq)\n        seq_arr[start_idx: start_idx + seq_len] = sorted_seq[:, 1:]  # assign sequence to numpy array\n```\n\n\n1. all nan frame ratio\n\n```python\nall_nan_rows = np.all(np.isnan(seq_arr), axis=-1).sum()\n\nall_nan_row_ratio = all_nan_rows / len(seq_arr)\nprint(f'all nan row ratio: {all_nan_row_ratio * 100} % ({all_nan_rows}/{len(seq_arr)})')\n```\n\n**all nan row ratio: 40.23106495159174 % (7992066/19865410)**"
  }
}