{
  "id": 421022,
  "title": "Frame length distribution about \"ALL\" sequence",
  "url": "/competitions/asl-fingerspelling/discussion/421022",
  "author_name": "Aurora_blue",
  "post_date": "2023-07-03T15:36:44.372000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>As everyone knows, out of memory error occurs when loading all sequence data.<br>\nI got the info about \"ALL\" frame length by implementing following code repeatedly and concatenating, so will share the data and notebook.<br>\nPlease upvote if you like. Thanks.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/datasets/clearwaterkzk/asl-fingerspelling-eda\" target=\"_blank\">data with frame length</a></li>\n<li><a href=\"https://www.kaggle.com/clearwaterkzk/asl-vis-all-frame-length\" target=\"_blank\">simple eda notebook</a></li>\n</ul>\n<pre><code>sequence_ids = []\nlengths = []\n\n path  tqdm(train.path.unique()):\n    tmp_df = train.loc[train[] == path]\n    pq_data = pd.read_parquet(os.path.join(data_path, path))\n     sequence_id  tmp_df.sequence_id:\n        data = pq_data.loc[sequence_id]\n        lengths.append((data))\n        sequence_ids.append(sequence_id)\n         data; gc.collect()\n     pq_data; gc.collect()\n\n    base_path = os.path.basename(path).split()[]\n     (, mode=)  fo:\n        pickle.dump(sequence_ids, fo)\n\n     (, mode=)  fo:\n        pickle.dump(lengths, fo)\n</code></pre>",
  "messages": [
    {
      "id": 2328461,
      "postDate": "2023-07-03T15:36:44.373Z",
      "content": "<p>As everyone knows, out of memory error occurs when loading all sequence data.<br>\nI got the info about \"ALL\" frame length by implementing following code repeatedly and concatenating, so will share the data and notebook.<br>\nPlease upvote if you like. Thanks.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/datasets/clearwaterkzk/asl-fingerspelling-eda\" target=\"_blank\">data with frame length</a></li>\n<li><a href=\"https://www.kaggle.com/clearwaterkzk/asl-vis-all-frame-length\" target=\"_blank\">simple eda notebook</a></li>\n</ul>\n<pre><code>sequence_ids = []\nlengths = []\n\n path  tqdm(train.path.unique()):\n    tmp_df = train.loc[train[] == path]\n    pq_data = pd.read_parquet(os.path.join(data_path, path))\n     sequence_id  tmp_df.sequence_id:\n        data = pq_data.loc[sequence_id]\n        lengths.append((data))\n        sequence_ids.append(sequence_id)\n         data; gc.collect()\n     pq_data; gc.collect()\n\n    base_path = os.path.basename(path).split()[]\n     (, mode=)  fo:\n        pickle.dump(sequence_ids, fo)\n\n     (, mode=)  fo:\n        pickle.dump(lengths, fo)\n</code></pre>",
      "rawMarkdown": "As everyone knows, out of memory error occurs when loading all sequence data.\nI got the info about \"ALL\" frame length by implementing following code repeatedly and concatenating, so will share the data and notebook.\nPlease upvote if you like. Thanks.\n\n- [data with frame length](https://www.kaggle.com/datasets/clearwaterkzk/asl-fingerspelling-eda)\n- [simple eda notebook](https://www.kaggle.com/clearwaterkzk/asl-vis-all-frame-length)\n\n\n```python\nsequence_ids = []\nlengths = []\n\nfor path in tqdm(train.path.unique()):\n    tmp_df = train.loc[train[\"path\"] == path]\n    pq_data = pd.read_parquet(os.path.join(data_path, path))\n    for sequence_id in tmp_df.sequence_id:\n        data = pq_data.loc[sequence_id]\n        lengths.append(len(data))\n        sequence_ids.append(sequence_id)\n        del data; gc.collect()\n    del pq_data; gc.collect()\n\n    base_path = os.path.basename(path).split(\".\")[0]\n    with open(f\"../output/eda/sequence_ids_{base_path}.pickle\", mode=\"wb\") as fo:\n        pickle.dump(sequence_ids, fo)\n\n    with open(f\"../output/eda/lengths_{base_path}.pickle\", mode=\"wb\") as fo:\n        pickle.dump(lengths, fo)\n\n```"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2328461": "As everyone knows, out of memory error occurs when loading all sequence data.\nI got the info about \"ALL\" frame length by implementing following code repeatedly and concatenating, so will share the data and notebook.\nPlease upvote if you like. Thanks.\n\n- [data with frame length](https://www.kaggle.com/datasets/clearwaterkzk/asl-fingerspelling-eda)\n- [simple eda notebook](https://www.kaggle.com/clearwaterkzk/asl-vis-all-frame-length)\n\n\n```python\nsequence_ids = []\nlengths = []\n\nfor path in tqdm(train.path.unique()):\n    tmp_df = train.loc[train[\"path\"] == path]\n    pq_data = pd.read_parquet(os.path.join(data_path, path))\n    for sequence_id in tmp_df.sequence_id:\n        data = pq_data.loc[sequence_id]\n        lengths.append(len(data))\n        sequence_ids.append(sequence_id)\n        del data; gc.collect()\n    del pq_data; gc.collect()\n\n    base_path = os.path.basename(path).split(\".\")[0]\n    with open(f\"../output/eda/sequence_ids_{base_path}.pickle\", mode=\"wb\") as fo:\n        pickle.dump(sequence_ids, fo)\n\n    with open(f\"../output/eda/lengths_{base_path}.pickle\", mode=\"wb\") as fo:\n        pickle.dump(lengths, fo)\n\n```"
  }
}