{
  "id": 413104,
  "title": "TFRecords Format for Parquet Files",
  "url": "/competitions/asl-fingerspelling/discussion/413104",
  "author_name": "Miltiades general",
  "post_date": "2023-05-26T22:13:51.524000",
  "votes": 0,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Does anyone have knowledge of datasets containing the landmarks in .tfrecords format for this competition? </p>",
  "messages": [
    {
      "id": 2278972,
      "postDate": "2023-05-29T05:59:34.593Z",
      "content": "<p>You can use something like this to create tfrecord files</p>\n<pre><code> def load_relevant_data_subset(pq_path):\n    return pd.read_parquet(pq_path, columns=SEL_COLS)\n\nfor file_id in tqdm(df.file_id.unique()):\n    pqfile = f\"{inpdir}/train_landmarks/{file_id}.parquet\"\n    if not os.path.isdir(\"tfds\"): os.mkdir(\"tfds\")\n    tffile = f\"tfds/{file_id}.tfrecord\"\n    seq_refs = df.loc[df.file_id == file_id]\n    seqs = load_relevant_data_subset(pqfile)\n\n    with tf.io.TFRecordWriter(tffile) as file_writer:\n        for seq_id, phrase in zip(seq_refs.sequence_id, seq_refs.phrase_bytes):\n            frames = seqs.iloc[seqs.index == seq_id]\n            frames128 = frames.fillna(0).to_numpy()\n            frames128 = resize(frames128, (FRAME_LEN, len(SEL_COLS)))\n            frames = pd.DataFrame(data = frames128, columns=frames.columns)\n\n            features = {COL: tf.train.Feature(float_list=tf.train.FloatList(value=frames[COL])) for COL in SEL_COLS}\n            features[\"phrase\"] = tf.train.Feature(bytes_list=tf.train.BytesList(value=[phrase]))\n            record_bytes = tf.train.Example(features=tf.train.Features(feature=features)).SerializeToString()\n            file_writer.write(record_bytes)\n</code></pre>",
      "rawMarkdown": "You can use something like this to create tfrecord files\n\n     def load_relevant_data_subset(pq_path):\n        return pd.read_parquet(pq_path, columns=SEL_COLS)\n\n    for file_id in tqdm(df.file_id.unique()):\n        pqfile = f\"{inpdir}/train_landmarks/{file_id}.parquet\"\n        if not os.path.isdir(\"tfds\"): os.mkdir(\"tfds\")\n        tffile = f\"tfds/{file_id}.tfrecord\"\n        seq_refs = df.loc[df.file_id == file_id]\n        seqs = load_relevant_data_subset(pqfile)\n    \n        with tf.io.TFRecordWriter(tffile) as file_writer:\n            for seq_id, phrase in zip(seq_refs.sequence_id, seq_refs.phrase_bytes):\n                frames = seqs.iloc[seqs.index == seq_id]\n                frames128 = frames.fillna(0).to_numpy()\n                frames128 = resize(frames128, (FRAME_LEN, len(SEL_COLS)))\n                frames = pd.DataFrame(data = frames128, columns=frames.columns)\n            \n                features = {COL: tf.train.Feature(float_list=tf.train.FloatList(value=frames[COL])) for COL in SEL_COLS}\n                features[\"phrase\"] = tf.train.Feature(bytes_list=tf.train.BytesList(value=[phrase]))\n                record_bytes = tf.train.Example(features=tf.train.Features(feature=features)).SerializeToString()\n                file_writer.write(record_bytes)\n",
      "votes": 1
    },
    {
      "id": 2276553,
      "postDate": "2023-05-27T01:27:11.293Z",
      "content": "<p>That's not something provided by the host. You can create by yourself.<br>\nThis notebook may help.<br>\n<a href=\"https://www.kaggle.com/code/ryanholbrook/tfrecords-basics\" target=\"_blank\">https://www.kaggle.com/code/ryanholbrook/tfrecords-basics</a></p>",
      "rawMarkdown": "That's not something provided by the host. You can create by yourself.\nThis notebook may help.\nhttps://www.kaggle.com/code/ryanholbrook/tfrecords-basics"
    },
    {
      "id": 2276494,
      "postDate": "2023-05-26T22:13:51.523Z",
      "content": "<p>Does anyone have knowledge of datasets containing the landmarks in .tfrecords format for this competition? </p>",
      "rawMarkdown": "Does anyone have knowledge of datasets containing the landmarks in .tfrecords format for this competition? "
    }
  ],
  "comments": [
    {
      "id": 2278972,
      "author_name": "Rohith Ingilela",
      "author_url": "",
      "post_date": "2023-05-29T05:59:34.593000",
      "content": "<p>You can use something like this to create tfrecord files</p>\n<pre><code> def load_relevant_data_subset(pq_path):\n    return pd.read_parquet(pq_path, columns=SEL_COLS)\n\nfor file_id in tqdm(df.file_id.unique()):\n    pqfile = f\"{inpdir}/train_landmarks/{file_id}.parquet\"\n    if not os.path.isdir(\"tfds\"): os.mkdir(\"tfds\")\n    tffile = f\"tfds/{file_id}.tfrecord\"\n    seq_refs = df.loc[df.file_id == file_id]\n    seqs = load_relevant_data_subset(pqfile)\n\n    with tf.io.TFRecordWriter(tffile) as file_writer:\n        for seq_id, phrase in zip(seq_refs.sequence_id, seq_refs.phrase_bytes):\n            frames = seqs.iloc[seqs.index == seq_id]\n            frames128 = frames.fillna(0).to_numpy()\n            frames128 = resize(frames128, (FRAME_LEN, len(SEL_COLS)))\n            frames = pd.DataFrame(data = frames128, columns=frames.columns)\n\n            features = {COL: tf.train.Feature(float_list=tf.train.FloatList(value=frames[COL])) for COL in SEL_COLS}\n            features[\"phrase\"] = tf.train.Feature(bytes_list=tf.train.BytesList(value=[phrase]))\n            record_bytes = tf.train.Example(features=tf.train.Features(feature=features)).SerializeToString()\n            file_writer.write(record_bytes)\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2276553,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2023-05-27T01:27:11.293000",
      "content": "<p>That's not something provided by the host. You can create by yourself.<br>\nThis notebook may help.<br>\n<a href=\"https://www.kaggle.com/code/ryanholbrook/tfrecords-basics\" target=\"_blank\">https://www.kaggle.com/code/ryanholbrook/tfrecords-basics</a></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2278972": "You can use something like this to create tfrecord files\n\n     def load_relevant_data_subset(pq_path):\n        return pd.read_parquet(pq_path, columns=SEL_COLS)\n\n    for file_id in tqdm(df.file_id.unique()):\n        pqfile = f\"{inpdir}/train_landmarks/{file_id}.parquet\"\n        if not os.path.isdir(\"tfds\"): os.mkdir(\"tfds\")\n        tffile = f\"tfds/{file_id}.tfrecord\"\n        seq_refs = df.loc[df.file_id == file_id]\n        seqs = load_relevant_data_subset(pqfile)\n    \n        with tf.io.TFRecordWriter(tffile) as file_writer:\n            for seq_id, phrase in zip(seq_refs.sequence_id, seq_refs.phrase_bytes):\n                frames = seqs.iloc[seqs.index == seq_id]\n                frames128 = frames.fillna(0).to_numpy()\n                frames128 = resize(frames128, (FRAME_LEN, len(SEL_COLS)))\n                frames = pd.DataFrame(data = frames128, columns=frames.columns)\n            \n                features = {COL: tf.train.Feature(float_list=tf.train.FloatList(value=frames[COL])) for COL in SEL_COLS}\n                features[\"phrase\"] = tf.train.Feature(bytes_list=tf.train.BytesList(value=[phrase]))\n                record_bytes = tf.train.Example(features=tf.train.Features(feature=features)).SerializeToString()\n                file_writer.write(record_bytes)\n",
    "2276553": "That's not something provided by the host. You can create by yourself.\nThis notebook may help.\nhttps://www.kaggle.com/code/ryanholbrook/tfrecords-basics",
    "2276494": "Does anyone have knowledge of datasets containing the landmarks in .tfrecords format for this competition? "
  }
}