{
  "id": 359657,
  "title": "Suggestions on training on the dataset?",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/359657",
  "author_name": "Jacob Dawson",
  "post_date": "2022-10-13T02:09:44.761000",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/code/chazzer/how-to-read-the-hdf5-files\" target=\"_blank\">https://www.kaggle.com/code/chazzer/how-to-read-the-hdf5-files</a> has provided some really helpful info on how to extract numpy arrays from the hdf5 format. I tried making my setup and I ran into a problem: how do I bunch these up into a dataset for processing? One way that I've found is that I can process each hdf5 one-by-one like so:</p>\n<p><code>train_dataset = tf.data.Dataset.list_files(train_path+\"*.hdf5\")\nprint(train_dataset.cardinality())\nprint(train_dataset.take(1).get_single_element())</code></p>\n<p>where that last line outputs:<br>\n<code>tf.Tensor(b'..\\\\g2net-detecting-continuous-gravitational-waves\\\\train\\\\123594dc7.hdf5', shape=(), dtype=string)</code></p>\n<p>I intend to read the hdf5 files one-by-one from these strings using h5py.File(), but I was curious--is there a better way to do this? Would it be better to create a generator/dataset of hdf5 files themselves which can then be retrieved (obviously we can't store all 200 gb in RAM!)?</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": 1984953,
      "postDate": "2022-10-13T02:09:44.760Z",
      "content": "<p><a href=\"https://www.kaggle.com/code/chazzer/how-to-read-the-hdf5-files\" target=\"_blank\">https://www.kaggle.com/code/chazzer/how-to-read-the-hdf5-files</a> has provided some really helpful info on how to extract numpy arrays from the hdf5 format. I tried making my setup and I ran into a problem: how do I bunch these up into a dataset for processing? One way that I've found is that I can process each hdf5 one-by-one like so:</p>\n<p><code>train_dataset = tf.data.Dataset.list_files(train_path+\"*.hdf5\")\nprint(train_dataset.cardinality())\nprint(train_dataset.take(1).get_single_element())</code></p>\n<p>where that last line outputs:<br>\n<code>tf.Tensor(b'..\\\\g2net-detecting-continuous-gravitational-waves\\\\train\\\\123594dc7.hdf5', shape=(), dtype=string)</code></p>\n<p>I intend to read the hdf5 files one-by-one from these strings using h5py.File(), but I was curious--is there a better way to do this? Would it be better to create a generator/dataset of hdf5 files themselves which can then be retrieved (obviously we can't store all 200 gb in RAM!)?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "https://www.kaggle.com/code/chazzer/how-to-read-the-hdf5-files has provided some really helpful info on how to extract numpy arrays from the hdf5 format. I tried making my setup and I ran into a problem: how do I bunch these up into a dataset for processing? One way that I've found is that I can process each hdf5 one-by-one like so:\n\n`train_dataset = tf.data.Dataset.list_files(train_path+\"*.hdf5\")\nprint(train_dataset.cardinality())\nprint(train_dataset.take(1).get_single_element())`\n\nwhere that last line outputs:\n`tf.Tensor(b'..\\\\g2net-detecting-continuous-gravitational-waves\\\\train\\\\123594dc7.hdf5', shape=(), dtype=string)`\n\nI intend to read the hdf5 files one-by-one from these strings using h5py.File(), but I was curious--is there a better way to do this? Would it be better to create a generator/dataset of hdf5 files themselves which can then be retrieved (obviously we can't store all 200 gb in RAM!)?\n\nThanks!",
      "votes": 6
    },
    {
      "id": 1985969,
      "postDate": "2022-10-13T17:58:54.943Z",
      "content": "<p>What do you want to do? hdf5 is pretty quick to read. </p>",
      "rawMarkdown": "What do you want to do? hdf5 is pretty quick to read. \n",
      "replies": [
        {
          "id": 1987892,
          "postDate": "2022-10-14T21:58:56.877Z",
          "content": "<p>Perhaps I didn't phrase my thoughts clearly enough. I can see how I'd parse through each hdf5 file individually and feed them to my model, but I don't see how I could make a dataset out of them so that I can run a batch in parallel. Is there a good way to do this in Tensorflow that I'm missing?</p>",
          "rawMarkdown": "Perhaps I didn't phrase my thoughts clearly enough. I can see how I'd parse through each hdf5 file individually and feed them to my model, but I don't see how I could make a dataset out of them so that I can run a batch in parallel. Is there a good way to do this in Tensorflow that I'm missing?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1985969,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2022-10-13T17:58:54.943000",
      "content": "<p>What do you want to do? hdf5 is pretty quick to read. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1987892,
          "author_name": "Jacob Dawson",
          "author_url": "",
          "post_date": "2022-10-14T21:58:56.877000",
          "content": "<p>Perhaps I didn't phrase my thoughts clearly enough. I can see how I'd parse through each hdf5 file individually and feed them to my model, but I don't see how I could make a dataset out of them so that I can run a batch in parallel. Is there a good way to do this in Tensorflow that I'm missing?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1984953": "https://www.kaggle.com/code/chazzer/how-to-read-the-hdf5-files has provided some really helpful info on how to extract numpy arrays from the hdf5 format. I tried making my setup and I ran into a problem: how do I bunch these up into a dataset for processing? One way that I've found is that I can process each hdf5 one-by-one like so:\n\n`train_dataset = tf.data.Dataset.list_files(train_path+\"*.hdf5\")\nprint(train_dataset.cardinality())\nprint(train_dataset.take(1).get_single_element())`\n\nwhere that last line outputs:\n`tf.Tensor(b'..\\\\g2net-detecting-continuous-gravitational-waves\\\\train\\\\123594dc7.hdf5', shape=(), dtype=string)`\n\nI intend to read the hdf5 files one-by-one from these strings using h5py.File(), but I was curious--is there a better way to do this? Would it be better to create a generator/dataset of hdf5 files themselves which can then be retrieved (obviously we can't store all 200 gb in RAM!)?\n\nThanks!",
    "1985969": "What do you want to do? hdf5 is pretty quick to read. \n"
  }
}