{
  "id": 363715,
  "title": "Insufficient RAM for test set",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363715",
  "author_name": "Ran Gladstone",
  "post_date": "2022-11-02T21:10:12.963000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I've looked at some notebooks posted here to see how to read .hdf5 files which were very helpful and made an initial model that worked on the files in the train folder and seemed to do ok. However, once I try to predict the labels for the test folder my notebook crashes due to a lack of RAM. There is more than an order of magnitude more test data than train data. I came to realize that I should come up with a way to create small batches of data and feed them to my model instead of creating a Pandas DataFrame for all the data. I cannot directly adopt the available notebooks DataLoader approach because I am using TensorFlow.The only way I currently know to do that in TensorFlow is through generators, but it seems clunky.</p>\n<p>I have seen custom generators for .hdf5 images available, but since the data for the competition doesn't have a typical image format I suspect they will not work. I have tried writing my own little generator but I have very little experience with such programming at the moment and so far it keeps giving me errors.</p>\n<p>Has anyone run into this issue of creating batches of data and labels from the .hdf5 files in TensorFlow?</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": 2014797,
      "postDate": "2022-11-02T21:10:12.963Z",
      "content": "<p>I've looked at some notebooks posted here to see how to read .hdf5 files which were very helpful and made an initial model that worked on the files in the train folder and seemed to do ok. However, once I try to predict the labels for the test folder my notebook crashes due to a lack of RAM. There is more than an order of magnitude more test data than train data. I came to realize that I should come up with a way to create small batches of data and feed them to my model instead of creating a Pandas DataFrame for all the data. I cannot directly adopt the available notebooks DataLoader approach because I am using TensorFlow.The only way I currently know to do that in TensorFlow is through generators, but it seems clunky.</p>\n<p>I have seen custom generators for .hdf5 images available, but since the data for the competition doesn't have a typical image format I suspect they will not work. I have tried writing my own little generator but I have very little experience with such programming at the moment and so far it keeps giving me errors.</p>\n<p>Has anyone run into this issue of creating batches of data and labels from the .hdf5 files in TensorFlow?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "I've looked at some notebooks posted here to see how to read .hdf5 files which were very helpful and made an initial model that worked on the files in the train folder and seemed to do ok. However, once I try to predict the labels for the test folder my notebook crashes due to a lack of RAM. There is more than an order of magnitude more test data than train data. I came to realize that I should come up with a way to create small batches of data and feed them to my model instead of creating a Pandas DataFrame for all the data. I cannot directly adopt the available notebooks DataLoader approach because I am using TensorFlow.The only way I currently know to do that in TensorFlow is through generators, but it seems clunky.\n\nI have seen custom generators for .hdf5 images available, but since the data for the competition doesn't have a typical image format I suspect they will not work. I have tried writing my own little generator but I have very little experience with such programming at the moment and so far it keeps giving me errors.\n\nHas anyone run into this issue of creating batches of data and labels from the .hdf5 files in TensorFlow?\n\nThanks!"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2014797": "I've looked at some notebooks posted here to see how to read .hdf5 files which were very helpful and made an initial model that worked on the files in the train folder and seemed to do ok. However, once I try to predict the labels for the test folder my notebook crashes due to a lack of RAM. There is more than an order of magnitude more test data than train data. I came to realize that I should come up with a way to create small batches of data and feed them to my model instead of creating a Pandas DataFrame for all the data. I cannot directly adopt the available notebooks DataLoader approach because I am using TensorFlow.The only way I currently know to do that in TensorFlow is through generators, but it seems clunky.\n\nI have seen custom generators for .hdf5 images available, but since the data for the competition doesn't have a typical image format I suspect they will not work. I have tried writing my own little generator but I have very little experience with such programming at the moment and so far it keeps giving me errors.\n\nHas anyone run into this issue of creating batches of data and labels from the .hdf5 files in TensorFlow?\n\nThanks!"
  }
}