{
  "id": 72334,
  "title": "validation set memory error",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/72334",
  "author_name": "MarcYueZhao",
  "post_date": "2018-11-22T10:34:28.064000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I used 80 per class as validation set (27200 total). When I used (256, 256, 3) image, the validation set memory exceeded 32G. Any one know how to deal with this when you use 256 * 256? Thanks a lot.</p>",
  "messages": [
    {
      "id": 426120,
      "postDate": "2018-11-22T16:49:51.327Z",
      "content": "<p>I'm learning keras, with this keras sequence,  <a href=\"https://keras.io/utils/#sequence\">https://keras.io/utils/#sequence</a>, I can  fit/evaluate/predict by batch, load data once a batch. So we can avoid memory error.</p>",
      "rawMarkdown": "I'm learning keras, with this keras sequence,  https://keras.io/utils/#sequence, I can  fit/evaluate/predict by batch, load data once a batch. So we can avoid memory error.",
      "votes": 1
    },
    {
      "id": 425934,
      "postDate": "2018-11-22T10:34:28.063Z",
      "content": "<p>I used 80 per class as validation set (27200 total). When I used (256, 256, 3) image, the validation set memory exceeded 32G. Any one know how to deal with this when you use 256 * 256? Thanks a lot.</p>",
      "rawMarkdown": "I used 80 per class as validation set (27200 total). When I used (256, 256, 3) image, the validation set memory exceeded 32G. Any one know how to deal with this when you use 256 * 256? Thanks a lot.",
      "votes": 1
    },
    {
      "id": 425943,
      "postDate": "2018-11-22T10:46:33.800Z",
      "content": "<p>use generator to generate validation data, the same as training data, see the following code:</p>\n\n<p>```python\ndef training_data_generator():\n    while True:\n        for k in np.random.permutation(range(NUM_FOLDS-1)): # the last fold is validation set\n            file_path = os.path.join(OUTPUT_DIR, 'k_fold', 'train_k{0:02d}.csv.gz'.format(k))\n            for df in pd.read_csv(file_path, chunksize=BATCH_SIZE):\n                X = df_to_image_array_xd(df, WIDTH, IS_RGB)\n                Y = keras.utils.to_categorical(df.y, num_classes=C)\n                yield X, Y</p>\n\n<p>def validation_data_generator():\n    while True:\n        file_path = os.path.join(OUTPUT_DIR, 'k_fold', 'train_k{0:02d}.csv.gz'.format(NUM_FOLDS-1))\n        for df in pd.read_csv(file_path, chunksize=BATCH_SIZE):\n            X = df_to_image_array_xd(df, WIDTH, IS_RGB)\n            Y = keras.utils.to_categorical(df.y, num_classes=C)\n            yield X, Y\n```</p>",
      "rawMarkdown": "use generator to generate validation data, the same as training data, see the following code:\n\n```python\ndef training_data_generator():\n    while True:\n        for k in np.random.permutation(range(NUM_FOLDS-1)): # the last fold is validation set\n            file_path = os.path.join(OUTPUT_DIR, 'k_fold', 'train_k{0:02d}.csv.gz'.format(k))\n            for df in pd.read_csv(file_path, chunksize=BATCH_SIZE):\n                X = df_to_image_array_xd(df, WIDTH, IS_RGB)\n                Y = keras.utils.to_categorical(df.y, num_classes=C)\n                yield X, Y\n\ndef validation_data_generator():\n    while True:\n        file_path = os.path.join(OUTPUT_DIR, 'k_fold', 'train_k{0:02d}.csv.gz'.format(NUM_FOLDS-1))\n        for df in pd.read_csv(file_path, chunksize=BATCH_SIZE):\n            X = df_to_image_array_xd(df, WIDTH, IS_RGB)\n            Y = keras.utils.to_categorical(df.y, num_classes=C)\n            yield X, Y\n```",
      "votes": 2,
      "replies": [
        {
          "id": 425996,
          "postDate": "2018-11-22T12:14:41.387Z",
          "content": "<p>Thx!</p>",
          "rawMarkdown": "Thx!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 426120,
      "author_name": "yyqing",
      "author_url": "",
      "post_date": "2018-11-22T16:49:51.327000",
      "content": "<p>I'm learning keras, with this keras sequence,  <a href=\"https://keras.io/utils/#sequence\">https://keras.io/utils/#sequence</a>, I can  fit/evaluate/predict by batch, load data once a batch. So we can avoid memory error.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 425943,
      "author_name": "[he.ai]soulmachine",
      "author_url": "",
      "post_date": "2018-11-22T10:46:33.800000",
      "content": "<p>use generator to generate validation data, the same as training data, see the following code:</p>\n\n<p>```python\ndef training_data_generator():\n    while True:\n        for k in np.random.permutation(range(NUM_FOLDS-1)): # the last fold is validation set\n            file_path = os.path.join(OUTPUT_DIR, 'k_fold', 'train_k{0:02d}.csv.gz'.format(k))\n            for df in pd.read_csv(file_path, chunksize=BATCH_SIZE):\n                X = df_to_image_array_xd(df, WIDTH, IS_RGB)\n                Y = keras.utils.to_categorical(df.y, num_classes=C)\n                yield X, Y</p>\n\n<p>def validation_data_generator():\n    while True:\n        file_path = os.path.join(OUTPUT_DIR, 'k_fold', 'train_k{0:02d}.csv.gz'.format(NUM_FOLDS-1))\n        for df in pd.read_csv(file_path, chunksize=BATCH_SIZE):\n            X = df_to_image_array_xd(df, WIDTH, IS_RGB)\n            Y = keras.utils.to_categorical(df.y, num_classes=C)\n            yield X, Y\n```</p>",
      "votes": 2,
      "replies": [
        {
          "id": 425996,
          "author_name": "MarcYueZhao",
          "author_url": "",
          "post_date": "2018-11-22T12:14:41.387000",
          "content": "<p>Thx!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "426120": "I'm learning keras, with this keras sequence,  https://keras.io/utils/#sequence, I can  fit/evaluate/predict by batch, load data once a batch. So we can avoid memory error.",
    "425934": "I used 80 per class as validation set (27200 total). When I used (256, 256, 3) image, the validation set memory exceeded 32G. Any one know how to deal with this when you use 256 * 256? Thanks a lot.",
    "425943": "use generator to generate validation data, the same as training data, see the following code:\n\n```python\ndef training_data_generator():\n    while True:\n        for k in np.random.permutation(range(NUM_FOLDS-1)): # the last fold is validation set\n            file_path = os.path.join(OUTPUT_DIR, 'k_fold', 'train_k{0:02d}.csv.gz'.format(k))\n            for df in pd.read_csv(file_path, chunksize=BATCH_SIZE):\n                X = df_to_image_array_xd(df, WIDTH, IS_RGB)\n                Y = keras.utils.to_categorical(df.y, num_classes=C)\n                yield X, Y\n\ndef validation_data_generator():\n    while True:\n        file_path = os.path.join(OUTPUT_DIR, 'k_fold', 'train_k{0:02d}.csv.gz'.format(NUM_FOLDS-1))\n        for df in pd.read_csv(file_path, chunksize=BATCH_SIZE):\n            X = df_to_image_array_xd(df, WIDTH, IS_RGB)\n            Y = keras.utils.to_categorical(df.y, num_classes=C)\n            yield X, Y\n```"
  }
}