{
  "id": 74768,
  "title": "Tricks to load the whole simplified dataset into memory",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/74768",
  "author_name": "[he.ai]soulmachine",
  "post_date": "2018-12-15T08:52:08.149000",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Tried a few ways, most of them failed but there is one way succeeded. Assume that you have unzipped the <code>train_simplified.zip</code> into directory <code>train_simplified/</code> which contains 340 csv files. </p>\n\n<p>First, merge 340 csv files into a single csv file and only keep two columns, <code>drawing</code> and <code>y</code>, <code>y</code> is an integer converted from the <code>word</code> column, the file size of <code>train_simplfied.csv</code> is 21GB. Then you can just load the whole dataset by <code>df = pd.read_csv('train_simplified.csv')</code>, which will consume around 25GB RAM.</p>\n\n<p>The following are a few failed ways I've tried:</p>\n\n<ol>\n<li><p>Load the <code>train_simplified.csv</code> using <code>df = pd.read_csv('train_simplified.csv', converters={'drawing': json.loads})</code>, this method ended with consuming up all RAM and failed.</p></li>\n<li><p>Then I optimized the parsing function:</p>\n\n<p><code>python\ndef parse_strokes(s):\nx = json.loads(s)\nreturn list(map(np.uint8, x))\n</code></p>\n\n<p>Because all the values in <code>drawing</code> column don't exceed 256 they can be stored in <code>np.uint8</code>, however <code>df = pd.read_csv('train_simplified.csv', converters={'drawing': parse_strokes})</code> ended with consuming up all RAM and failed again.</p></li>\n<li><p>Merge 340 csv files into a single hdf5 file, only keep two columns <code>drawing</code> and <code>y</code>, and parse the <code>drawing</code> by  either <code>json.loads()</code>  or <code>parse_strokes()</code> . After parsing strings to lists, the file size does reduces to <code>12G</code>, but when I tried to load the dataset into memory, <code>df = pd.read_hdf('train_simplified.hdf5')</code>, the RAM consumption grew over 50GB,  since my machine only has 52GB RAM it failed too.</p></li>\n</ol>\n\n<p>The key takeaway is that <strong>do NOT parse the <code>drawing</code> column in advance, instead you should parse them on the fly during training</strong>. It seems list of integers consume much more memory than plain strings.</p>",
  "messages": [
    {
      "id": 439595,
      "postDate": "2018-12-15T21:55:01.103Z",
      "content": "<p>If you store <code>drawing</code> as a string, there is a pretty dumb trick to save about 20% memory without content loss...</p>\n\n<p>... just remove whitespaces 😁</p>\n\n<p><img src=\"https://i.ibb.co/xhndhMQ/memory-trick.png\" alt=\"memory-trick\"></p>",
      "rawMarkdown": "If you store `drawing` as a string, there is a pretty dumb trick to save about 20% memory without content loss...\n\n... just remove whitespaces 😁\n\n![memory-trick](https://i.ibb.co/xhndhMQ/memory-trick.png)",
      "votes": 1,
      "replies": [
        {
          "id": 439732,
          "postDate": "2018-12-16T08:14:05.617Z",
          "content": "<p>Impressive, thanks for sharing! Is there any way to load all data into memory after parsing <code>drawing</code> from string to list of integers?</p>",
          "rawMarkdown": "Impressive, thanks for sharing! Is there any way to load all data into memory after parsing `drawing` from string to list of integers?",
          "votes": 1
        },
        {
          "id": 439996,
          "postDate": "2018-12-16T20:38:41.900Z",
          "content": "<p>My PC doesn't have even the half memory of yours, so I did not even try to load everything in memory. :)</p>\n\n<p>I came up with the common approach to lazily load data. While my GPU is crunching a batch, workers on CPU are loading and generating the images for the next batch. It keeps my GPU busy at 98%.\nThe way I implemented it, is probably less common: I am using SQLite file for data and a list of image index in memory. I gives some details <a href=\"https://towardsdatascience.com/10-lessons-learned-from-participating-to-google-ai-challenge-268b4aa87efa\">here</a> and source code is on <a href=\"https://github.com/ebouteillon/kaggle-quickdraw-doodle-recognition-challenge\">github</a>.</p>\n\n<p>Curious to know why you want to preload everything. I imagine that you have access to more GPU power than myself. 😁</p>",
          "rawMarkdown": "My PC doesn't have even the half memory of yours, so I did not even try to load everything in memory. :)\n\nI came up with the common approach to lazily load data. While my GPU is crunching a batch, workers on CPU are loading and generating the images for the next batch. It keeps my GPU busy at 98%.\nThe way I implemented it, is probably less common: I am using SQLite file for data and a list of image index in memory. I gives some details [here](https://towardsdatascience.com/10-lessons-learned-from-participating-to-google-ai-challenge-268b4aa87efa) and source code is on [github](https://github.com/ebouteillon/kaggle-quickdraw-doodle-recognition-challenge).\n\nCurious to know why you want to preload everything. I imagine that you have access to more GPU power than myself. 😁"
        },
        {
          "id": 440227,
          "postDate": "2018-12-17T08:54:25.273Z",
          "content": "<p>Currently I'm using a method that is in the middle of  your way and the way that loads all data to memory, I split training data into 100 hdf5 files, every time I load one HDF5 file into memory,  generating batches of images and feed to GPUs, and then load the next HDF5. By doing so I only need to access disk 100 times instead of accessing disk every mini-batch. The <code>fit_generator()</code> didn't enable <code>multiprocessing</code>.</p>\n\n<p>You're correct, I'm going to get a 4-GPU machine, so I'm concerned that CPU will not be fast enough to feed data to GPUs, thus I want to write a data generator that inherit <code>keras.util.Sequence</code> so that I can enable <code>multiprocessing</code> without duplicating data. To write a class inheriting from <code>keras.util.Sequence</code> I feel that this class would be much simpler if it can load the whole data into a single DataFrame.</p>",
          "rawMarkdown": "Currently I'm using a method that is in the middle of  your way and the way that loads all data to memory, I split training data into 100 hdf5 files, every time I load one HDF5 file into memory,  generating batches of images and feed to GPUs, and then load the next HDF5. By doing so I only need to access disk 100 times instead of accessing disk every mini-batch. The `fit_generator()` didn't enable `multiprocessing`.\n\nYou're correct, I'm going to get a 4-GPU machine, so I'm concerned that CPU will not be fast enough to feed data to GPUs, thus I want to write a data generator that inherit `keras.util.Sequence` so that I can enable `multiprocessing` without duplicating data. To write a class inheriting from `keras.util.Sequence` I feel that this class would be much simpler if it can load the whole data into a single DataFrame."
        }
      ]
    },
    {
      "id": 439367,
      "postDate": "2018-12-15T08:52:08.150Z",
      "content": "<p>Tried a few ways, most of them failed but there is one way succeeded. Assume that you have unzipped the <code>train_simplified.zip</code> into directory <code>train_simplified/</code> which contains 340 csv files. </p>\n\n<p>First, merge 340 csv files into a single csv file and only keep two columns, <code>drawing</code> and <code>y</code>, <code>y</code> is an integer converted from the <code>word</code> column, the file size of <code>train_simplfied.csv</code> is 21GB. Then you can just load the whole dataset by <code>df = pd.read_csv('train_simplified.csv')</code>, which will consume around 25GB RAM.</p>\n\n<p>The following are a few failed ways I've tried:</p>\n\n<ol>\n<li><p>Load the <code>train_simplified.csv</code> using <code>df = pd.read_csv('train_simplified.csv', converters={'drawing': json.loads})</code>, this method ended with consuming up all RAM and failed.</p></li>\n<li><p>Then I optimized the parsing function:</p>\n\n<p><code>python\ndef parse_strokes(s):\nx = json.loads(s)\nreturn list(map(np.uint8, x))\n</code></p>\n\n<p>Because all the values in <code>drawing</code> column don't exceed 256 they can be stored in <code>np.uint8</code>, however <code>df = pd.read_csv('train_simplified.csv', converters={'drawing': parse_strokes})</code> ended with consuming up all RAM and failed again.</p></li>\n<li><p>Merge 340 csv files into a single hdf5 file, only keep two columns <code>drawing</code> and <code>y</code>, and parse the <code>drawing</code> by  either <code>json.loads()</code>  or <code>parse_strokes()</code> . After parsing strings to lists, the file size does reduces to <code>12G</code>, but when I tried to load the dataset into memory, <code>df = pd.read_hdf('train_simplified.hdf5')</code>, the RAM consumption grew over 50GB,  since my machine only has 52GB RAM it failed too.</p></li>\n</ol>\n\n<p>The key takeaway is that <strong>do NOT parse the <code>drawing</code> column in advance, instead you should parse them on the fly during training</strong>. It seems list of integers consume much more memory than plain strings.</p>",
      "rawMarkdown": "Tried a few ways, most of them failed but there is one way succeeded. Assume that you have unzipped the `train_simplified.zip` into directory `train_simplified/` which contains 340 csv files. \n\nFirst, merge 340 csv files into a single csv file and only keep two columns, `drawing` and `y`, `y` is an integer converted from the `word` column, the file size of `train_simplfied.csv` is 21GB. Then you can just load the whole dataset by `df = pd.read_csv('train_simplified.csv')`, which will consume around 25GB RAM.\n\nThe following are a few failed ways I've tried:\n\n1. Load the `train_simplified.csv` using `df = pd.read_csv('train_simplified.csv', converters={'drawing': json.loads})`, this method ended with consuming up all RAM and failed.\n\n2. Then I optimized the parsing function:\n\n    ```python\n    def parse_strokes(s):\n    x = json.loads(s)\n    return list(map(np.uint8, x))\n    ```\n\n    Because all the values in `drawing` column don't exceed 256 they can be stored in `np.uint8`, however `df = pd.read_csv('train_simplified.csv', converters={'drawing': parse_strokes})` ended with consuming up all RAM and failed again.\n\n3. Merge 340 csv files into a single hdf5 file, only keep two columns `drawing` and `y`, and parse the `drawing` by  either `json.loads()`  or `parse_strokes()` . After parsing strings to lists, the file size does reduces to `12G`, but when I tried to load the dataset into memory, `df = pd.read_hdf('train_simplified.hdf5')`, the RAM consumption grew over 50GB,  since my machine only has 52GB RAM it failed too.\n\nThe key takeaway is that **do NOT parse the `drawing` column in advance, instead you should parse them on the fly during training**. It seems list of integers consume much more memory than plain strings.\n",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 439595,
      "author_name": "Eric Bouteillon",
      "author_url": "",
      "post_date": "2018-12-15T21:55:01.103000",
      "content": "<p>If you store <code>drawing</code> as a string, there is a pretty dumb trick to save about 20% memory without content loss...</p>\n\n<p>... just remove whitespaces 😁</p>\n\n<p><img src=\"https://i.ibb.co/xhndhMQ/memory-trick.png\" alt=\"memory-trick\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 439732,
          "author_name": "[he.ai]soulmachine",
          "author_url": "",
          "post_date": "2018-12-16T08:14:05.617000",
          "content": "<p>Impressive, thanks for sharing! Is there any way to load all data into memory after parsing <code>drawing</code> from string to list of integers?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 439996,
          "author_name": "Eric Bouteillon",
          "author_url": "",
          "post_date": "2018-12-16T20:38:41.900000",
          "content": "<p>My PC doesn't have even the half memory of yours, so I did not even try to load everything in memory. :)</p>\n\n<p>I came up with the common approach to lazily load data. While my GPU is crunching a batch, workers on CPU are loading and generating the images for the next batch. It keeps my GPU busy at 98%.\nThe way I implemented it, is probably less common: I am using SQLite file for data and a list of image index in memory. I gives some details <a href=\"https://towardsdatascience.com/10-lessons-learned-from-participating-to-google-ai-challenge-268b4aa87efa\">here</a> and source code is on <a href=\"https://github.com/ebouteillon/kaggle-quickdraw-doodle-recognition-challenge\">github</a>.</p>\n\n<p>Curious to know why you want to preload everything. I imagine that you have access to more GPU power than myself. 😁</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 440227,
          "author_name": "[he.ai]soulmachine",
          "author_url": "",
          "post_date": "2018-12-17T08:54:25.273000",
          "content": "<p>Currently I'm using a method that is in the middle of  your way and the way that loads all data to memory, I split training data into 100 hdf5 files, every time I load one HDF5 file into memory,  generating batches of images and feed to GPUs, and then load the next HDF5. By doing so I only need to access disk 100 times instead of accessing disk every mini-batch. The <code>fit_generator()</code> didn't enable <code>multiprocessing</code>.</p>\n\n<p>You're correct, I'm going to get a 4-GPU machine, so I'm concerned that CPU will not be fast enough to feed data to GPUs, thus I want to write a data generator that inherit <code>keras.util.Sequence</code> so that I can enable <code>multiprocessing</code> without duplicating data. To write a class inheriting from <code>keras.util.Sequence</code> I feel that this class would be much simpler if it can load the whole data into a single DataFrame.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "439595": "If you store `drawing` as a string, there is a pretty dumb trick to save about 20% memory without content loss...\n\n... just remove whitespaces 😁\n\n![memory-trick](https://i.ibb.co/xhndhMQ/memory-trick.png)",
    "439367": "Tried a few ways, most of them failed but there is one way succeeded. Assume that you have unzipped the `train_simplified.zip` into directory `train_simplified/` which contains 340 csv files. \n\nFirst, merge 340 csv files into a single csv file and only keep two columns, `drawing` and `y`, `y` is an integer converted from the `word` column, the file size of `train_simplfied.csv` is 21GB. Then you can just load the whole dataset by `df = pd.read_csv('train_simplified.csv')`, which will consume around 25GB RAM.\n\nThe following are a few failed ways I've tried:\n\n1. Load the `train_simplified.csv` using `df = pd.read_csv('train_simplified.csv', converters={'drawing': json.loads})`, this method ended with consuming up all RAM and failed.\n\n2. Then I optimized the parsing function:\n\n    ```python\n    def parse_strokes(s):\n    x = json.loads(s)\n    return list(map(np.uint8, x))\n    ```\n\n    Because all the values in `drawing` column don't exceed 256 they can be stored in `np.uint8`, however `df = pd.read_csv('train_simplified.csv', converters={'drawing': parse_strokes})` ended with consuming up all RAM and failed again.\n\n3. Merge 340 csv files into a single hdf5 file, only keep two columns `drawing` and `y`, and parse the `drawing` by  either `json.loads()`  or `parse_strokes()` . After parsing strings to lists, the file size does reduces to `12G`, but when I tried to load the dataset into memory, `df = pd.read_hdf('train_simplified.hdf5')`, the RAM consumption grew over 50GB,  since my machine only has 52GB RAM it failed too.\n\nThe key takeaway is that **do NOT parse the `drawing` column in advance, instead you should parse them on the fly during training**. It seems list of integers consume much more memory than plain strings.\n"
  }
}