{
  "id": 71703,
  "title": "data generation speed",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/71703",
  "author_name": "Colin",
  "post_date": "2018-11-15T19:35:42.348000",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Let's say I want to use all of the data for training (sine validation set).  There are ~50M instances.</p>\n\n<p>I can generate batches totalling 1_000_000 instances (read and converted to 64x64 images) in ~3 minutes, drawn using the code below...</p>\n\n<p>That means for all 50M drawings, it'd take ~2.5 hours just to generate the data once through, and there's way too much to cache all of it, at least in kaggle kernel space.</p>\n\n<p>Anybody doing something nicer? </p>\n\n<p>Thinking about parallelization - anybody tried it? I believe that numpy is configured to use multiple cores - any chance the code below is already parallel somehow through numpy? Asking because I've attempted to parallelize this sort of thing before with little success, only to find out that seemingly serial numpy operations were actually using multiple cores anyway. Watching the kaggle kernel cpu%, it never exceeds 101%. Not sure if I trust that.</p>\n\n<p>Happy to share all generation code if anybody's interested.</p>\n\n<pre><code>def draw_cv2(raw_strokes, size=256, linewidth=6, time_color=True):\n    img = np.zeros((256, 256), np.uint8)\n    for t, stroke in enumerate(raw_strokes):\n        for i in range(len(stroke[0]) - 1):\n            color = 255 - min(t, 10) * 13 if time_color else 255\n            _ = cv2.line(img, (stroke[0][i], stroke[1][i]), \n                         (stroke[0][i + 1], stroke[1][i + 1]), color, linewidth)\n    if size != 256:\n        return cv2.resize(img, (size, size))\n    else:\n        return img\n\ndef one_hot(a, num_classes):\n    return np.squeeze(np.eye(num_classes)[a.reshape(-1)])\n\ndef generate_batch_data_from_df(df, num_classes, batch_size, image_size, line_width):\n    for i in range(0, df.shape[0], batch_size):\n        y = one_hot(df['word_index'][i:i+batch_size].values, num_classes)\n\n        x = np.zeros((y.shape[0], image_size, image_size))\n        for j, raw_strokes in enumerate(df.drawing[i:i+batch_size]):\n            x[j] = draw_cv2(raw_strokes, size=image_size, linewidth=line_width)\n        x = x / 255.\n        x = x.reshape((y.shape[0], image_size, image_size, 1)).astype(np.float32)\n\n        yield (x, y)\n</code></pre>",
  "messages": [
    {
      "id": 422103,
      "postDate": "2018-11-15T19:44:05.853Z",
      "content": "<p>50M drawings in 2.5 hours is just fine. Your main bottleneck will be the deep learning part on the GPU.</p>",
      "rawMarkdown": "50M drawings in 2.5 hours is just fine. Your main bottleneck will be the deep learning part on the GPU.",
      "votes": 1
    },
    {
      "id": 422097,
      "postDate": "2018-11-15T19:35:42.350Z",
      "content": "<p>Let's say I want to use all of the data for training (sine validation set).  There are ~50M instances.</p>\n\n<p>I can generate batches totalling 1_000_000 instances (read and converted to 64x64 images) in ~3 minutes, drawn using the code below...</p>\n\n<p>That means for all 50M drawings, it'd take ~2.5 hours just to generate the data once through, and there's way too much to cache all of it, at least in kaggle kernel space.</p>\n\n<p>Anybody doing something nicer? </p>\n\n<p>Thinking about parallelization - anybody tried it? I believe that numpy is configured to use multiple cores - any chance the code below is already parallel somehow through numpy? Asking because I've attempted to parallelize this sort of thing before with little success, only to find out that seemingly serial numpy operations were actually using multiple cores anyway. Watching the kaggle kernel cpu%, it never exceeds 101%. Not sure if I trust that.</p>\n\n<p>Happy to share all generation code if anybody's interested.</p>\n\n<pre><code>def draw_cv2(raw_strokes, size=256, linewidth=6, time_color=True):\n    img = np.zeros((256, 256), np.uint8)\n    for t, stroke in enumerate(raw_strokes):\n        for i in range(len(stroke[0]) - 1):\n            color = 255 - min(t, 10) * 13 if time_color else 255\n            _ = cv2.line(img, (stroke[0][i], stroke[1][i]), \n                         (stroke[0][i + 1], stroke[1][i + 1]), color, linewidth)\n    if size != 256:\n        return cv2.resize(img, (size, size))\n    else:\n        return img\n\ndef one_hot(a, num_classes):\n    return np.squeeze(np.eye(num_classes)[a.reshape(-1)])\n\ndef generate_batch_data_from_df(df, num_classes, batch_size, image_size, line_width):\n    for i in range(0, df.shape[0], batch_size):\n        y = one_hot(df['word_index'][i:i+batch_size].values, num_classes)\n\n        x = np.zeros((y.shape[0], image_size, image_size))\n        for j, raw_strokes in enumerate(df.drawing[i:i+batch_size]):\n            x[j] = draw_cv2(raw_strokes, size=image_size, linewidth=line_width)\n        x = x / 255.\n        x = x.reshape((y.shape[0], image_size, image_size, 1)).astype(np.float32)\n\n        yield (x, y)\n</code></pre>",
      "rawMarkdown": "Let's say I want to use all of the data for training (sine validation set).  There are ~50M instances.\n\nI can generate batches totalling 1_000_000 instances (read and converted to 64x64 images) in ~3 minutes, drawn using the code below...\n\nThat means for all 50M drawings, it'd take ~2.5 hours just to generate the data once through, and there's way too much to cache all of it, at least in kaggle kernel space.\n\nAnybody doing something nicer? \n\nThinking about parallelization - anybody tried it? I believe that numpy is configured to use multiple cores - any chance the code below is already parallel somehow through numpy? Asking because I've attempted to parallelize this sort of thing before with little success, only to find out that seemingly serial numpy operations were actually using multiple cores anyway. Watching the kaggle kernel cpu%, it never exceeds 101%. Not sure if I trust that.\n\nHappy to share all generation code if anybody's interested.\n\n    def draw_cv2(raw_strokes, size=256, linewidth=6, time_color=True):\n        img = np.zeros((256, 256), np.uint8)\n        for t, stroke in enumerate(raw_strokes):\n            for i in range(len(stroke[0]) - 1):\n                color = 255 - min(t, 10) * 13 if time_color else 255\n                _ = cv2.line(img, (stroke[0][i], stroke[1][i]), \n                             (stroke[0][i + 1], stroke[1][i + 1]), color, linewidth)\n        if size != 256:\n            return cv2.resize(img, (size, size))\n        else:\n            return img\n        \n    def one_hot(a, num_classes):\n        return np.squeeze(np.eye(num_classes)[a.reshape(-1)])\n    \n    def generate_batch_data_from_df(df, num_classes, batch_size, image_size, line_width):\n        for i in range(0, df.shape[0], batch_size):\n            y = one_hot(df['word_index'][i:i+batch_size].values, num_classes)\n            \n            x = np.zeros((y.shape[0], image_size, image_size))\n            for j, raw_strokes in enumerate(df.drawing[i:i+batch_size]):\n                x[j] = draw_cv2(raw_strokes, size=image_size, linewidth=line_width)\n            x = x / 255.\n            x = x.reshape((y.shape[0], image_size, image_size, 1)).astype(np.float32)\n            \n            yield (x, y)",
      "votes": 1
    },
    {
      "id": 423755,
      "postDate": "2018-11-19T02:07:08.003Z",
      "content": "<p>For those with computers that have multiple physical cpus (i.e. not kaggle kernels):</p>\n\n<p><a href=\"https://github.com/colinator/doodle_generator/blob/master/data_generator_uniform_final.ipynb\">https://github.com/colinator/doodle_generator/blob/master/data_generator_uniform_final.ipynb</a></p>",
      "rawMarkdown": "For those with computers that have multiple physical cpus (i.e. not kaggle kernels):\n\nhttps://github.com/colinator/doodle_generator/blob/master/data_generator_uniform_final.ipynb"
    }
  ],
  "comments": [
    {
      "id": 422103,
      "author_name": "beluga",
      "author_url": "",
      "post_date": "2018-11-15T19:44:05.853000",
      "content": "<p>50M drawings in 2.5 hours is just fine. Your main bottleneck will be the deep learning part on the GPU.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 423755,
      "author_name": "Colin",
      "author_url": "",
      "post_date": "2018-11-19T02:07:08.003000",
      "content": "<p>For those with computers that have multiple physical cpus (i.e. not kaggle kernels):</p>\n\n<p><a href=\"https://github.com/colinator/doodle_generator/blob/master/data_generator_uniform_final.ipynb\">https://github.com/colinator/doodle_generator/blob/master/data_generator_uniform_final.ipynb</a></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "422103": "50M drawings in 2.5 hours is just fine. Your main bottleneck will be the deep learning part on the GPU.",
    "422097": "Let's say I want to use all of the data for training (sine validation set).  There are ~50M instances.\n\nI can generate batches totalling 1_000_000 instances (read and converted to 64x64 images) in ~3 minutes, drawn using the code below...\n\nThat means for all 50M drawings, it'd take ~2.5 hours just to generate the data once through, and there's way too much to cache all of it, at least in kaggle kernel space.\n\nAnybody doing something nicer? \n\nThinking about parallelization - anybody tried it? I believe that numpy is configured to use multiple cores - any chance the code below is already parallel somehow through numpy? Asking because I've attempted to parallelize this sort of thing before with little success, only to find out that seemingly serial numpy operations were actually using multiple cores anyway. Watching the kaggle kernel cpu%, it never exceeds 101%. Not sure if I trust that.\n\nHappy to share all generation code if anybody's interested.\n\n    def draw_cv2(raw_strokes, size=256, linewidth=6, time_color=True):\n        img = np.zeros((256, 256), np.uint8)\n        for t, stroke in enumerate(raw_strokes):\n            for i in range(len(stroke[0]) - 1):\n                color = 255 - min(t, 10) * 13 if time_color else 255\n                _ = cv2.line(img, (stroke[0][i], stroke[1][i]), \n                             (stroke[0][i + 1], stroke[1][i + 1]), color, linewidth)\n        if size != 256:\n            return cv2.resize(img, (size, size))\n        else:\n            return img\n        \n    def one_hot(a, num_classes):\n        return np.squeeze(np.eye(num_classes)[a.reshape(-1)])\n    \n    def generate_batch_data_from_df(df, num_classes, batch_size, image_size, line_width):\n        for i in range(0, df.shape[0], batch_size):\n            y = one_hot(df['word_index'][i:i+batch_size].values, num_classes)\n            \n            x = np.zeros((y.shape[0], image_size, image_size))\n            for j, raw_strokes in enumerate(df.drawing[i:i+batch_size]):\n                x[j] = draw_cv2(raw_strokes, size=image_size, linewidth=line_width)\n            x = x / 255.\n            x = x.reshape((y.shape[0], image_size, image_size, 1)).astype(np.float32)\n            \n            yield (x, y)",
    "423755": "For those with computers that have multiple physical cpus (i.e. not kaggle kernels):\n\nhttps://github.com/colinator/doodle_generator/blob/master/data_generator_uniform_final.ipynb"
  }
}