{
  "id": 67502,
  "title": "Applying function to pd.dataframe efficiently to get images",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/67502",
  "author_name": "Prajjwal",
  "post_date": "2018-10-03T11:05:26.286000",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Many people have shared their kernels, and many of them are using ConvNet approach, or at some point. When we get a pd.dataframe, The way to extract images can be many, but I've used\nthis <code>df['drawing'].apply(strokes2img)</code> where <code>strokes2img</code> returns the image. My concern is is there any efficient way to carry out this operation more efficiently. I've also used bag from dask but then converting it back to list is very expensive operation again which nullifies the benefit received at the first point. I'm using 4 CPUs with P100, and this above mentioned operation crashes the kernel every time. Is there any way can this operation be ported to GPU or more efficient aproach?</p>",
  "messages": [
    {
      "id": 398252,
      "postDate": "2018-10-03T18:47:20.100Z",
      "content": "<p>You may want to consider moving away from dataframes for this, and just using a numpy array.</p>\n\n<p>Something like</p>\n\n<pre><code>np.array([strokes2img(x) for x in df['drawing'].values])\n</code></pre>\n\n<p>would return a single 4D array (assuming strokes2img returns a 3D array) containing all the drawings. This would be easy to then use for training a convnet for example. <code>df[col].apply()</code> is an expensive operation in my experience, it is much faster to iterate over <code>df[col].values</code>, although I suspect your main computational constraint is in the strokes2img function and actually nothing to do with pandas.</p>\n\n<p>Personally, I am using a multi-threaded approach to convert the data to images, something like:</p>\n\n<pre><code>from multiprocessing import Pool\npool = Pool(8) # Number of threads in your system\nimgs = pool.map(strokes2img, df['drawing'].values)\n</code></pre>\n\n<p>which will apply the function across all your cores. (I am personally using mlcrate.SuperPool instead of the built in multiprocessing Pool as it also gives you a progressbar during the function mapping)</p>",
      "rawMarkdown": "You may want to consider moving away from dataframes for this, and just using a numpy array.\n\nSomething like\n\n    np.array([strokes2img(x) for x in df['drawing'].values])\n\nwould return a single 4D array (assuming strokes2img returns a 3D array) containing all the drawings. This would be easy to then use for training a convnet for example. `df[col].apply()` is an expensive operation in my experience, it is much faster to iterate over `df[col].values`, although I suspect your main computational constraint is in the strokes2img function and actually nothing to do with pandas.\n\nPersonally, I am using a multi-threaded approach to convert the data to images, something like:\n\n    from multiprocessing import Pool\n    pool = Pool(8) # Number of threads in your system\n    imgs = pool.map(strokes2img, df['drawing'].values)\n\nwhich will apply the function across all your cores. (I am personally using mlcrate.SuperPool instead of the built in multiprocessing Pool as it also gives you a progressbar during the function mapping)",
      "votes": 5,
      "replies": [
        {
          "id": 398435,
          "postDate": "2018-10-04T05:41:34.533Z",
          "content": "<p>Thanks ! This was really helpful. How much time does it take for this op ? It still crashes my kernel.</p>",
          "rawMarkdown": "Thanks ! This was really helpful. How much time does it take for this op ? It still crashes my kernel."
        },
        {
          "id": 398442,
          "postDate": "2018-10-04T05:52:32.217Z",
          "content": "<p>You seem to be using stoke based classification and images at the same time, if you have any resource in mind, it would be really helpful if you could share it. I haven't seen any implementation of such method.</p>",
          "rawMarkdown": "You seem to be using stoke based classification and images at the same time, if you have any resource in mind, it would be really helpful if you could share it. I haven't seen any implementation of such method."
        },
        {
          "id": 400731,
          "postDate": "2018-10-08T19:31:13.327Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 406016,
          "postDate": "2018-10-18T13:52:36.343Z",
          "content": "<p>Im using cv2 to make the conversion, check that too maybe, would not be hard to make a quick test to compare the 3 methods!</p>",
          "rawMarkdown": "Im using cv2 to make the conversion, check that too maybe, would not be hard to make a quick test to compare the 3 methods!"
        }
      ]
    },
    {
      "id": 397940,
      "postDate": "2018-10-03T11:05:26.287Z",
      "content": "<p>Many people have shared their kernels, and many of them are using ConvNet approach, or at some point. When we get a pd.dataframe, The way to extract images can be many, but I've used\nthis <code>df['drawing'].apply(strokes2img)</code> where <code>strokes2img</code> returns the image. My concern is is there any efficient way to carry out this operation more efficiently. I've also used bag from dask but then converting it back to list is very expensive operation again which nullifies the benefit received at the first point. I'm using 4 CPUs with P100, and this above mentioned operation crashes the kernel every time. Is there any way can this operation be ported to GPU or more efficient aproach?</p>",
      "rawMarkdown": "Many people have shared their kernels, and many of them are using ConvNet approach, or at some point. When we get a pd.dataframe, The way to extract images can be many, but I've used\nthis ```df['drawing'].apply(strokes2img)``` where `strokes2img` returns the image. My concern is is there any efficient way to carry out this operation more efficiently. I've also used bag from dask but then converting it back to list is very expensive operation again which nullifies the benefit received at the first point. I'm using 4 CPUs with P100, and this above mentioned operation crashes the kernel every time. Is there any way can this operation be ported to GPU or more efficient aproach?",
      "votes": 3
    },
    {
      "id": 427596,
      "postDate": "2018-11-25T20:03:03.200Z",
      "content": "<p>I did a benchmark and it shows that cv2 is 9x faster than pillow, so get rid of pillow and use cv2 instead.</p>\n\n<p>The following is my drawing code:</p>\n\n<p><code>python\ndef draw_simplified_strokes_cv2(strokes, width=256, lw=6, ink_decay=True):\n    img = np.full((256, 256), 255, np.uint8)\n    for t, stroke in enumerate(strokes):\n        color = min(t, 10) * 13 if ink_decay else 0\n        for i in range(len(stroke[0]) - 1):\n            _ = cv2.line(img, (stroke[0][i], stroke[1][i]),\n                         (stroke[0][i + 1], stroke[1][i + 1]), color, lw)\n    if width != 256:\n        return cv2.resize(img, (width, width))\n    else:\n        return img\n</code></p>",
      "rawMarkdown": "I did a benchmark and it shows that cv2 is 9x faster than pillow, so get rid of pillow and use cv2 instead.\n\nThe following is my drawing code:\n\n```python\ndef draw_simplified_strokes_cv2(strokes, width=256, lw=6, ink_decay=True):\n    img = np.full((256, 256), 255, np.uint8)\n    for t, stroke in enumerate(strokes):\n        color = min(t, 10) * 13 if ink_decay else 0\n        for i in range(len(stroke[0]) - 1):\n            _ = cv2.line(img, (stroke[0][i], stroke[1][i]),\n                         (stroke[0][i + 1], stroke[1][i + 1]), color, lw)\n    if width != 256:\n        return cv2.resize(img, (width, width))\n    else:\n        return img\n```",
      "votes": 1
    },
    {
      "id": 398200,
      "postDate": "2018-10-03T16:45:36.910Z",
      "content": "<p>Maybe try a GPU dataframe? I've never used it but the benchmark data is really impressive. <a href=\"https://github.com/gpuopenanalytics/pygdf\">https://github.com/gpuopenanalytics/pygdf</a> </p>\n\n<p>Another option might be to dump the image plot to a tensor and go forward with tensorflow or pytorch.</p>",
      "rawMarkdown": "Maybe try a GPU dataframe? I've never used it but the benchmark data is really impressive. https://github.com/gpuopenanalytics/pygdf \n\nAnother option might be to dump the image plot to a tensor and go forward with tensorflow or pytorch.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 398252,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "2018-10-03T18:47:20.100000",
      "content": "<p>You may want to consider moving away from dataframes for this, and just using a numpy array.</p>\n\n<p>Something like</p>\n\n<pre><code>np.array([strokes2img(x) for x in df['drawing'].values])\n</code></pre>\n\n<p>would return a single 4D array (assuming strokes2img returns a 3D array) containing all the drawings. This would be easy to then use for training a convnet for example. <code>df[col].apply()</code> is an expensive operation in my experience, it is much faster to iterate over <code>df[col].values</code>, although I suspect your main computational constraint is in the strokes2img function and actually nothing to do with pandas.</p>\n\n<p>Personally, I am using a multi-threaded approach to convert the data to images, something like:</p>\n\n<pre><code>from multiprocessing import Pool\npool = Pool(8) # Number of threads in your system\nimgs = pool.map(strokes2img, df['drawing'].values)\n</code></pre>\n\n<p>which will apply the function across all your cores. (I am personally using mlcrate.SuperPool instead of the built in multiprocessing Pool as it also gives you a progressbar during the function mapping)</p>",
      "votes": 5,
      "replies": [
        {
          "id": 398435,
          "author_name": "Prajjwal",
          "author_url": "",
          "post_date": "2018-10-04T05:41:34.533000",
          "content": "<p>Thanks ! This was really helpful. How much time does it take for this op ? It still crashes my kernel.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 398442,
          "author_name": "Prajjwal",
          "author_url": "",
          "post_date": "2018-10-04T05:52:32.217000",
          "content": "<p>You seem to be using stoke based classification and images at the same time, if you have any resource in mind, it would be really helpful if you could share it. I haven't seen any implementation of such method.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 400731,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-10-08T19:31:13.327000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 406016,
          "author_name": "André Neves",
          "author_url": "",
          "post_date": "2018-10-18T13:52:36.343000",
          "content": "<p>Im using cv2 to make the conversion, check that too maybe, would not be hard to make a quick test to compare the 3 methods!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 427596,
      "author_name": "[he.ai]soulmachine",
      "author_url": "",
      "post_date": "2018-11-25T20:03:03.200000",
      "content": "<p>I did a benchmark and it shows that cv2 is 9x faster than pillow, so get rid of pillow and use cv2 instead.</p>\n\n<p>The following is my drawing code:</p>\n\n<p><code>python\ndef draw_simplified_strokes_cv2(strokes, width=256, lw=6, ink_decay=True):\n    img = np.full((256, 256), 255, np.uint8)\n    for t, stroke in enumerate(strokes):\n        color = min(t, 10) * 13 if ink_decay else 0\n        for i in range(len(stroke[0]) - 1):\n            _ = cv2.line(img, (stroke[0][i], stroke[1][i]),\n                         (stroke[0][i + 1], stroke[1][i + 1]), color, lw)\n    if width != 256:\n        return cv2.resize(img, (width, width))\n    else:\n        return img\n</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 398200,
      "author_name": "JohnM",
      "author_url": "",
      "post_date": "2018-10-03T16:45:36.910000",
      "content": "<p>Maybe try a GPU dataframe? I've never used it but the benchmark data is really impressive. <a href=\"https://github.com/gpuopenanalytics/pygdf\">https://github.com/gpuopenanalytics/pygdf</a> </p>\n\n<p>Another option might be to dump the image plot to a tensor and go forward with tensorflow or pytorch.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "398252": "You may want to consider moving away from dataframes for this, and just using a numpy array.\n\nSomething like\n\n    np.array([strokes2img(x) for x in df['drawing'].values])\n\nwould return a single 4D array (assuming strokes2img returns a 3D array) containing all the drawings. This would be easy to then use for training a convnet for example. `df[col].apply()` is an expensive operation in my experience, it is much faster to iterate over `df[col].values`, although I suspect your main computational constraint is in the strokes2img function and actually nothing to do with pandas.\n\nPersonally, I am using a multi-threaded approach to convert the data to images, something like:\n\n    from multiprocessing import Pool\n    pool = Pool(8) # Number of threads in your system\n    imgs = pool.map(strokes2img, df['drawing'].values)\n\nwhich will apply the function across all your cores. (I am personally using mlcrate.SuperPool instead of the built in multiprocessing Pool as it also gives you a progressbar during the function mapping)",
    "397940": "Many people have shared their kernels, and many of them are using ConvNet approach, or at some point. When we get a pd.dataframe, The way to extract images can be many, but I've used\nthis ```df['drawing'].apply(strokes2img)``` where `strokes2img` returns the image. My concern is is there any efficient way to carry out this operation more efficiently. I've also used bag from dask but then converting it back to list is very expensive operation again which nullifies the benefit received at the first point. I'm using 4 CPUs with P100, and this above mentioned operation crashes the kernel every time. Is there any way can this operation be ported to GPU or more efficient aproach?",
    "427596": "I did a benchmark and it shows that cv2 is 9x faster than pillow, so get rid of pillow and use cv2 instead.\n\nThe following is my drawing code:\n\n```python\ndef draw_simplified_strokes_cv2(strokes, width=256, lw=6, ink_decay=True):\n    img = np.full((256, 256), 255, np.uint8)\n    for t, stroke in enumerate(strokes):\n        color = min(t, 10) * 13 if ink_decay else 0\n        for i in range(len(stroke[0]) - 1):\n            _ = cv2.line(img, (stroke[0][i], stroke[1][i]),\n                         (stroke[0][i + 1], stroke[1][i + 1]), color, lw)\n    if width != 256:\n        return cv2.resize(img, (width, width))\n    else:\n        return img\n```",
    "398200": "Maybe try a GPU dataframe? I've never used it but the benchmark data is really impressive. https://github.com/gpuopenanalytics/pygdf \n\nAnother option might be to dump the image plot to a tensor and go forward with tensorflow or pytorch."
  }
}