{
  "id": 67036,
  "title": "Transforming the data into images?",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/67036",
  "author_name": "Daniel Möller",
  "post_date": "2018-09-27T22:39:53.412000",
  "votes": 12,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Hello, everyone!</p>\n\n<p>Are there preset tools capable of creating pixel images from the given data?\nWhich other approaches are interesting?</p>",
  "messages": [
    {
      "id": 395069,
      "postDate": "2018-09-27T22:39:53.413Z",
      "content": "<p>Hello, everyone!</p>\n\n<p>Are there preset tools capable of creating pixel images from the given data?\nWhich other approaches are interesting?</p>",
      "rawMarkdown": "Hello, everyone!\n\nAre there preset tools capable of creating pixel images from the given data?\nWhich other approaches are interesting?\n\n",
      "votes": 12
    },
    {
      "id": 396480,
      "postDate": "2018-09-30T20:55:08.300Z",
      "content": "<p>matplotlib might be inefficient. Alternative option is using PIL/pillow :</p>\n\n<pre><code>import numpy as np\nfrom PIL import Image, ImageDraw\n\ndef draw_it(raw_strokes):\n    image = Image.new(\"P\", (255,255), color=255)\n    image_draw = ImageDraw.Draw(image)\n\n    for stroke in eval(raw_strokes):\n        for i in range(len(stroke[0])-1):\n\n            image_draw.line([stroke[0][i], \n                             stroke[1][i],\n                             stroke[0][i+1], \n                             stroke[1][i+1]],\n                            fill=0, width=6)\n    return np.array(image)\n</code></pre>",
      "rawMarkdown": "matplotlib might be inefficient. Alternative option is using PIL/pillow :\n\n    import numpy as np\n    from PIL import Image, ImageDraw\n\n    def draw_it(raw_strokes):\n        image = Image.new(\"P\", (255,255), color=255)\n        image_draw = ImageDraw.Draw(image)\n\n        for stroke in eval(raw_strokes):\n            for i in range(len(stroke[0])-1):\n\n                image_draw.line([stroke[0][i], \n                                 stroke[1][i],\n                                 stroke[0][i+1], \n                                 stroke[1][i+1]],\n                                fill=0, width=6)\n        return np.array(image)",
      "votes": 8,
      "replies": [
        {
          "id": 397749,
          "postDate": "2018-10-03T03:22:18.010Z",
          "content": "<p>Pillow was much faster for me than matplotlib and without the memory spikes - thanks!</p>",
          "rawMarkdown": "Pillow was much faster for me than matplotlib and without the memory spikes - thanks!",
          "votes": 1
        },
        {
          "id": 398030,
          "postDate": "2018-10-03T13:18:31.987Z",
          "content": "<p>What's cline ? Is it dataframe ?</p>",
          "rawMarkdown": "What's cline ? Is it dataframe ?",
          "votes": 1
        },
        {
          "id": 398038,
          "postDate": "2018-10-03T13:26:45.053Z",
          "content": "<p>It was a typo, this should be \"raw_strokes\", the stroke string in dataframe.</p>\n\n<p>It's fixed now. Thanks for pointing it out.</p>",
          "rawMarkdown": "It was a typo, this should be \"raw_strokes\", the stroke string in dataframe.\n\nIt's fixed now. Thanks for pointing it out."
        },
        {
          "id": 400057,
          "postDate": "2018-10-07T13:31:44.560Z",
          "content": "<p>Your approach is really fast. <br>\nThere is a problem that the image size (255x255) maybe not enough to draw some images. Therefore, you can get white image then.</p>",
          "rawMarkdown": "Your approach is really fast.  \nThere is a problem that the image size (255x255) maybe not enough to draw some images. Therefore, you can get white image then."
        },
        {
          "id": 400799,
          "postDate": "2018-10-08T22:58:35.100Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 402208,
          "postDate": "2018-10-11T10:27:14.287Z",
          "content": "<p>Your approach is really good, thank you for the sharing). BTW, looks like json.loads should be faster than ast.literal_eval or eval (<a href=\"https://stackoverflow.com/questions/9949533/python-eval-vs-ast-literal-eval-vs-json-decode\">https://stackoverflow.com/questions/9949533/python-eval-vs-ast-literal-eval-vs-json-decode</a>), so I think it could be even faster.</p>",
          "rawMarkdown": " Your approach is really good, thank you for the sharing). BTW, looks like json.loads should be faster than ast.literal_eval or eval (https://stackoverflow.com/questions/9949533/python-eval-vs-ast-literal-eval-vs-json-decode), so I think it could be even faster.",
          "votes": 1
        },
        {
          "id": 417692,
          "postDate": "2018-11-08T16:38:22.277Z",
          "content": "<p>how do you determine the image size ?  seems the 256*256 is not big enough,  and different drawings has different image size</p>",
          "rawMarkdown": "how do you determine the image size ?  seems the 256*256 is not big enough,  and different drawings has different image size"
        }
      ]
    },
    {
      "id": 395135,
      "postDate": "2018-09-28T02:37:18.837Z",
      "content": "<p>Just one update, just got a working code with help of one of the written kernels, it's basicaly what RDizzl3 said:</p>\n\n<pre><code>def drawing_to_np(drawing, shape=(64, 64)):\n    drawing = eval(drawing)\n    fig, ax = plt.subplots()\n    for x,y in drawing:\n        ax.plot(x, y, marker='.')\n        ax.axis('off')\n    fig.canvas.draw()\n    # Convert images to numpy arrat\n    np_drawing = np.array(fig.canvas.renderer._renderer)\n\n    return cv2.resize(np_drawing, shape) # Resize array\n</code></pre>\n\n<p>How you could apply:</p>\n\n<pre><code>train['drawing_converted'] = train['drawing'].map(drawing_to_np)\n</code></pre>\n\n<p>but maybe what you may want to do is use the function with batch loading, hope i could help.</p>",
      "rawMarkdown": "Just one update, just got a working code with help of one of the written kernels, it's basicaly what RDizzl3 said:\n\n    def drawing_to_np(drawing, shape=(64, 64)):\n        drawing = eval(drawing)\n        fig, ax = plt.subplots()\n        for x,y in drawing:\n            ax.plot(x, y, marker='.')\n            ax.axis('off')\n        fig.canvas.draw()\n        # Convert images to numpy arrat\n        np_drawing = np.array(fig.canvas.renderer._renderer)\n\n        return cv2.resize(np_drawing, shape) # Resize array\n\nHow you could apply:\n\n    train['drawing_converted'] = train['drawing'].map(drawing_to_np)\n\nbut maybe what you may want to do is use the function with batch loading, hope i could help.",
      "votes": 3,
      "replies": [
        {
          "id": 395187,
          "postDate": "2018-09-28T05:21:25.040Z",
          "content": "<p>That's great for starting :)</p>",
          "rawMarkdown": "That's great for starting :)",
          "votes": 1
        },
        {
          "id": 395359,
          "postDate": "2018-09-28T12:29:49.527Z",
          "content": "<p>In case anyone want a working code as a demonstration of one way you can do this, <a href=\"https://www.kaggle.com/dimitreoliveira/quick-tips-converting-drawings-to-numpy-arrays\">checkout this kernel.</a></p>",
          "rawMarkdown": "In case anyone want a working code as a demonstration of one way you can do this, [checkout this kernel.](https://www.kaggle.com/dimitreoliveira/quick-tips-converting-drawings-to-numpy-arrays)"
        },
        {
          "id": 395526,
          "postDate": "2018-09-28T18:51:03.183Z",
          "content": "<p>I just wanted to make one more comment here. I ran this method now a few times and it was unable to finish every time. The issue? Leaving the figures open inside the for loop causes a MASSIVE memory leak! I order to fix this I had to add </p>\n\n<pre><code>plt.close('all')\nplt.gcf()\n</code></pre>\n\n<p>after every figure creation. Here is the reference.\n<a href=\"https://stackoverflow.com/questions/2364945/matplotlib-runs-out-of-memory-when-plotting-in-a-loop\">https://stackoverflow.com/questions/2364945/matplotlib-runs-out-of-memory-when-plotting-in-a-loop</a></p>",
          "rawMarkdown": "I just wanted to make one more comment here. I ran this method now a few times and it was unable to finish every time. The issue? Leaving the figures open inside the for loop causes a MASSIVE memory leak! I order to fix this I had to add \n\n    plt.close('all')\n    plt.gcf()\n\nafter every figure creation. Here is the reference.\nhttps://stackoverflow.com/questions/2364945/matplotlib-runs-out-of-memory-when-plotting-in-a-loop",
          "votes": 6
        },
        {
          "id": 398031,
          "postDate": "2018-10-03T13:19:11.743Z",
          "content": "<p>Even after adding these two snippet, kernel crashes everytime. Is there any more optimized way ?</p>",
          "rawMarkdown": "Even after adding these two snippet, kernel crashes everytime. Is there any more optimized way ?"
        },
        {
          "id": 399445,
          "postDate": "2018-10-05T19:41:14.740Z",
          "content": "<p>great code, worked well for me in ipython, but I have some problems running in python: <a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/67820\">https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/67820</a> does anybody knows why?</p>",
          "rawMarkdown": "great code, worked well for me in ipython, but I have some problems running in python: https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/67820 does anybody knows why?"
        }
      ]
    },
    {
      "id": 395080,
      "postDate": "2018-09-27T23:34:44.747Z",
      "content": "<p>Hi Daniel, currently i'm using the following snippet.</p>\n\n<pre><code>drawings = [ast.literal_eval(pts) for pts in train[:1]['drawing'].values]\n\nfor drawing in drawings:\n    for x,y in drawing:\n        plt.plot(x, y, marker='.')\n</code></pre>\n\n<p>where train is the training set, and you get \"ast\" lib with \"import ast\"</p>",
      "rawMarkdown": "Hi Daniel, currently i'm using the following snippet.\n\n    drawings = [ast.literal_eval(pts) for pts in train[:1]['drawing'].values]\n\n    for drawing in drawings:\n        for x,y in drawing:\n            plt.plot(x, y, marker='.')\n\nwhere train is the training set, and you get \"ast\" lib with \"import ast\"",
      "votes": 1,
      "replies": [
        {
          "id": 395084,
          "postDate": "2018-09-28T00:01:27.963Z",
          "content": "<p>Thanks for the tip :)</p>\n\n<p>But I meant not only to \"visualize\", but use as image data, a numpy image, for instance.</p>",
          "rawMarkdown": "Thanks for the tip :)\n\nBut I meant not only to \"visualize\", but use as image data, a numpy image, for instance.",
          "votes": 1
        },
        {
          "id": 395085,
          "postDate": "2018-09-28T00:11:38.480Z",
          "content": "<p>I am using something similar :) From this step you can also extract the matplotlib plot to a numpy array then you can use opencv to do some resizing and convert to grayscale if you wanted too. For now at least I feel like this would be much more efficient than loading the raw dataset. </p>\n\n<p>For context here is what the subsequent steps might look like:</p>\n\n<pre><code>fig, ax = plt.subplots()\n\ncode from above ...\n\nfig.canvas.draw()\ndata = np.fromstring(fig.canvas.tostring_rgb(), dtype=np.uint8, sep='')\ndata = data.reshape(fig.canvas.get_width_height()[::-1] + (3,))\ndata = cv2.resize(data, (28, 28), interpolation = cv2.INTER_AREA)\ndata =  cv2.cvtColor(data, cv2.COLOR_BGR2GRAY)\n</code></pre>",
          "rawMarkdown": "I am using something similar :) From this step you can also extract the matplotlib plot to a numpy array then you can use opencv to do some resizing and convert to grayscale if you wanted too. For now at least I feel like this would be much more efficient than loading the raw dataset. \n\nFor context here is what the subsequent steps might look like:\n\n    fig, ax = plt.subplots()\n    \n    code from above ...\n    \n    fig.canvas.draw()\n    data = np.fromstring(fig.canvas.tostring_rgb(), dtype=np.uint8, sep='')\n    data = data.reshape(fig.canvas.get_width_height()[::-1] + (3,))\n    data = cv2.resize(data, (28, 28), interpolation = cv2.INTER_AREA)\n    data =  cv2.cvtColor(data, cv2.COLOR_BGR2GRAY)",
          "votes": 3
        },
        {
          "id": 395114,
          "postDate": "2018-09-28T01:42:51.383Z",
          "content": "<p>Oh Daniel, now i got what you meant, i'm still writing my CNN code, but this may help you, it's a way to <a href=\"https://matplotlib.org/gallery/misc/agg_buffer_to_array.html\">convert plot images into numpy arrays</a>.</p>",
          "rawMarkdown": "Oh Daniel, now i got what you meant, i'm still writing my CNN code, but this may help you, it's a way to [convert plot images into numpy arrays](https://matplotlib.org/gallery/misc/agg_buffer_to_array.html).",
          "votes": 1
        }
      ]
    },
    {
      "id": 397681,
      "postDate": "2018-10-02T22:54:07.723Z",
      "content": "<p>There's also the direct conversion method like in <a href=\"https://www.kaggle.com/xmaayy/converting-vector-images-to-binary-images\">Xander May's kernel</a>. It's fast and memory efficient and should work well. I didn't get good results with classification though.</p>",
      "rawMarkdown": "There's also the direct conversion method like in [Xander May's kernel](https://www.kaggle.com/xmaayy/converting-vector-images-to-binary-images). It's fast and memory efficient and should work well. I didn't get good results with classification though."
    }
  ],
  "comments": [
    {
      "id": 396480,
      "author_name": "Miha Skalic",
      "author_url": "",
      "post_date": "2018-09-30T20:55:08.300000",
      "content": "<p>matplotlib might be inefficient. Alternative option is using PIL/pillow :</p>\n\n<pre><code>import numpy as np\nfrom PIL import Image, ImageDraw\n\ndef draw_it(raw_strokes):\n    image = Image.new(\"P\", (255,255), color=255)\n    image_draw = ImageDraw.Draw(image)\n\n    for stroke in eval(raw_strokes):\n        for i in range(len(stroke[0])-1):\n\n            image_draw.line([stroke[0][i], \n                             stroke[1][i],\n                             stroke[0][i+1], \n                             stroke[1][i+1]],\n                            fill=0, width=6)\n    return np.array(image)\n</code></pre>",
      "votes": 8,
      "replies": [
        {
          "id": 397749,
          "author_name": "JohnM",
          "author_url": "",
          "post_date": "2018-10-03T03:22:18.010000",
          "content": "<p>Pillow was much faster for me than matplotlib and without the memory spikes - thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 398030,
          "author_name": "Prajjwal",
          "author_url": "",
          "post_date": "2018-10-03T13:18:31.987000",
          "content": "<p>What's cline ? Is it dataframe ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 398038,
          "author_name": "Miha Skalic",
          "author_url": "",
          "post_date": "2018-10-03T13:26:45.053000",
          "content": "<p>It was a typo, this should be \"raw_strokes\", the stroke string in dataframe.</p>\n\n<p>It's fixed now. Thanks for pointing it out.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 400057,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2018-10-07T13:31:44.560000",
          "content": "<p>Your approach is really fast. <br>\nThere is a problem that the image size (255x255) maybe not enough to draw some images. Therefore, you can get white image then.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 400799,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-10-08T22:58:35.100000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 402208,
          "author_name": "Danylo Honcharov",
          "author_url": "",
          "post_date": "2018-10-11T10:27:14.287000",
          "content": "<p>Your approach is really good, thank you for the sharing). BTW, looks like json.loads should be faster than ast.literal_eval or eval (<a href=\"https://stackoverflow.com/questions/9949533/python-eval-vs-ast-literal-eval-vs-json-decode\">https://stackoverflow.com/questions/9949533/python-eval-vs-ast-literal-eval-vs-json-decode</a>), so I think it could be even faster.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 417692,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "2018-11-08T16:38:22.277000",
          "content": "<p>how do you determine the image size ?  seems the 256*256 is not big enough,  and different drawings has different image size</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 395135,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2018-09-28T02:37:18.837000",
      "content": "<p>Just one update, just got a working code with help of one of the written kernels, it's basicaly what RDizzl3 said:</p>\n\n<pre><code>def drawing_to_np(drawing, shape=(64, 64)):\n    drawing = eval(drawing)\n    fig, ax = plt.subplots()\n    for x,y in drawing:\n        ax.plot(x, y, marker='.')\n        ax.axis('off')\n    fig.canvas.draw()\n    # Convert images to numpy arrat\n    np_drawing = np.array(fig.canvas.renderer._renderer)\n\n    return cv2.resize(np_drawing, shape) # Resize array\n</code></pre>\n\n<p>How you could apply:</p>\n\n<pre><code>train['drawing_converted'] = train['drawing'].map(drawing_to_np)\n</code></pre>\n\n<p>but maybe what you may want to do is use the function with batch loading, hope i could help.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 395187,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-09-28T05:21:25.040000",
          "content": "<p>That's great for starting :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 395359,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2018-09-28T12:29:49.527000",
          "content": "<p>In case anyone want a working code as a demonstration of one way you can do this, <a href=\"https://www.kaggle.com/dimitreoliveira/quick-tips-converting-drawings-to-numpy-arrays\">checkout this kernel.</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 395526,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2018-09-28T18:51:03.183000",
          "content": "<p>I just wanted to make one more comment here. I ran this method now a few times and it was unable to finish every time. The issue? Leaving the figures open inside the for loop causes a MASSIVE memory leak! I order to fix this I had to add </p>\n\n<pre><code>plt.close('all')\nplt.gcf()\n</code></pre>\n\n<p>after every figure creation. Here is the reference.\n<a href=\"https://stackoverflow.com/questions/2364945/matplotlib-runs-out-of-memory-when-plotting-in-a-loop\">https://stackoverflow.com/questions/2364945/matplotlib-runs-out-of-memory-when-plotting-in-a-loop</a></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 398031,
          "author_name": "Prajjwal",
          "author_url": "",
          "post_date": "2018-10-03T13:19:11.743000",
          "content": "<p>Even after adding these two snippet, kernel crashes everytime. Is there any more optimized way ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 399445,
          "author_name": "Vitaliy",
          "author_url": "",
          "post_date": "2018-10-05T19:41:14.740000",
          "content": "<p>great code, worked well for me in ipython, but I have some problems running in python: <a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/67820\">https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/67820</a> does anybody knows why?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 395080,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2018-09-27T23:34:44.747000",
      "content": "<p>Hi Daniel, currently i'm using the following snippet.</p>\n\n<pre><code>drawings = [ast.literal_eval(pts) for pts in train[:1]['drawing'].values]\n\nfor drawing in drawings:\n    for x,y in drawing:\n        plt.plot(x, y, marker='.')\n</code></pre>\n\n<p>where train is the training set, and you get \"ast\" lib with \"import ast\"</p>",
      "votes": 1,
      "replies": [
        {
          "id": 395084,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-09-28T00:01:27.963000",
          "content": "<p>Thanks for the tip :)</p>\n\n<p>But I meant not only to \"visualize\", but use as image data, a numpy image, for instance.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 395085,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2018-09-28T00:11:38.480000",
          "content": "<p>I am using something similar :) From this step you can also extract the matplotlib plot to a numpy array then you can use opencv to do some resizing and convert to grayscale if you wanted too. For now at least I feel like this would be much more efficient than loading the raw dataset. </p>\n\n<p>For context here is what the subsequent steps might look like:</p>\n\n<pre><code>fig, ax = plt.subplots()\n\ncode from above ...\n\nfig.canvas.draw()\ndata = np.fromstring(fig.canvas.tostring_rgb(), dtype=np.uint8, sep='')\ndata = data.reshape(fig.canvas.get_width_height()[::-1] + (3,))\ndata = cv2.resize(data, (28, 28), interpolation = cv2.INTER_AREA)\ndata =  cv2.cvtColor(data, cv2.COLOR_BGR2GRAY)\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 395114,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2018-09-28T01:42:51.383000",
          "content": "<p>Oh Daniel, now i got what you meant, i'm still writing my CNN code, but this may help you, it's a way to <a href=\"https://matplotlib.org/gallery/misc/agg_buffer_to_array.html\">convert plot images into numpy arrays</a>.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 397681,
      "author_name": "JohnM",
      "author_url": "",
      "post_date": "2018-10-02T22:54:07.723000",
      "content": "<p>There's also the direct conversion method like in <a href=\"https://www.kaggle.com/xmaayy/converting-vector-images-to-binary-images\">Xander May's kernel</a>. It's fast and memory efficient and should work well. I didn't get good results with classification though.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "395069": "Hello, everyone!\n\nAre there preset tools capable of creating pixel images from the given data?\nWhich other approaches are interesting?\n\n",
    "396480": "matplotlib might be inefficient. Alternative option is using PIL/pillow :\n\n    import numpy as np\n    from PIL import Image, ImageDraw\n\n    def draw_it(raw_strokes):\n        image = Image.new(\"P\", (255,255), color=255)\n        image_draw = ImageDraw.Draw(image)\n\n        for stroke in eval(raw_strokes):\n            for i in range(len(stroke[0])-1):\n\n                image_draw.line([stroke[0][i], \n                                 stroke[1][i],\n                                 stroke[0][i+1], \n                                 stroke[1][i+1]],\n                                fill=0, width=6)\n        return np.array(image)",
    "395135": "Just one update, just got a working code with help of one of the written kernels, it's basicaly what RDizzl3 said:\n\n    def drawing_to_np(drawing, shape=(64, 64)):\n        drawing = eval(drawing)\n        fig, ax = plt.subplots()\n        for x,y in drawing:\n            ax.plot(x, y, marker='.')\n            ax.axis('off')\n        fig.canvas.draw()\n        # Convert images to numpy arrat\n        np_drawing = np.array(fig.canvas.renderer._renderer)\n\n        return cv2.resize(np_drawing, shape) # Resize array\n\nHow you could apply:\n\n    train['drawing_converted'] = train['drawing'].map(drawing_to_np)\n\nbut maybe what you may want to do is use the function with batch loading, hope i could help.",
    "395080": "Hi Daniel, currently i'm using the following snippet.\n\n    drawings = [ast.literal_eval(pts) for pts in train[:1]['drawing'].values]\n\n    for drawing in drawings:\n        for x,y in drawing:\n            plt.plot(x, y, marker='.')\n\nwhere train is the training set, and you get \"ast\" lib with \"import ast\"",
    "397681": "There's also the direct conversion method like in [Xander May's kernel](https://www.kaggle.com/xmaayy/converting-vector-images-to-binary-images). It's fast and memory efficient and should work well. I didn't get good results with classification though."
  }
}