{"cells":[{"metadata":{"_uuid":"8c05de53a72b8ffdf89a9ce56dbcb5a11e589e9e"},"cell_type":"markdown","source":"# Making Binary Images"},{"metadata":{"_uuid":"23f01a3f9d8e7efd017b4d77bd77a14630915b3a"},"cell_type":"markdown","source":"By using this you need to understand that you are throwing away data about how the user drew the image, you're going to get thin lines, that might not be the best with larget kernel sizes. Anyway, heres how I'm suggesting you can process these images from their given format into binary images"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport re\nfrom PIL import Image\nimport matplotlib.pyplot as plt\nimport os","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"draw_type=\"airplane\"\nout_dir = \"./images/{0}/\".format(draw_type)\n\ndata = pd.read_csv(\"../input/train_simplified/{0}.csv\".format(draw_type))\ndata.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5d31081d6831849a9fe81106c2be2e74937386db"},"cell_type":"markdown","source":"### Creating a line\nThe gist of it is that you pass it two tuples - the start and end points of a line - and you get back all the intermediate points on that line\n\nTaken from here for the sake of speed\nhttp://www.roguebasin.com/index.php?title=Bresenham%27s_Line_Algorithm#Python"},{"metadata":{"trusted":true,"_uuid":"10aabb1a0178453b8c9e463d5f33fc9ac4e7e0e7"},"cell_type":"code","source":"def get_line(start, end):\n    \"\"\"Bresenham's Line Algorithm\n    Produces a list of tuples from start and end\n \n    >>> points1 = get_line((0, 0), (3, 4))\n    >>> points2 = get_line((3, 4), (0, 0))\n    >>> assert(set(points1) == set(points2))\n    >>> print points1\n    [(0, 0), (1, 1), (1, 2), (2, 3), (3, 4)]\n    >>> print points2\n    [(3, 4), (2, 3), (1, 2), (1, 1), (0, 0)]\n    \"\"\"\n    # Setup initial conditions\n    x1, y1 = start\n    x2, y2 = end\n    dx = x2 - x1\n    dy = y2 - y1\n \n    # Determine how steep the line is\n    is_steep = abs(dy) > abs(dx)\n \n    # Rotate line\n    if is_steep:\n        x1, y1 = y1, x1\n        x2, y2 = y2, x2\n \n    # Swap start and end points if necessary and store swap state\n    swapped = False\n    if x1 > x2:\n        x1, x2 = x2, x1\n        y1, y2 = y2, y1\n        swapped = True\n \n    # Recalculate differentials\n    dx = x2 - x1\n    dy = y2 - y1\n \n    # Calculate error\n    error = int(dx / 2.0)\n    ystep = 1 if y1 < y2 else -1\n \n    # Iterate over bounding box generating points between start and end\n    y = y1\n    points = []\n    for x in range(x1, x2 + 1):\n        coord = (y, x) if is_steep else (x, y)\n        points.append(coord)\n        error -= abs(dy)\n        if error < 0:\n            y += ystep\n            error += dx\n \n    # Reverse the list if the coordinates were swapped\n    if swapped:\n        points.reverse()\n    return points","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5e6996732a4213ae2483f83e97b687a414fac6ef"},"cell_type":"markdown","source":"## The magic\nHeres where we render a binary image for each vector image in the specified csv file"},{"metadata":{"trusted":true,"_uuid":"aa44a51643740a926be29d9c7c40cd944fb8e032"},"cell_type":"code","source":"# First get the drawing data, of course\ndrawings = data[\"drawing\"].values\ndrawing_ids = data[\"key_id\"].values\n\n# Make a new output directory for the images\ntry:\n    os.makedirs(out_dir)\nexcept:\n    pass\n\n# This is set to range(1) so you can see what an output image looks like, but to actually run it change it\n# to :                                      which will iterate though all the drawings in the specified csv\n#               range(len(drawings))\nfor draw_idx in range(1):\n    # Next get it out of the nasty 'line by line' format that actually contains useful data about the drawing and\n    # just make it into a large array with regex :)\n    coords = re.findall(r\"\\[[^\\[\\]]+\\]\", drawings[draw_idx])\n\n    # Again split off each line into its co-ordinates with more string ops\n    sets = []\n    for co_set in coords: \n        sets.append(np.int_(co_set.strip('[ ]').split(',')))\n\n    # Initialize\n    image = np.zeros([255,255])\n    pixels = []\n    endpair = []\n    \n    # While we still have sets of co-ordinates left\n    while len(sets) > 0:\n        x = sets.pop() # Get the X set\n        y = sets.pop() # Get the Y set\n        \n        # Put it into [x1, y1] form because thats how I like it\n        #             [x2, y2] \n        pairs = np.hstack([np.transpose(x).reshape(len(x),1), np.transpose(y).reshape(len(y),1)])\n        \n        # You could heavily optimize this by running all the line generation in parallel but thats\n        # out of the scope of this terrible kernel\n        for i in range(len(pairs)-1):\n            pixels.extend(get_line(pairs[i],pairs[i+1]))\n            \n    # Set each pixel in the image to 1 if its one of the co-ordintates \n    # I'm not yet well versed enough in python to do this properly, so again here is a place you could optimize\n    for pixel in pixels:\n        image[pixel[0]-1,pixel[1]-1]=1\n\n    # The fruit of our labor\n    fig = plt.imshow(image, cmap='binary')\n    \n    # And if you want to save it and use all these as pre-processed images for a network\n    plt.imsave(os.path.join(out_dir,'{0}.png'.format(drawing_ids[draw_idx])), image, cmap='binary')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a9a25b48a1ffa21e03e53a534f7a9caadc980cc9"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}