{"cells":[{"metadata":{"_uuid":"e87c1f259b32d4b67c09cfbf51ad39abd5c55904"},"cell_type":"markdown","source":"## This Notebook Demonstrates:\n1. Reading the data in python, preparing it for analysis, and adjusting the labels to contain underscores\n2. The code that simplfies a Raw drawing to the Simplified drawing\n3. How to make a submission file with predictions in the required format"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore') # to suppress some matplotlib deprecation warnings\n\nimport ast\nimport math\n\n# Have you installed your own package in Kernels yet? \n# If you need to, you can use the \"Settings\" bar on the right to install `simplification`\nfrom simplification.cutil import simplify_coords\n\nimport matplotlib.pyplot as plt\nimport matplotlib.style as style\n\n%matplotlib inline\n%config InlineBackend.figure_format = 'retina'\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"77f7515fd4ad8b9ee8395ceaacea4895f46d6e28"},"cell_type":"markdown","source":"### Let's read some of the training data"},{"metadata":{"trusted":true,"_uuid":"87b728e5a92ac1939ab9c0a2ba5bd554b1f1ddb4"},"cell_type":"code","source":"data = pd.read_csv('../input/train_simplified/roller coaster.csv',\n                   index_col='key_id',\n                   nrows=100)\ndata.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"14bf565221faad9702137b302f61e97d23623eea"},"cell_type":"markdown","source":"### Fixing labels\n\nNotice the `word` values for this label have a space. This aligns with the original Quick Draw data set that is already public. The Kaggle metric code for MAP@K requires each label to be separated by a single space. It parses \"roller coaster\" as two labels, \"roller\" and \"coaster\". Thus, `word` values with spaces need to be updated to use underscores. This is easily done, e.g.,"},{"metadata":{"trusted":true,"_uuid":"bf64c4a8edccce3dbe54c20fa6965c97cf5f2d90"},"cell_type":"code","source":"data['word'] = data['word'].replace(' ', '_', regex=True)\ndata.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"29e02e4643fc58d5e7410e2550a22d8ba21131f3"},"cell_type":"markdown","source":"### Let's look at some images\nWe're going to grab the first 10 images from the `test_raw.csv` file. Since the `word` values are read as a string, we need to convert them to a list using the `ast.literal_eval` function."},{"metadata":{"trusted":true,"_uuid":"f9afcea049ae88b7e6ad350652e794d5fc0b8207"},"cell_type":"code","source":"test_raw = pd.read_csv('../input/test_raw.csv', index_col='key_id')\nfirst_ten_ids = test_raw.iloc[:10].index\nraw_images = [ast.literal_eval(lst) for lst in test_raw.loc[first_ten_ids, 'drawing'].values]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8161f1f77f65af0c211e8ad4e44faf269ed5c2a8"},"cell_type":"markdown","source":"## From Raw to Simplified\n\nThis code demonstrates how the `simplified` data was generated from the `raw` data.\n\n(Code by Jonas Jongejan)"},{"metadata":{"trusted":true,"_uuid":"7e747bcff55c1c02ccba9488eb5af03cfcf62389"},"cell_type":"code","source":"def resample(x, y, spacing=1.0):\n    output = []\n    n = len(x)\n    px = x[0]\n    py = y[0]\n    cumlen = 0\n    pcumlen = 0\n    offset = 0\n    for i in range(1, n):\n        cx = x[i]\n        cy = y[i]\n        dx = cx - px\n        dy = cy - py\n        curlen = math.sqrt(dx*dx + dy*dy)\n        cumlen += curlen\n        while offset < cumlen:\n            t = (offset - pcumlen) / curlen\n            invt = 1 - t\n            tx = px * invt + cx * t\n            ty = py * invt + cy * t\n            output.append((tx, ty))\n            offset += spacing\n        pcumlen = cumlen\n        px = cx\n        py = cy\n    output.append((x[-1], y[-1]))\n    return output\n  \ndef normalize_resample_simplify(strokes, epsilon=1.0, resample_spacing=1.0):\n    if len(strokes) == 0:\n        raise ValueError('empty image')\n\n    # find min and max\n    amin = None\n    amax = None\n    for x, y, _ in strokes:\n        cur_min = [np.min(x), np.min(y)]\n        cur_max = [np.max(x), np.max(y)]\n        amin = cur_min if amin is None else np.min([amin, cur_min], axis=0)\n        amax = cur_max if amax is None else np.max([amax, cur_max], axis=0)\n\n    # drop any drawings that are linear along one axis\n    arange = np.array(amax) - np.array(amin)\n    if np.min(arange) == 0:\n        raise ValueError('bad range of values')\n\n    arange = np.max(arange)\n    output = []\n    for x, y, _ in strokes:\n        xy = np.array([x, y], dtype=float).T\n        xy -= amin\n        xy *= 255.\n        xy /= arange\n        resampled = resample(xy[:, 0], xy[:, 1], resample_spacing)\n        simplified = simplify_coords(resampled, epsilon)\n        xy = np.around(simplified).astype(np.uint8)\n        output.append(xy.T.tolist())\n\n    return output","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f5a71af1e2207cb137d58d3c5c907f374e92461f"},"cell_type":"code","source":"simplified_drawings = []\nfor drawing in raw_images:\n    simplified_drawing = normalize_resample_simplify(drawing)\n    simplified_drawings.append(simplified_drawing)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8450231b569efa269dda20ee10cf8b7730077c03"},"cell_type":"markdown","source":"## Viewing  Drawings\nAren't these fun to plot?!?!"},{"metadata":{"trusted":true,"_uuid":"f3cd12fa58226fd081e710f1909db45c140af5a5"},"cell_type":"code","source":"for index, raw_drawing in enumerate(raw_images, 0):\n    \n    plt.figure(figsize=(6,3))\n    \n    for x,y,t in raw_drawing:\n        plt.subplot(1,2,1)\n        plt.plot(x, y, marker='.')\n        plt.axis('off')\n\n    plt.gca().invert_yaxis()\n    plt.axis('equal')\n\n    for x,y in simplified_drawings[index]:\n        plt.subplot(1,2,2)\n        plt.plot(x, y, marker='.')\n        plt.axis('off')\n\n    plt.gca().invert_yaxis()\n    plt.axis('equal')\n    plt.show()  ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8dd9622a3f004ecc4b422d39dacd5f1f3302396a"},"cell_type":"markdown","source":"## Making a Submission"},{"metadata":{"trusted":true,"_uuid":"3d52b0bd272556f6da8f021b19350b7b050a249c"},"cell_type":"code","source":"submission = pd.read_csv('../input/sample_submission.csv', index_col='key_id')\n# Don't forget, your multi-word labels need underscores instead of spaces!\nmy_favorite_words = ['donut', 'roller_coaster', 'smiley_face']  \nsubmission['word'] = \" \".join(my_favorite_words)\nsubmission.to_csv('my_favorite_words.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b05834b538958d2fa4d32e4e77f09473b81804b1"},"cell_type":"code","source":"submission.head()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}