{"cells":[{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"import json\nfrom functools import partial\nimport ast\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n%matplotlib inline\nfrom matplotlib import collections  as mc\nimport seaborn as sns\n\nimport warnings\nwarnings.filterwarnings('ignore')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6e99cefa7d7f7a3be40044bb3fde754fc7e17f18"},"cell_type":"markdown","source":"When asked to draw a simple shape, do you usually draw it clockwise or counterclockwise? Does it depend on the shape in question? Does it depend on your cultural background? These are the questions we rarely ask ourselves, but can potentially lead to interesting discoveries. Maybe before reading this notebook, try this yourself. Draw a circle, square and hexagon each, and see instinctively, which direction your hand is going, and then find out below if you draw the same way the majority of people from your region do.\n\nTo determine how people from different parts of the world draw simple shapes, we will be using the QuickDraw dataset, which is a collection of various simple shapes people drew online when asked to do a sketch of a given object. Here we are most interested in how people draw the most basic shapes (circle, square, hexagon). The dataset includes the country code of the user as an attribute, allowing us to perform some geography / culture-based analysis on our results.\n\nThe goal of our analysis is to find out if there are any differences in the direction people draw simple shapes around the world. To do so, we will have to come up with a measure of 'counterclockwiseness' for each of the sktches in the dataset. But first, let us look at in what format the drawings appear in the dataset. Let us load the hexagon drawings first:"},{"metadata":{"trusted":true,"_uuid":"9bf3d4e249a1a846cdc48fb831036de7bbf06685"},"cell_type":"code","source":"hex_df = pd.read_csv('../input/train_simplified/hexagon.csv')\nhex_df['drawing'] = hex_df['drawing'].apply(ast.literal_eval)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d048058a635625458e531fe4135398dcb084d53f"},"cell_type":"code","source":"hex_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ed3dd162cf22bf2c30b4834a48aeca7c40116ea0"},"cell_type":"markdown","source":"We see from below that the drawings are stored as lists of continous strokes, where each stroke is further broken down into line segments and stored as [(x1, x2, x3,...), (y1, y2, y3...)]. We have to remember that image coordinates are different from typical coordinate systems in that the y axis points downwards (so (1, 1) is at the top left corner not bottom left corner). It is easier to process and visualise the data if we convert it to the ordinary coordinate format by setting y = 255 - y."},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"fc704ae61e5a4826493f6e24c2ee15fc465a78d7"},"cell_type":"code","source":"hex_df.loc[0, 'drawing']","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1530eebaf0fe44f8070a6f3fe091247ca62beba0"},"cell_type":"markdown","source":"## Hexagons:"},{"metadata":{"_uuid":"b83987030194a91d4781bb6a31d32b965b28a468"},"cell_type":"markdown","source":"Let us first visualise a few examples from the hexagon drawings to see what a typical sketch looks like."},{"metadata":{"trusted":true,"_uuid":"ef4e97ccaf5b9faf6b5b9ac9318d62b37eea0f1a"},"cell_type":"code","source":"def to_line_collection(stroke):\n    points = list(zip(stroke[0], [255 - x for x in stroke[1]]))\n    lc = [[points[i], points[i + 1]] for i in range(len(points) - 1)]\n    return mc.LineCollection(lc, linewidth=2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"14597636f66fc4eaae5c2e6b5069e2b7b9ec626a"},"cell_type":"code","source":"def visualise_drawing(drawing, ax):\n    for stroke in drawing:\n        ax.add_collection(to_line_collection(stroke))\n    ax.autoscale()\n    return ax","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"35344c17ddca2167cd93e2f82ece2150058a6b83"},"cell_type":"code","source":"f, ax = plt.subplots(2, 5, figsize=(16, 6))\nfor i in range(10):\n    visualise_drawing(hex_df.loc[i, 'drawing'], ax=ax[i//5, i%5])\n    ax[i//5, i%5].axis('off')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"011e577171b5916010e0ff4cc6b2ea99406bf01c"},"cell_type":"markdown","source":"#### Define clockwise / counterclockwise stroke\nNow we need to find a measure for 'counterclockwiseness'. How do we analyse something so abstract? Intuitively, we know that something goes clockwise if it keeps turning right, and counterclockwise if it keeps going left. Therefore, it would be a good start to compare each segment in a stroke to see if it is going left or right compared to the last segment. To do so, we need to separate segments from a stroke and compare the angles between consecutive segments. However, the angle alone would likely not help us much, as we will then only be counting how many 'loops' there are in the curves drawn, regardless of whether it is a tiny turn, a minor jitter or an entire cirlce around the drawing area. We need to take into account the relative lengths of each line segments. For each turn, we need to know how much it turned left or right compared to the last line segment. The sine of the turn angle is a good candidate for this, as it can be interpreted as the sideways component of the second segment divided by the length of the first segment. It is still not perfect (e.g. a small turn has the same score as a large turn if the length proportions are the same), but it has a few nice properties, such as:\n\n* It can be both positive and negative, and when summed up, equivalent turns in opposite directions cancel each other out\n* It is resistant to tiny jitters, so a brief change in stroke direction will not greatly impact the net score\n* It is bounded between -1 and 1, so one single turn cannot dominate the score of the whole drawing\n\nTo put it simply, **postitive score -> counterclockwise preference, negative score -> clockwise preference**.\n\nHere we calculate the sine score for each turn in the drawings, then sum up the scores by drawings."},{"metadata":{"trusted":true,"_uuid":"667e37528181f234c171e53c12b7fb1d4da8eda6"},"cell_type":"code","source":"def invert_y(strokes):\n    strokes[:, 1] = 255 - strokes[:, 1]\n    return strokes\n\ndef decompose_drawing(drawing):\n    strokes = [invert_y(np.array(stroke).T) for stroke in drawing]\n    return strokes","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"939cee54ee4c7e78d50f92259859542c1559b5b4"},"cell_type":"code","source":"def relative_turn_distance(three_points):\n    vector_1 = three_points[1, :] - three_points[0, :]\n    vector_2 = three_points[2, :] - three_points[1, :]\n    distance = np.cross(vector_1, vector_2) / (np.linalg.norm(vector_1) * np.linalg.norm(vector_2))\n    return distance if not np.isnan(distance) else 0.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d8c6bfe4c05c1e962077b0736180e0d3acb2c59f"},"cell_type":"code","source":"def stroke_relative_turn_distance(stroke):\n    if stroke.shape[0] < 3:\n        distance = 0\n    else:\n        distance = np.sum([relative_turn_distance(stroke[start: start + 3, :]) for start in range(0, stroke.shape[0] - 2)])\n    return distance","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1b12c2757db5b7f67290b665080a900819796e18"},"cell_type":"code","source":"def drawing_relative_turn_distance(drawing, connect=False):\n    strokes = decompose_drawing(drawing)\n    if connect:\n        distance = stroke_relative_turn_distance(np.concatenate(strokes, axis=0))\n    else:\n        distance = np.sum([stroke_relative_turn_distance(stroke) for stroke in strokes])\n    return distance","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d953d73fd3a9e92d98699fd107bd1902e06ff9a4"},"cell_type":"markdown","source":"When performing the sum of scores for each drawing, we can treat each strokes separately or treat then as if they were chained to each other head to tail (so the change in position from stroke 1 to stroke 2 is also considered a line segment for the purpose of the analysis). We will mainly be doing our analysis using the 'chained' approach, but the analysis using 'per stoke' produces similar results."},{"metadata":{"trusted":true,"_uuid":"a8c0b81886914320a02726f8d0232dd9b539ae2e"},"cell_type":"code","source":"# per_stroke_score = hex_df['drawing'].apply(partial(drawing_relative_turn_distance, connect=False))\nper_drawing_score = hex_df['drawing'].apply(partial(drawing_relative_turn_distance, connect=True))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"96ed47e47384b50378c07732b5c1afb5e422699d"},"cell_type":"code","source":"# hex_df['sum_per_stroke_score'] = per_stroke_score\nhex_df['score'] = per_drawing_score","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"de0907af456fddcda32b424585a94d5c0cd78797"},"cell_type":"code","source":"hex_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"50c344b1ef7280fe4866d9b4d02a7df93c0018b7"},"cell_type":"markdown","source":"Let us first look at the most 'unusual' drawings based on the score, from both end of the scale:"},{"metadata":{"trusted":true,"_uuid":"45796cd0e2b9238027564b8a3a4deac6f1eba59e"},"cell_type":"code","source":"f, ax = plt.subplots(2, 5, figsize=(16, 6))\nordered_subset = hex_df.sort_values('score').iloc[:5, :]\nfor i, drawing in enumerate(ordered_subset['drawing']):\n    visualise_drawing(drawing, ax=ax[0, i])\n    ax[0, i].axis('off')\nordered_subset = hex_df.sort_values('score').iloc[-5:, :]\nfor i, drawing in enumerate(ordered_subset['drawing']):\n    visualise_drawing(drawing, ax=ax[1, i])\n    ax[1, i].axis('off')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"27645d6aea95bbb5e4f06db02aae1b5b4371ae7c"},"cell_type":"markdown","source":"As we see, the most extreme cases are typically random drawings that do not make much sense, so it is likely safe to exclude the most extreme cases in our analysis later. As we see below, most of the score values fall between -5 and 5."},{"metadata":{"trusted":true,"_uuid":"5683a8f2159eba95806a031fad9c9e60bbc2de09"},"cell_type":"code","source":"hex_df.describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ff90c28c56b57dfc89ae99d2ae5f1285add8aafb"},"cell_type":"markdown","source":"Now let us look at the drawings that are closest to the median score of the drawings. As we see below, having a median score for the 'counterclockwiseness' does not necessarily mean having a drawing closest to a standard shape."},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"d5141dd5a9500e4c88b5e2880c616250ff38c2e5"},"cell_type":"code","source":"f, ax = plt.subplots(2, 5, figsize=(16, 6))\ntemp = hex_df.copy()\ntemp['dev'] = np.abs(temp['score'] - temp['score'].median())\nordered_subset = temp.sort_values('dev', ascending=True).iloc[:10, :]\nfor i, drawing in enumerate(ordered_subset['drawing']):\n    visualise_drawing(drawing, ax=ax[i//5, i%5])\n    ax[i//5, i%5].axis('off')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7ed253c114ee75758ac5a00bdcafa7911b8197f0"},"cell_type":"markdown","source":"Let us get back to our original question. Can we expect to observe different patterns from different parts of the world in terms of drawing counterclockwiseness?"},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"_uuid":"73b205dc62bafd2f4bea2a4f43e9ba1bd498751d"},"cell_type":"code","source":"country_codes = '''\nCountry Name;ISO 3166-1-alpha-2 code\nAFGHANISTAN;AF\nÅLAND ISLANDS;AX\nALBANIA;AL\nALGERIA;DZ\nAMERICAN SAMOA;AS\nANDORRA;AD\nANGOLA;AO\nANGUILLA;AI\nANTARCTICA;AQ\nANTIGUA AND BARBUDA;AG\nARGENTINA;AR\nARMENIA;AM\nARUBA;AW\nAUSTRALIA;AU\nAUSTRIA;AT\nAZERBAIJAN;AZ\nBAHAMAS;BS\nBAHRAIN;BH\nBANGLADESH;BD\nBARBADOS;BB\nBELARUS;BY\nBELGIUM;BE\nBELIZE;BZ\nBENIN;BJ\nBERMUDA;BM\nBHUTAN;BT\nBOLIVIA, PLURINATIONAL STATE OF;BO\nBONAIRE, SINT EUSTATIUS AND SABA;BQ\nBOSNIA AND HERZEGOVINA;BA\nBOTSWANA;BW\nBOUVET ISLAND;BV\nBRAZIL;BR\nBRITISH INDIAN OCEAN TERRITORY;IO\nBRUNEI DARUSSALAM;BN\nBULGARIA;BG\nBURKINA FASO;BF\nBURUNDI;BI\nCAMBODIA;KH\nCAMEROON;CM\nCANADA;CA\nCAPE VERDE;CV\nCAYMAN ISLANDS;KY\nCENTRAL AFRICAN REPUBLIC;CF\nCHAD;TD\nCHILE;CL\nCHINA;CN\nCHRISTMAS ISLAND;CX\nCOCOS (KEELING) ISLANDS;CC\nCOLOMBIA;CO\nCOMOROS;KM\nCONGO;CG\nCONGO, THE DEMOCRATIC REPUBLIC OF THE;CD\nCOOK ISLANDS;CK\nCOSTA RICA;CR\nCÔTE D'IVOIRE;CI\nCROATIA;HR\nCUBA;CU\nCURAÇAO;CW\nCYPRUS;CY\nCZECH REPUBLIC;CZ\nDENMARK;DK\nDJIBOUTI;DJ\nDOMINICA;DM\nDOMINICAN REPUBLIC;DO\nECUADOR;EC\nEGYPT;EG\nEL SALVADOR;SV\nEQUATORIAL GUINEA;GQ\nERITREA;ER\nESTONIA;EE\nETHIOPIA;ET\nFALKLAND ISLANDS (MALVINAS);FK\nFAROE ISLANDS;FO\nFIJI;FJ\nFINLAND;FI\nFRANCE;FR\nFRENCH GUIANA;GF\nFRENCH POLYNESIA;PF\nFRENCH SOUTHERN TERRITORIES;TF\nGABON;GA\nGAMBIA;GM\nGEORGIA;GE\nGERMANY;DE\nGHANA;GH\nGIBRALTAR;GI\nGREECE;GR\nGREENLAND;GL\nGRENADA;GD\nGUADELOUPE;GP\nGUAM;GU\nGUATEMALA;GT\nGUERNSEY;GG\nGUINEA;GN\nGUINEA-BISSAU;GW\nGUYANA;GY\nHAITI;HT\nHEARD ISLAND AND MCDONALD ISLANDS;HM\nHOLY SEE (VATICAN CITY STATE);VA\nHONDURAS;HN\nHONG KONG;HK\nHUNGARY;HU\nICELAND;IS\nINDIA;IN\nINDONESIA;ID\nIRAN, ISLAMIC REPUBLIC OF;IR\nIRAQ;IQ\nIRELAND;IE\nISLE OF MAN;IM\nISRAEL;IL\nITALY;IT\nJAMAICA;JM\nJAPAN;JP\nJERSEY;JE\nJORDAN;JO\nKAZAKHSTAN;KZ\nKENYA;KE\nKIRIBATI;KI\nKOREA, DEMOCRATIC PEOPLE'S REPUBLIC OF;KP\nKOREA, REPUBLIC OF;KR\nKUWAIT;KW\nKYRGYZSTAN;KG\nLAO PEOPLE'S DEMOCRATIC REPUBLIC;LA\nLATVIA;LV\nLEBANON;LB\nLESOTHO;LS\nLIBERIA;LR\nLIBYA;LY\nLIECHTENSTEIN;LI\nLITHUANIA;LT\nLUXEMBOURG;LU\nMACAO;MO\nMACEDONIA, THE FORMER YUGOSLAV REPUBLIC OF;MK\nMADAGASCAR;MG\nMALAWI;MW\nMALAYSIA;MY\nMALDIVES;MV\nMALI;ML\nMALTA;MT\nMARSHALL ISLANDS;MH\nMARTINIQUE;MQ\nMAURITANIA;MR\nMAURITIUS;MU\nMAYOTTE;YT\nMEXICO;MX\nMICRONESIA, FEDERATED STATES OF;FM\nMOLDOVA, REPUBLIC OF;MD\nMONACO;MC\nMONGOLIA;MN\nMONTENEGRO;ME\nMONTSERRAT;MS\nMOROCCO;MA\nMOZAMBIQUE;MZ\nMYANMAR;MM\nNAMIBIA;NA\nNAURU;NR\nNEPAL;NP\nNETHERLANDS;NL\nNEW CALEDONIA;NC\nNEW ZEALAND;NZ\nNICARAGUA;NI\nNIGER;NE\nNIGERIA;NG\nNIUE;NU\nNORFOLK ISLAND;NF\nNORTHERN MARIANA ISLANDS;MP\nNORWAY;NO\nOMAN;OM\nPAKISTAN;PK\nPALAU;PW\nPALESTINE, STATE OF;PS\nPANAMA;PA\nPAPUA NEW GUINEA;PG\nPARAGUAY;PY\nPERU;PE\nPHILIPPINES;PH\nPITCAIRN;PN\nPOLAND;PL\nPORTUGAL;PT\nPUERTO RICO;PR\nQATAR;QA\nRÉUNION;RE\nROMANIA;RO\nRUSSIAN FEDERATION;RU\nRWANDA;RW\nSAINT BARTHÉLEMY;BL\nSAINT HELENA, ASCENSION AND TRISTAN DA CUNHA;SH\nSAINT KITTS AND NEVIS;KN\nSAINT LUCIA;LC\nSAINT MARTIN (FRENCH PART);MF\nSAINT PIERRE AND MIQUELON;PM\nSAINT VINCENT AND THE GRENADINES;VC\nSAMOA;WS\nSAN MARINO;SM\nSAO TOME AND PRINCIPE;ST\nSAUDI ARABIA;SA\nSENEGAL;SN\nSERBIA;RS\nSEYCHELLES;SC\nSIERRA LEONE;SL\nSINGAPORE;SG\nSINT MAARTEN (DUTCH PART);SX\nSLOVAKIA;SK\nSLOVENIA;SI\nSOLOMON ISLANDS;SB\nSOMALIA;SO\nSOUTH AFRICA;ZA\nSOUTH GEORGIA AND THE SOUTH SANDWICH ISLANDS;GS\nSOUTH SUDAN;SS\nSPAIN;ES\nSRI LANKA;LK\nSUDAN;SD\nSURINAME;SR\nSVALBARD AND JAN MAYEN;SJ\nSWAZILAND;SZ\nSWEDEN;SE\nSWITZERLAND;CH\nSYRIAN ARAB REPUBLIC;SY\nTAIWAN;TW\nTAJIKISTAN;TJ\nTANZANIA, UNITED REPUBLIC OF;TZ\nTHAILAND;TH\nTIMOR-LESTE;TL\nTOGO;TG\nTOKELAU;TK\nTONGA;TO\nTRINIDAD AND TOBAGO;TT\nTUNISIA;TN\nTURKEY;TR\nTURKMENISTAN;TM\nTURKS AND CAICOS ISLANDS;TC\nTUVALU;TV\nUGANDA;UG\nUKRAINE;UA\nUNITED ARAB EMIRATES;AE\nUNITED KINGDOM;GB\nUNITED STATES;US\nUNITED STATES MINOR OUTLYING ISLANDS;UM\nURUGUAY;UY\nUZBEKISTAN;UZ\nVANUATU;VU\nVENEZUELA, BOLIVARIAN REPUBLIC OF;VE\nVIET NAM;VN\nVIRGIN ISLANDS, BRITISH;VG\nVIRGIN ISLANDS, U.S.;VI\nWALLIS AND FUTUNA;WF\nWESTERN SAHARA;EH\nYEMEN;YE\nZAMBIA;ZM\nZIMBABWE;ZW\n'''\nfrom io import StringIO\ncountry_codes_df = pd.read_csv(StringIO(country_codes), sep=';')\ncountry_codes_df.columns = ['countryname', 'countrycode']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e4596b3ef1f49af1c096aeac22787114cc24a805"},"cell_type":"code","source":"country_codes_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"83b0ab6c9b00568a5acb40cfb1884e425b432887"},"cell_type":"code","source":"hex_df = pd.merge(hex_df, country_codes_df, on='countrycode', how='left')\ntop_countries = hex_df['countrycode'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"88e846cd9a9140aa873207b34510960ca6cc0335"},"cell_type":"code","source":"f, ax = plt.subplots(figsize=(14, 14))\nsubset = hex_df[hex_df['countrycode'].isin(top_countries.index[:40]) & (np.abs(hex_df['score'] < 10))]\nsub_order = subset.groupby('countryname')['score'].mean().sort_values().index\nsns.barplot(data=subset, y='countryname', x='score', order=sub_order)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6430edc52c43b0439e9fda6a0f6bf96498bdffa8"},"cell_type":"markdown","source":"What a surprise! While we see that most user around the world draw a hexagon clockwise (negative score), the average user from Japan actually draws the shape counterclockwise. Also, two nearby countries, Thailand and Viet Nam, share the top two spots on the clockwiseness scale. We do havr to take into account sample size though, as these two countries' score have a rather high variance."},{"metadata":{"_uuid":"601ebbe4caec7100d668d10e1443c10a753844e0"},"cell_type":"markdown","source":"## Circles:"},{"metadata":{"trusted":true,"_uuid":"8a39ed41fdcabaeeead0e5b121130c7c4ce818bb"},"cell_type":"code","source":"circle_df = pd.read_csv('../input/train_simplified/circle.csv')\ncircle_df['drawing'] = circle_df['drawing'].apply(ast.literal_eval)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f5e5af0b979a507eb3335063a0388098668d7cdf"},"cell_type":"code","source":"f, ax = plt.subplots(2, 5, figsize=(16, 6))\nfor i in range(10):\n    visualise_drawing(circle_df.loc[i, 'drawing'], ax=ax[i//5, i%5])\n    ax[i//5, i%5].axis('off')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"47c63ba8172a9c7ce8e17ba6f61ac1b64cb95377"},"cell_type":"code","source":"circle_df['score'] = circle_df['drawing'].apply(partial(drawing_relative_turn_distance, connect=True))\ncircle_df = pd.merge(circle_df, country_codes_df, on='countrycode', how='left')\ntop_countries = circle_df['countrycode'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a04eabd9ea63476130b6e5ecbf097e6135e4766a"},"cell_type":"code","source":"f, ax = plt.subplots(figsize=(14, 14))\nsubset = circle_df[circle_df['countrycode'].isin(top_countries.index[:40]) & (np.abs(circle_df['score'] < 10))]\nsub_order = subset.groupby('countryname')['score'].mean().sort_values().index\nsns.barplot(data=subset, y='countryname', x='score', order=sub_order)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3d51ea924ca65d357f4decf7e5168f29417bdc0a"},"cell_type":"markdown","source":"Another purprise! This time, most users around the world tend to draw circles counterclockwise, except Taiwan and (again!) Japan, whose user typically draw it clockwise. We noticed that even for simple shapes like hexagon and circles, the type of the shape still has an impact on the usual direction people draw them. What about other polygons? Are they more like hexagons or are they more like circles? Which of the two shapes above is the norm and which is the exception, or is it just different for each individual shape?\n\nWe also learned that at least for the two tasks above, the Japanese users always do things differently. I wonder why."},{"metadata":{"trusted":true,"_uuid":"194b2d2ed8f98753e95aed5f9b3bbc035f8b0062"},"cell_type":"markdown","source":"## Squares:"},{"metadata":{"trusted":true,"_uuid":"82d250e34106f0ee4b3aceb3031bf96087eb6e0e"},"cell_type":"code","source":"square_df = pd.read_csv('../input/train_simplified/square.csv')\nsquare_df['drawing'] = square_df['drawing'].apply(ast.literal_eval)\n\nf, ax = plt.subplots(2, 5, figsize=(16, 6))\nfor i in range(10):\n    visualise_drawing(square_df.loc[i, 'drawing'], ax=ax[i//5, i%5])\n    ax[i//5, i%5].axis('off')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c1dd3a525765941805d30109aac5c59da26d953d"},"cell_type":"code","source":"square_df['score'] = square_df['drawing'].apply(partial(drawing_relative_turn_distance, connect=True))\nsquare_df = pd.merge(square_df, country_codes_df, on='countrycode', how='left')\ntop_countries = square_df['countrycode'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"24ee0df0959667cdd63eec1ea03bad09613d699a"},"cell_type":"code","source":"f, ax = plt.subplots(figsize=(14, 14))\nsubset = square_df[square_df['countrycode'].isin(top_countries.index[:40]) & (np.abs(square_df['score'] < 10))]\nsub_order = subset.groupby('countryname')['score'].mean().sort_values().index\nsns.barplot(data=subset, y='countryname', x='score', order=sub_order)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d81d91e2d9f8a60b0a9e1b875140a50ac07aaae6"},"cell_type":"markdown","source":"Here we see the results for squares. Japan is not the odd one out this time, but we see an interesting pattern here. While most users draw a square counterclockwise, people from European countries are more likely to do so, whereas shapes drawn by Asian users (including Indian and Middle Eastern) seem less likely so. Maybe people from different parts of the world just have specific preferences for each of the common shapes. What about a less common shape like an octagon?"},{"metadata":{"_uuid":"5468157b1ff596b57dcf4ce41b463f667cb93ca9"},"cell_type":"markdown","source":"## Octagon:"},{"metadata":{"trusted":true,"_uuid":"c95c3619431cb839693ccf4d39750d4fe9badcf5"},"cell_type":"code","source":"octagon_df = pd.read_csv('../input/train_simplified/octagon.csv')\noctagon_df['drawing'] = octagon_df['drawing'].apply(ast.literal_eval)\n\nf, ax = plt.subplots(2, 5, figsize=(16, 6))\nfor i in range(10):\n    visualise_drawing(octagon_df.loc[i, 'drawing'], ax=ax[i//5, i%5])\n    ax[i//5, i%5].axis('off')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5ce8a4912c2628ab7df46a16c29600ce12ab7762"},"cell_type":"markdown","source":"Look at the samples above. Looks like quite a few users have no idea what an octagon looks like!"},{"metadata":{"trusted":true,"_uuid":"a26e74044402d586e2617070bb30e8f814fd0b8f"},"cell_type":"code","source":"octagon_df['score'] = octagon_df['drawing'].apply(partial(drawing_relative_turn_distance, connect=False))\noctagon_df = pd.merge(octagon_df, country_codes_df, on='countrycode', how='left')\ntop_countries = octagon_df['countrycode'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b743b855821591cb2073d0978ca2e872b83608f9"},"cell_type":"code","source":"f, ax = plt.subplots(figsize=(14, 14))\nsubset = octagon_df[octagon_df['countrycode'].isin(top_countries.index[:40]) & (np.abs(octagon_df['score'] < 10))]\nsub_order = subset.groupby('countryname')['score'].mean().sort_values().index\nsns.barplot(data=subset, y='countryname', x='score', order=sub_order)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"99ae5c75dbc5a3b1f33c171f05ca33f8176c439b"},"cell_type":"markdown","source":"We see a much more random pattern here. It looks like the strong preferences only exist for simpler shapes that are drawn often. When it comes to less common shapes, the direction of drawing is less determined by the user's location and the variance in personal perferences become larger."},{"metadata":{"_uuid":"0240e681b35b7aa5737c1152af02ffd947224255"},"cell_type":"markdown","source":"## Finally, let us compre the four shapes:"},{"metadata":{"trusted":true,"_uuid":"d02d82dcf08471597498ab5a871f58488169226a"},"cell_type":"code","source":"dfs = [square_df, hex_df, octagon_df, circle_df]\ncombined = pd.concat(dfs, axis=0, sort=False)\ncombined['n_segs'] = combined['drawing'].apply(lambda x: np.sum([len(s[0]) for s in x]))\ncombined['per_seg_score'] = combined['score'] / combined['n_segs']\ncombined.groupby('word')['n_segs'].describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1a1ab20a409ac2557b744e0d9c93b22b9c4421da"},"cell_type":"code","source":"combined.groupby('word')['score'].describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f8affbfaba053ea414604bf8bde140767616079d"},"cell_type":"code","source":"f, ax = plt.subplots(figsize=(14, 10))\nfor word in combined['word'].unique():\n    sns.kdeplot(data=combined[(combined['word'] == word) & (np.abs(combined['score']) < 10)]['score'], ax=ax, label=word)\nax.set_xlabel('counterclockwiseness score')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5ac2551421b8d22c865446cb9226c42520fd2ac5"},"cell_type":"markdown","source":"We see the distributions of counterclockwiseness for each of the shapes above. We see that for almost of them, there is a divide between the clockwise users vs counterclockwise users, as seen by the bimodal shape of the distributions. However, octagons, we actually see a *trimodal* pattern, with a significant number of drawings showing no particular direction preference. Do people actually draw from both sides and meet in the middle?\n\nThe distibutions of different shapes are slightly shifted away from each other due to the different number of segments for different shapes. If we adjust for number of segments, we see the distribution below:"},{"metadata":{"trusted":true,"_uuid":"edabe23e12855c3a14dbc8e1e7d362cd956c5ed6"},"cell_type":"code","source":"f, ax = plt.subplots(figsize=(14, 10))\nfor word in combined['word'].unique():\n    sns.kdeplot(data=combined[(combined['word'] == word) & (np.abs(combined['per_seg_score']) < 2)]['per_seg_score'], ax=ax, label=word)\nax.set_xlabel('counterclockwiseness score per segment')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"437efc98768296ea8e8240c5d721093623f50444"},"cell_type":"markdown","source":"Here we can more clearly see the bimodal / trimodal patterns. As most of the shapes drawn complete a full turn from start to finish, as expected, we see similar (absolute) values of the most typical score per segment for all of the shapes."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}