{"cells":[{"metadata":{"_uuid":"e3ddec1af184dc488f27bab93a2d40de91fcdd63"},"cell_type":"markdown","source":"# Country clustering with hdbscan\nIf you think that country info in Doodle competition is too specific and [continent](https://www.kaggle.com/qlasty/localization-context-country-continent) info is too general why not to try something in between and to cluster them in order to take advantage of possible cultural context. In this notebook simple clusterization based on localization is proposed."},{"metadata":{"trusted":true,"_uuid":"ddee66a3642c220c26f756e9ef7cd7ab84e7cfc8","_kg_hide-output":true,"_kg_hide-input":true},"cell_type":"code","source":"!pip install hdbscan\n!pip install pycountry-convert\n\nfrom geopy.geocoders import Nominatim\nimport pycountry_convert\nimport geopandas as gp\nimport pandas as pd\nimport numpy as np\nimport hdbscan","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"37e4b4fb23383afc23752d989e1f8ea987fcf6fa"},"cell_type":"markdown","source":"Take country names along with their 2 and 3 letter codes available in the pycountry_convert lib."},{"metadata":{"trusted":true,"_uuid":"1d411c3fc6cd2649f5374aa326558f5941df9c5b"},"cell_type":"code","source":"valid_countries_dict = pycountry_convert.map_countries(cn_name_format=\"default\")\nvalid_countries = [(k, v['alpha_2'], v['alpha_3']) for k, v in valid_countries_dict.items()]\ncountryDF = pd.DataFrame(valid_countries, columns=[\"country\",\"alpha2\",\"alpha3\"])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3514ddeb778f314ac6d3641b81b1fbb578349d05"},"cell_type":"markdown","source":"Most countries appear two times (probably their names are in two languages: official and English) so let's get rid of the duplicates."},{"metadata":{"trusted":true,"_uuid":"a96c9daa2b9de614fa7948b4026f897bef4a988a"},"cell_type":"code","source":"countryDF.drop_duplicates('alpha2', inplace=True)\ncountryDF.reset_index(inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d14638b388be22e1e178d159a21a89aea4430934"},"cell_type":"markdown","source":"Following four countries names need to be corrected manually in order to be recognized by the geopy lib:"},{"metadata":{"trusted":true,"_uuid":"563edf54a40435494a98442a90c41eb7cbabc218"},"cell_type":"code","source":"countryDF.loc[countryDF['alpha3']=='BOL','country']='Bolivia'\ncountryDF.loc[countryDF['alpha3']=='VAT','country']='Vatican'\ncountryDF.loc[countryDF['alpha3']=='TWN','country']='Taiwan'\ncountryDF.loc[countryDF['alpha3']=='VIR','country']='Virgin Islands'","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e9ef53a215516d6e6a46607b49b822d096e2f0f6"},"cell_type":"markdown","source":"Function for finding localization of a country, given its name. The returned latitude, longitude points the center of a country."},{"metadata":{"trusted":true,"_uuid":"bcec182a457e550073c00557ca067d4f3cf8db0a"},"cell_type":"code","source":"geoloc = Nominatim(user_agent=\"area_clustering\")\n\ndef getCoordinates(countryName):    \n    try:\n        info = geoloc.geocode(countryName)    \n        return [info.latitude, info.longitude]\n    except:\n        print('Error: Country {} not found'.format(countryName))        \n        return [0, 90]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e50a3aef9f79325cd403fc260c0635ed263bdd09","scrolled":true},"cell_type":"code","source":"coordinates = countryDF[\"country\"].apply(getCoordinates)\ncountryDF = pd.concat([countryDF, pd.DataFrame(coordinates.tolist(), columns=['latitude', 'longitude'])], axis=1)\ncountryDF[\"latitude\"] = countryDF[\"latitude\"].apply(np.radians)\ncountryDF[\"longitude\"] = countryDF[\"longitude\"].apply(np.radians)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8de6e07551869d9c5224018e3afc6f62094993e0"},"cell_type":"markdown","source":"Function, which groups countries based on their localization. \"Haversine\" metric is the one appropriate for clusterization having longitute/latitude info."},{"metadata":{"trusted":true,"_uuid":"781089044f4a32305c61ccdac5f880154615be4d"},"cell_type":"code","source":"def cluster_data(dataframe, min_cluster_size, min_samples):\n    clus = hdbscan.HDBSCAN(metric='haversine', min_cluster_size=min_cluster_size, min_samples=min_samples)    \n    dataframe.loc[:,'groups'] = clus.fit_predict(dataframe[[\"latitude\",\"longitude\"]])    \n    print('n_groups: {}, unclustered objects: {}'.format(max(dataframe['groups']), sum(dataframe['groups']==-1)))\n    return dataframe","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"744b368445b47103d7df16c91d37ee143f39b8ce"},"cell_type":"code","source":"countryDF = cluster_data(dataframe=countryDF, min_cluster_size=5, min_samples=3)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5ef8099d763d58ff9124165c902c7c03a7247054"},"cell_type":"markdown","source":"Select unclustered countries and make clustering again, this time with settings allowing for smaller groups."},{"metadata":{"trusted":true,"_uuid":"81e3d3dd486d2904c7dd8874aa6c6fb20e072205"},"cell_type":"code","source":"rec = countryDF.loc[countryDF['groups']==-1].copy()\nrec = cluster_data(dataframe=rec, min_cluster_size=3, min_samples=1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"292783b493278426e8ed447b2f80f46279c2162e"},"cell_type":"markdown","source":"Merge results from both clusterings and display groups quantities."},{"metadata":{"trusted":true,"_uuid":"c36326ff6a5e4fa07bc39543b813ffc373ec754b"},"cell_type":"code","source":"rec.loc[rec['groups']>=0,'groups']+=np.max(countryDF['groups']+1)\ncountryDF[countryDF['groups']==-1]=rec\ncountryDF['groups'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c7633a0ece4fd562d96276319d304c924077545f"},"cell_type":"markdown","source":"Display unclustered countries"},{"metadata":{"trusted":true,"_uuid":"ae1985c305023c4a5a44da17ad0bf7638f0755f0"},"cell_type":"code","source":"countryDF[countryDF['groups']==-1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c0bf95945f94ae15653fa9caee01b4e394a2d7cf"},"cell_type":"code","source":"countryDF['groups']=countryDF['groups']+1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"005c1267c75ebc87680ea1b6d92844a0e277a67d"},"cell_type":"code","source":"countryDF.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"911a787f7b5cea0daaded7221daa5c66141d3eac"},"cell_type":"markdown","source":"Save results to file."},{"metadata":{"trusted":true,"_uuid":"31cc80bedd2ca050b757f9ff54e4e01624a1afeb"},"cell_type":"code","source":"countryDF.to_csv('area_mapping.csv', columns=['alpha3','groups'], index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bc559d6f68ccd0e62c30c380ed0a7a4637b479bf"},"cell_type":"markdown","source":"### Dict usage\nHow the file with countries groups can be used as dictionary."},{"metadata":{"trusted":true,"_uuid":"af91b2b495c4fb3a90c8901d775aee0cbc8ace16"},"cell_type":"code","source":"our_mapping = pd.read_csv('area_mapping.csv')\nour_mapping.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"67d81d678430cdd19b5db89f807ba1a59f50b281"},"cell_type":"markdown","source":"Unknown countries will be assigned as unclustered"},{"metadata":{"trusted":true,"_uuid":"8bb261939f0321343b9ab9161d1910ce437c1c01"},"cell_type":"code","source":"our_dict = pd.Series(our_mapping.groups.values, index = our_mapping.alpha3).to_dict()\n\ndef get_area(iso_code):\n    try:\n        return our_dict[iso_code]\n    except KeyError:\n        return 0","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a6d2a5acffd10dc73aa9eeffa8326c76b3bf5198"},"cell_type":"code","source":"print(get_area('POL'))\nprint(get_area('not valid'))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"72864dd33bfdce1ae70d53d3076a008e4a0dae1b"},"cell_type":"markdown","source":"### Visualize clustering result"},{"metadata":{"trusted":true,"_uuid":"b9b3fd4b20ccbaf9a7109975291f284ff9406c8c"},"cell_type":"code","source":"world = gp.read_file(gp.datasets.get_path('naturalearth_lowres'))\nworld['groups']=0\nworld.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ccf12214ee71d4e320ee6e4036d635727e928c54"},"cell_type":"code","source":"for _id in range(len(world)):\n    number=countryDF.index[countryDF['alpha3']==world.loc[_id,'iso_a3']]\n    \n    if len(number)>0:        \n        tmp = countryDF.loc[number,'groups']\n        world.at[_id,'groups'] = tmp","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"21c25da89a5dfcedc700eb718211f0ef0f5c1835"},"cell_type":"code","source":"world.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e73048765682a0b3344ad3d455c996d766206cab"},"cell_type":"markdown","source":"Show countries groups"},{"metadata":{"trusted":true,"_uuid":"437017ec9af798310c636057b528bcbe6dcb0efc"},"cell_type":"code","source":"n_groups = max(our_mapping.groups)+1","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f597d86671b5e398bea0ba722402f498f3cf98e3"},"cell_type":"markdown","source":"The only purpose of shuffling groups is to spread groups around the world map so similar colors won't be close to each other, what may be confusing."},{"metadata":{"trusted":true,"_uuid":"6006040a8289f5ddf26680b88edb465401ba76eb"},"cell_type":"code","source":"shuffled = np.random.permutation(n_groups)\ncategories_dict = {_id: shuffled[_id]  for _id in range(n_groups)}\nworld['groups'] = world['groups'].replace(categories_dict)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"465f3f18b8ee5969b8a0e912dc985ed690b8b326"},"cell_type":"code","source":"world.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9601d785ade8a15725471d23319f64abb3fa76f9"},"cell_type":"code","source":"ax = world.plot(color='white', edgecolor='black', figsize=(20,20))\nplot = world.plot(ax=ax, column='groups', cmap='jet')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cde4fcea7cd414160a1a5529a65f1ff569267696"},"cell_type":"markdown","source":"Show countries without group"},{"metadata":{"trusted":true,"_uuid":"19511fd2416d9b1389459afb4b30c38086c78b09"},"cell_type":"code","source":"wun=world[world['groups']==categories_dict[0]]\nax = world.plot(color='white', edgecolor='black', figsize=(20,20))\nplot = wun.plot(ax=ax, column='groups', cmap='jet', vmin=0,vmax=1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b79809aa87be015435a02600c2458b076e751391"},"cell_type":"markdown","source":"### Manual groups tunning\nFrom the first image (map) one can clearly see that e.g. the grouping of Russia, Mongolia, China and Japan may not be a perfect idea if we expect to cluster countries similar in terms of culture. One simple way of creating manually another group of some already assigned countries is presented below."},{"metadata":{"trusted":true,"_uuid":"f53a19b9ed61b180554a222562ceea0eae098a10"},"cell_type":"code","source":"world.loc[world.iso_a3=='RUS']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a80c8a62c1013c5f383af4a9156fc71f6e96c6b3"},"cell_type":"code","source":"world.loc[world.groups==world.loc[world.iso_a3=='RUS'].groups.values[0]]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b6b07b201c3bb3bd2f8a4935418da600da231e6a"},"cell_type":"code","source":"def reGroup(iso_a3_list):\n    new_group_id = max(world.groups)+1\n    \n    for item in iso_a3_list:\n        world.loc[world.iso_a3==item,'groups']=new_group_id    ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"aa201c5b06aeb7bca96ec9d122ffd359a1cae952"},"cell_type":"markdown","source":"If you think that e.g. Russia and Mongolia should be in another group, run the following cell:"},{"metadata":{"trusted":true,"_uuid":"4a2e301a919bfc24afde92c6b794be9c6f898788"},"cell_type":"code","source":"reGroup(['RUS', 'MNG'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2afc20582265b7004fd12e5f1c201bb663b9fcd1"},"cell_type":"code","source":"ax = world.plot(color='white', edgecolor='black', figsize=(20,20))\nplot = world.plot(ax=ax, column='groups', cmap='jet')","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}