{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction","metadata":{"_uuid":"23568bf6-b483-4001-ad1d-86e3be74eed4","_cell_guid":"3a936ab5-7f05-401c-a532-f32d2a34477f","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"Hi!\nThis is new playground competition and in this month we're having an insurance dataset. \n\nBefore diving deeper into code, let's describe a little topic in which we will work.\n\nFor start, a little meme which describes all this thing :)\n\n![](https://leadsurance.com/wp-content/uploads/2020/08/insurance-meme-6.jpg)\r\n","metadata":{}},{"cell_type":"markdown","source":"# Problem definition","metadata":{}},{"cell_type":"markdown","source":"## What is insurance?","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#5642C5;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n\n<p style=\"padding: 10px;\n              color:white;\">Insurance is a means of protection from financial loss in which, in exchange for a fee, a party agrees to compensate another party in the event of a certain loss, damage, or injury. It is a form of risk management, primarily used to protect against the risk of a contingent or uncertain loss. <br><br>\n    An entity which provides insurance is known as an insurer, insurance company, insurance carrier, or underwriter. A person or entity who buys insurance is known as a policyholder, while a person or entity covered under the policy is called an insured.\n\r\n</p>\n</div>","metadata":{}},{"cell_type":"markdown","source":"## How insurance Works\r<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#5642C5;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n\n<p style=\"padding: 10px;\n              color:white;\">\nWhen you buy a policy you make regular payments, known as premiums, to the insurer. If you make a claim your insurer will pay out for the loss that is covered under the policy. <br><br>\nIf you don’t make a claim, you won’t get your money back; instead it is pooled with the premiums of other policyholders who have taken out insurance with the same insurance company. If you make a claim the money comes from the pool of policyholders’ premiums.\n</p>\n</div>\n","metadata":{}},{"cell_type":"markdown","source":"## Insurance Policy Components\r\n","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;\r\n           display:fill;\r\n           border-radius:5px;\r\n           background-color:#5642C5;\r\n           font-size:110%;\r\n           font-family:Verdana;\r\n           letter-spacing:0.5px\">\r\n\r\n<ul>\r\n  <li>\r\n    <b>Insurance Premiums</b><br>\r\n    An insurance policy's premium is the amount you must pay to obtain a specified quantity of insurance coverage. It's usually described as a recurring fee that you pay as a lumpsum, or on a monthly, quarterly, half-yearly, or annual basis during the premium payment period.<br><br>\r\n    An insurance firm determines the premium of an insurance plan depending on a number of criteria. The goal is to determine if an insured person is eligible for the type of insurance plan he or she wants to purchase.<br><br>\r\n    For example, suppose you are fit and have no medical history of receiving treatment for serious physical disorders. In that case, you will likely be paying less for medical insurance or life insurance than those who have many ailments.<br><br>\r\n    You should also be aware that for comparable products, various insurance providers may charge varying prices. As a result, finding the correct one at a reasonable price takes some time and effort.\r\n  </li>\r\n\r\n  <li>\r\n    <b>Policy Restrictions</b><br>\r\n    It is defined as the maximum amount for which an insurance company is responsible for losses covered by the policy. It is calculated depending on the insurance term, loss or damage, and other comparable criteria.<br><br>\r\n    Typically, the greater the policy limit, the greater the premium. The sum assured is the maximum amount that an insurer will pay to the nominee under a life insurance policy.\r\n  </li>\r\n\r\n  <li>\r\n    <b>Deductible</b><br>\r\n    The amount or percentage that the insured decides to pay out of pocket before the insurer steps in to settle a claim is referred to as the deductible. The insurance company is liable to pay the claim amount only when it exceeds the deductible.<br><br>\r\n    Deductibles are determined by the provisions of a certain type of policy and are applicable per policy or per claim. In general, insurance policies with large deductibles are less expensive since fewer claims are filed due to the greater out-of-pocket price.\r\n  </li>\r\n</ul>\r\n\r\n</div>\r\n","metadata":{}},{"cell_type":"markdown","source":"# Import libraries","metadata":{"_uuid":"23e0c399-9c79-4374-8847-08df1b1784aa","_cell_guid":"096ad279-cbc5-49fd-b80c-9f34e221487e","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import mean_squared_log_error\nimport pandas as pd\nimport missingno as mnso\nimport numpy as np\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.preprocessing import StandardScaler, OneHotEncoder\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.metrics import mean_squared_error\nfrom sklearn.ensemble import RandomForestRegressor\nfrom scipy.stats import boxcox\nfrom scipy.special import boxcox1p\n","metadata":{"_uuid":"90871205-85e8-45d8-97a6-490aee449832","_cell_guid":"f9eafd57-6931-4703-8bb0-eef4e35e86e7","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-12-18T20:16:54.758488Z","iopub.execute_input":"2024-12-18T20:16:54.75885Z","iopub.status.idle":"2024-12-18T20:16:55.856378Z","shell.execute_reply.started":"2024-12-18T20:16:54.758796Z","shell.execute_reply":"2024-12-18T20:16:55.855447Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Load data","metadata":{"_uuid":"d005054f-58e9-4091-b065-00c10dbfee18","_cell_guid":"229240ab-e333-4f0a-a562-7749e4add3aa","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"After exploring discussions it seems reasonable to log transform our target variable to ensure that we optimizing directly the evaluation metric for this competition","metadata":{}},{"cell_type":"code","source":"X_train = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\nX_test = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')\nsample_submission = pd.read_csv('/kaggle/input/playground-series-s4e12/sample_submission.csv')\n\nX_train = X_train.drop(['id'], axis=1)\nid_test = X_test.id\nX_test = X_test.drop('id', axis=1)\n# y = X_train['Premium Amount']\n# y_log = np.log1p(y)\n\npd.set_option('display.max_columns', None)","metadata":{"_uuid":"cbbb9373-1d55-4852-8548-09338a7de6d5","_cell_guid":"b811f64a-b2c4-4899-bdcf-9f948b67f048","trusted":true,"execution":{"iopub.status.busy":"2024-12-18T20:16:55.859006Z","iopub.execute_input":"2024-12-18T20:16:55.859578Z","iopub.status.idle":"2024-12-18T20:17:02.12822Z","shell.execute_reply.started":"2024-12-18T20:16:55.85953Z","shell.execute_reply":"2024-12-18T20:17:02.127291Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_train.columns","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-18T20:17:02.129447Z","iopub.execute_input":"2024-12-18T20:17:02.129801Z","iopub.status.idle":"2024-12-18T20:17:02.13745Z","shell.execute_reply.started":"2024-12-18T20:17:02.12976Z","shell.execute_reply":"2024-12-18T20:17:02.136607Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# EDA","metadata":{"_uuid":"55ee458e-7685-4061-a945-0fb6664d5fab","_cell_guid":"ff6fdb2b-3dd8-4946-b5c3-d8388ecd9432","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"X_train","metadata":{"_uuid":"145de48c-b9a5-4e07-bf4c-bb5884e16d5e","_cell_guid":"4f6a074b-f0e5-4eab-99c1-3392b5f7f2f7","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:02.138634Z","iopub.execute_input":"2024-12-18T20:17:02.139224Z","iopub.status.idle":"2024-12-18T20:17:02.1662Z","shell.execute_reply.started":"2024-12-18T20:17:02.139167Z","shell.execute_reply":"2024-12-18T20:17:02.165456Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_train.info()","metadata":{"_uuid":"d4cc7adb-0d6f-49dd-b3ad-f5b2c95d5202","_cell_guid":"5b9d679d-48e4-45f6-bbc3-5ccb5b1eca9a","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:02.167395Z","iopub.execute_input":"2024-12-18T20:17:02.168007Z","iopub.status.idle":"2024-12-18T20:17:02.718331Z","shell.execute_reply.started":"2024-12-18T20:17:02.167966Z","shell.execute_reply":"2024-12-18T20:17:02.717463Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_train['Number of Dependents'] = X_train['Number of Dependents'].astype('Int64')\nX_train['Vehicle Age'] = X_train['Vehicle Age'].astype('Int64')\nX_train['Insurance Duration'] = X_train['Insurance Duration'].astype('Int64')\nX_train['Previous Claims'] = X_train['Previous Claims'].astype('Int64')","metadata":{"_uuid":"9a16ab6e-08a3-4034-8baf-3d0ca5e511c8","_cell_guid":"f720fbc2-241b-4aa1-b974-687e33c3f0d8","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:02.719488Z","iopub.execute_input":"2024-12-18T20:17:02.71976Z","iopub.status.idle":"2024-12-18T20:17:03.380897Z","shell.execute_reply.started":"2024-12-18T20:17:02.719734Z","shell.execute_reply":"2024-12-18T20:17:03.379832Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This features I decided to treat as discrete (maybe categorical), because they have fixed range of values","metadata":{}},{"cell_type":"markdown","source":"## Defining missing values","metadata":{"_uuid":"f5574e81-fad9-4a58-bb8d-fbaba0cd2b6f","_cell_guid":"bb69f33a-f5ff-46ab-8baf-ac2c5d5a66ab","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"mnso.bar(X_train)","metadata":{"_uuid":"ab35b427-5fb2-4c8d-8949-8ed2f3e85885","_cell_guid":"625d8c13-5048-486e-95c9-991df4c778fc","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:03.384765Z","iopub.execute_input":"2024-12-18T20:17:03.385613Z","iopub.status.idle":"2024-12-18T20:17:05.694922Z","shell.execute_reply.started":"2024-12-18T20:17:03.385568Z","shell.execute_reply":"2024-12-18T20:17:05.694045Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It seems that we're having a smaller number of missing values than in previous competition.","metadata":{}},{"cell_type":"markdown","source":"## Defining feature's types","metadata":{"_uuid":"2a8f2644-611b-442d-a7ca-81511d599a47","_cell_guid":"5620dc3d-1d83-4602-9a35-0721e4a230e4","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"numerical_features_total = X_train.select_dtypes(exclude='object').columns\nnumerical_features_continuos = X_train.select_dtypes(exclude=['object', 'int']).columns\nnumerical_features_discrete = X_train.select_dtypes(exclude=['object', 'float']).columns\ncategorical_features = X_train.select_dtypes(include='object').columns","metadata":{"_uuid":"a3ed8ec3-5a89-4c2a-a95f-5a34e39dc2ec","_cell_guid":"7dcc0717-8747-4568-b7fe-f9017a3a7b9f","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:05.696063Z","iopub.execute_input":"2024-12-18T20:17:05.696325Z","iopub.status.idle":"2024-12-18T20:17:05.985119Z","shell.execute_reply.started":"2024-12-18T20:17:05.696299Z","shell.execute_reply":"2024-12-18T20:17:05.984285Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Heatmap for numerical features","metadata":{"_uuid":"0c30f7eb-f7dd-4765-8b3f-5259196f4226","_cell_guid":"6367a821-3cd8-407c-a58d-51cb62aa22e5","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"plt.figure(figsize=(10, 10))\nsns.heatmap(X_train[numerical_features_total].corr(), annot=True)","metadata":{"_uuid":"a0ac4a0c-298c-4769-bdae-9111875b7bec","_cell_guid":"2ca32ae5-0b17-4754-b4b4-b6e5c77734fc","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:05.986182Z","iopub.execute_input":"2024-12-18T20:17:05.98648Z","iopub.status.idle":"2024-12-18T20:17:06.901255Z","shell.execute_reply.started":"2024-12-18T20:17:05.986452Z","shell.execute_reply":"2024-12-18T20:17:06.900385Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It seems that our features doesnt have a direct correlation with target variable (Premium Amount). In fact there's no strong correlation between every pair of features. Maybe situation will change after some feature engineering later.","metadata":{}},{"cell_type":"markdown","source":"## Numerical features summary","metadata":{"_uuid":"07c2007c-ab07-4e73-9351-ac3e23f42a50","_cell_guid":"2772a3ce-1bca-4d15-be2b-ed5c83708804","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"X_train[numerical_features_total].describe()","metadata":{"_uuid":"6685c723-cfff-4b6a-aace-eec17b8b834c","_cell_guid":"4179eaba-8d75-4baf-a0e1-d6bff698bf57","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:06.902503Z","iopub.execute_input":"2024-12-18T20:17:06.902851Z","iopub.status.idle":"2024-12-18T20:17:07.519228Z","shell.execute_reply.started":"2024-12-18T20:17:06.902813Z","shell.execute_reply":"2024-12-18T20:17:07.518273Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Vizualization of continuos features","metadata":{"_uuid":"64f3d234-c851-449a-a2fc-3e46b7c7abdd","_cell_guid":"bc2d3b2e-1646-4c3a-9411-b80f6ca02945","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"numerical_features_continuos","metadata":{"_uuid":"bd51bde0-fc63-491d-9861-f2e61b456a6c","_cell_guid":"03c15480-92d4-463d-bf72-8a5c216bbba5","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:07.520617Z","iopub.execute_input":"2024-12-18T20:17:07.521023Z","iopub.status.idle":"2024-12-18T20:17:07.528004Z","shell.execute_reply.started":"2024-12-18T20:17:07.520978Z","shell.execute_reply":"2024-12-18T20:17:07.527052Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, axes = plt.subplots(2, 3, figsize=(15,15))\nfig.suptitle('Distributions of continuous features')\n\n\nsns.kdeplot(ax=axes[0,0], data=X_train, x='Credit Score')\naxes[0,0] = 'Credit Score Distribution'\nsns.kdeplot(ax=axes[0,1], data=X_train, x='Age')\naxes[0,1] = 'Age Distribution'\nsns.kdeplot(ax=axes[0,2], data=X_train, x='Annual Income')\naxes[0,2] = 'Annual Income Distribution'\nsns.kdeplot(ax=axes[1,0], data=X_train, x='Health Score')\naxes[1,0] = 'Health Score Distribution'\nsns.kdeplot(ax=axes[1,1], data=X_train, x='Premium Amount')\naxes[1,1] = 'Premium Amount Distribution'","metadata":{"_uuid":"880db4c4-4a9e-4a9c-a7fc-1a9e01b4fbaa","_cell_guid":"b925e22d-12b8-4381-8c69-8fe406292859","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:07.529176Z","iopub.execute_input":"2024-12-18T20:17:07.529475Z","iopub.status.idle":"2024-12-18T20:17:28.784922Z","shell.execute_reply.started":"2024-12-18T20:17:07.529431Z","shell.execute_reply":"2024-12-18T20:17:28.784041Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Vizualization of discrete features","metadata":{"_uuid":"60754b8e-0825-4ded-bf27-13e360be1b58","_cell_guid":"bf136762-6bad-4967-8c5f-f51c2218bfc9","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"numerical_features_discrete","metadata":{"_uuid":"4e380837-46dc-4bb7-8cd2-a68e9cc7721f","_cell_guid":"d556efc7-5ea2-4621-964b-668657526961","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:28.786081Z","iopub.execute_input":"2024-12-18T20:17:28.786462Z","iopub.status.idle":"2024-12-18T20:17:28.792407Z","shell.execute_reply.started":"2024-12-18T20:17:28.786408Z","shell.execute_reply":"2024-12-18T20:17:28.791687Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, axes = plt.subplots(2, 2, figsize=(15,15))\nfig.suptitle('Distributions of discrete features')\n\n\nsns.countplot(ax=axes[0,0], data=X_train, x='Number of Dependents')\naxes[0,0] = 'Number of Dependents Distribution'\nsns.countplot(ax=axes[0,1], data=X_train, x='Previous Claims')\naxes[0,1] = 'Previous Claims Distribution'\nsns.countplot(ax=axes[1,0], data=X_train, x='Vehicle Age')\naxes[1,0] = 'Vehicle Age Distribution'\nsns.countplot(ax=axes[1,1], data=X_train, x='Insurance Duration')\naxes[1,1] = 'Insurance Duration Distribution'","metadata":{"_uuid":"c2567f4b-ec2d-4e67-ba57-6aa5b0323fed","_cell_guid":"d94e649a-c7f9-4c65-bda4-96923a01ff36","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:28.793826Z","iopub.execute_input":"2024-12-18T20:17:28.794515Z","iopub.status.idle":"2024-12-18T20:17:29.952247Z","shell.execute_reply.started":"2024-12-18T20:17:28.794485Z","shell.execute_reply":"2024-12-18T20:17:29.951306Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Vizualization of categorical features","metadata":{"_uuid":"1111df4e-66e3-424b-b90d-27ca5c5a38d8","_cell_guid":"143e8395-23fa-4be1-962a-59562cd05c00","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"categorical_features","metadata":{"_uuid":"73fbfe47-5ba9-4d86-aa00-883df3ffb357","_cell_guid":"218b6ea8-5a2a-40f2-a5bb-166e4e794244","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:29.953554Z","iopub.execute_input":"2024-12-18T20:17:29.954233Z","iopub.status.idle":"2024-12-18T20:17:29.961002Z","shell.execute_reply.started":"2024-12-18T20:17:29.954184Z","shell.execute_reply":"2024-12-18T20:17:29.959978Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"new_array = np.delete(categorical_features, np.where(categorical_features == 'Policy Start Date'))","metadata":{"_uuid":"2170d23a-9e45-408d-b808-d4baf7b8e666","_cell_guid":"86bb2be7-be6d-4177-89e4-bf4aad0511ac","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:29.962148Z","iopub.execute_input":"2024-12-18T20:17:29.962462Z","iopub.status.idle":"2024-12-18T20:17:29.969521Z","shell.execute_reply.started":"2024-12-18T20:17:29.962406Z","shell.execute_reply":"2024-12-18T20:17:29.968676Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"new_array","metadata":{"_uuid":"2d20f541-684f-47d1-9d8e-8d38d8a1df2d","_cell_guid":"cc130e89-e004-42ae-ab94-55c4ae9c56f2","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:29.970599Z","iopub.execute_input":"2024-12-18T20:17:29.970859Z","iopub.status.idle":"2024-12-18T20:17:29.981231Z","shell.execute_reply.started":"2024-12-18T20:17:29.970833Z","shell.execute_reply":"2024-12-18T20:17:29.980447Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_categorical_features(data, features):\n    fig, axes = plt.subplots(5, 2, figsize=(15,20))\n    fig.suptitle('Distributions of categorical features')\n    feat = 0\n    for a in axes.flat:\n        sns.countplot(ax=a, data=data, x=features[feat])\n        a = features[feat]\n        feat +=1 \n\n    plt.show()\nplot_categorical_features(X_train, new_array)","metadata":{"_uuid":"2a5d63a0-8638-4520-bb2b-eef7976a29ae","_cell_guid":"34d682e7-4668-4930-bf01-366832cc1ef6","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:29.98261Z","iopub.execute_input":"2024-12-18T20:17:29.982853Z","iopub.status.idle":"2024-12-18T20:17:36.50069Z","shell.execute_reply.started":"2024-12-18T20:17:29.982829Z","shell.execute_reply":"2024-12-18T20:17:36.499773Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Plot value counts for each feature","metadata":{"_uuid":"719c0c8c-2da7-4a21-89de-0c3ac30746e7","_cell_guid":"762fa5f9-d160-4ac1-bd83-749c853a7f2b","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"for cat in categorical_features:\n    print(X_train[cat].value_counts())\n    print()","metadata":{"_uuid":"1badf574-06e1-425a-a952-28ae34fdc06e","_cell_guid":"0c9d124b-9be9-4a88-98bb-132172fd3b29","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2024-12-18T20:17:36.501905Z","iopub.execute_input":"2024-12-18T20:17:36.502272Z","iopub.status.idle":"2024-12-18T20:17:37.588087Z","shell.execute_reply.started":"2024-12-18T20:17:36.50223Z","shell.execute_reply":"2024-12-18T20:17:37.587127Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As we can see categorical features approximately have the almost the same number of categories, except Policy Start Date. I've decided not to plot it, because it would take a lo-o-ot of time.\n\nFrom value counts we can assume types of encoding for our categorical features.\nBinary encoding:\n* Gender\n* Smoking Status\n\nOrdinal encoding:\n* Education Level\n* Policy Type\n* Exercise frequency\n* Customer Feedback\n\nAll other features we can encode by one-hot encoding","metadata":{"_uuid":"2a8e9c94-9e8d-4fa8-8b45-e7e07b98432a","_cell_guid":"a3461cf8-02d4-46d2-a0b3-2b2bd486b818","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"## Vizualization of features vs target","metadata":{}},{"cell_type":"markdown","source":"I've did this to see have features depend on target, so I can see which type of relationship is here","metadata":{}},{"cell_type":"code","source":"numerical_features_total","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-18T20:17:37.58947Z","iopub.execute_input":"2024-12-18T20:17:37.589888Z","iopub.status.idle":"2024-12-18T20:17:37.596956Z","shell.execute_reply.started":"2024-12-18T20:17:37.589832Z","shell.execute_reply":"2024-12-18T20:17:37.59586Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for n in numerical_features_total:\n    sns.scatterplot(X_train, x=n, y='Premium Amount')\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-18T20:17:37.598213Z","iopub.execute_input":"2024-12-18T20:17:37.598537Z","iopub.status.idle":"2024-12-18T20:17:56.839288Z","shell.execute_reply.started":"2024-12-18T20:17:37.598509Z","shell.execute_reply":"2024-12-18T20:17:56.838436Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Sort by Premium Amount","metadata":{}},{"cell_type":"markdown","source":"Here I decided to sort dataset by premium amount to get insight which type of users pay the highest premium","metadata":{}},{"cell_type":"code","source":"sorted_data = X_train.sort_values(by='Premium Amount', ascending=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-18T20:17:56.844305Z","iopub.execute_input":"2024-12-18T20:17:56.844723Z","iopub.status.idle":"2024-12-18T20:17:57.619948Z","shell.execute_reply.started":"2024-12-18T20:17:56.844695Z","shell.execute_reply":"2024-12-18T20:17:57.619205Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sorted_data.head(10)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-18T20:17:57.621182Z","iopub.execute_input":"2024-12-18T20:17:57.62156Z","iopub.status.idle":"2024-12-18T20:17:57.64346Z","shell.execute_reply.started":"2024-12-18T20:17:57.621521Z","shell.execute_reply":"2024-12-18T20:17:57.642478Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import seaborn as sns\n\n# Вибір змінних\nsubset = sorted_data[['Age', 'Annual Income', 'Premium Amount']]\n\n# Візуалізація залежності між віком та здоров'ям для премії\nsns.scatterplot(data=subset, x='Age', y='Annual Income', size='Premium Amount', hue='Premium Amount', palette='cool', alpha=0.7)\nplt.title('Залежність Premium Amount від Age і Health Score')\nplt.xlabel('Age')\nplt.ylabel('Health Score')\nplt.legend()\nplt.grid(True)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-18T20:17:57.644516Z","iopub.execute_input":"2024-12-18T20:17:57.644738Z","iopub.status.idle":"2024-12-18T20:19:00.776441Z","shell.execute_reply.started":"2024-12-18T20:17:57.644715Z","shell.execute_reply":"2024-12-18T20:19:00.775567Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sns.histplot(np.log(sorted_data['Premium Amount']), bins=30)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-18T20:19:00.777735Z","iopub.execute_input":"2024-12-18T20:19:00.778068Z","iopub.status.idle":"2024-12-18T20:19:01.704235Z","shell.execute_reply.started":"2024-12-18T20:19:00.778033Z","shell.execute_reply":"2024-12-18T20:19:01.703357Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"box_features = sorted_data.select_dtypes(include=['object']).columns.tolist()\nbox_features.remove('Policy Start Date')\nbox_features","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-18T20:19:01.705564Z","iopub.execute_input":"2024-12-18T20:19:01.706278Z","iopub.status.idle":"2024-12-18T20:19:01.898354Z","shell.execute_reply.started":"2024-12-18T20:19:01.706237Z","shell.execute_reply":"2024-12-18T20:19:01.897477Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for b in box_features:\n    sns.boxplot(data=sorted_data, x=b, y='Premium Amount')\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-18T20:19:01.899786Z","iopub.execute_input":"2024-12-18T20:19:01.900094Z","iopub.status.idle":"2024-12-18T20:19:09.139923Z","shell.execute_reply.started":"2024-12-18T20:19:01.900065Z","shell.execute_reply":"2024-12-18T20:19:09.13896Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_train.isna().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-18T20:19:09.143744Z","iopub.execute_input":"2024-12-18T20:19:09.144017Z","iopub.status.idle":"2024-12-18T20:19:09.695111Z","shell.execute_reply.started":"2024-12-18T20:19:09.143988Z","shell.execute_reply":"2024-12-18T20:19:09.694208Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}