{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n<h1 align=\"center\"><font color='#386068'> Insurance Premium Prediction </font></h1>\n    \n","metadata":{}},{"cell_type":"markdown","source":"**Aim**: is to develop a model that accurately forecasts the premium amount a policyholder is likely to pay based on various factors. By analyzing historical data, this model helps insurers optimize pricing strategies, reduce risks, and improve customer satisfaction by providing fair and tailored premium estimates.\n\n**Evaluation metrics**: RMSLE (Root Mean Squared Logarithmic Error) is a metric used to evaluate the performance of regression models, particularly when dealing with data that can vary over a large range, including exponential or skewed distributions. It focuses on reducing the impact of large errors that can occur in high-value predictions, which is often useful in cases where the target variable is positive and spans several orders of magnitude.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"border-radius:5px; border:#6E2A04 solid; padding: 15px; background-color: #BED4D2; font-size:100%; text-align:left\">\n\n<h2 align=\"center\"><font color='#6E2A04'>Importing Libraries </font></h2>","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom scipy import stats\nimport re\nfrom scipy.stats import norm, skew\n\n\nfrom sklearn.model_selection import StratifiedKFold, cross_val_score, KFold\nfrom sklearn.preprocessing import LabelEncoder, OrdinalEncoder, OneHotEncoder, StandardScaler, PowerTransformer\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.metrics import mean_squared_error, mean_squared_log_error\nfrom xgboost import XGBRegressor\nfrom lightgbm import LGBMRegressor, early_stopping\nfrom sklearn.model_selection import train_test_split\nfrom catboost import  CatBoostRegressor\nimport optuna\nfrom optuna import Trial\nfrom sklearn.metrics import accuracy_score, classification_report, roc_auc_score\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.pipeline import Pipeline\n\nimport gc\nimport warnings\n\n# Suppress all warnings\nwarnings.filterwarnings('ignore')\n# select color palette\npalette_color = sns.color_palette('dark')\n\n","metadata":{"execution":{"iopub.status.busy":"2024-12-16T08:37:18.65887Z","iopub.execute_input":"2024-12-16T08:37:18.659118Z","iopub.status.idle":"2024-12-16T08:37:23.186989Z","shell.execute_reply.started":"2024-12-16T08:37:18.659086Z","shell.execute_reply":"2024-12-16T08:37:23.18625Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"border-radius:5px; border:#6E2A04 solid; padding: 15px; background-color: #BED4D2; font-size:100%; text-align:left\">\n\n<h2 align=\"center\"><font color='#6E2A04'>Loading Dataset</font></h2>","metadata":{}},{"cell_type":"code","source":"# load data\ndf_train=pd.read_csv(\"/kaggle/input/playground-series-s4e12/train.csv\")\ndf_test=pd.read_csv(\"/kaggle/input/playground-series-s4e12/test.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-12-16T08:37:23.189183Z","iopub.execute_input":"2024-12-16T08:37:23.190194Z","iopub.status.idle":"2024-12-16T08:37:31.0102Z","shell.execute_reply.started":"2024-12-16T08:37:23.190148Z","shell.execute_reply":"2024-12-16T08:37:31.009234Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# quick look at the training data\ndf_train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:37:31.011205Z","iopub.execute_input":"2024-12-16T08:37:31.011485Z","iopub.status.idle":"2024-12-16T08:37:31.045024Z","shell.execute_reply.started":"2024-12-16T08:37:31.011459Z","shell.execute_reply":"2024-12-16T08:37:31.044283Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# quick look at the testing data\ndf_test.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:37:31.045884Z","iopub.execute_input":"2024-12-16T08:37:31.04613Z","iopub.status.idle":"2024-12-16T08:37:31.062814Z","shell.execute_reply.started":"2024-12-16T08:37:31.046105Z","shell.execute_reply":"2024-12-16T08:37:31.061911Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n<div style=\"border-radius:5px; border:#6E2A04 solid; padding: 15px; background-color: #BED4D2; font-size:100%; text-align:left\">\n\n<h2 align=\"center\"><font color='#6E2A04'>Summarize the data</font></h2>\n","metadata":{}},{"cell_type":"code","source":"# Number of rows [examples], Number of columns [Features] \nprint(\"Total rows in train data: {0}, Total columns in train data: {1}\".\n      format(df_train.shape[0], df_train.shape[1]))\n      \nprint(\"Total rows in test data: {0}, Total columns in test data: {1}\".\n      format(df_test.shape[0], df_test.shape[1]))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:37:31.06378Z","iopub.execute_input":"2024-12-16T08:37:31.064045Z","iopub.status.idle":"2024-12-16T08:37:31.073647Z","shell.execute_reply.started":"2024-12-16T08:37:31.064007Z","shell.execute_reply":"2024-12-16T08:37:31.072905Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Information about our training set \ndf_train.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:37:31.074593Z","iopub.execute_input":"2024-12-16T08:37:31.07517Z","iopub.status.idle":"2024-12-16T08:37:31.627047Z","shell.execute_reply.started":"2024-12-16T08:37:31.075129Z","shell.execute_reply":"2024-12-16T08:37:31.626179Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train.dtypes.value_counts()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:37:31.629321Z","iopub.execute_input":"2024-12-16T08:37:31.629604Z","iopub.status.idle":"2024-12-16T08:37:31.636365Z","shell.execute_reply.started":"2024-12-16T08:37:31.629576Z","shell.execute_reply":"2024-12-16T08:37:31.635509Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n<h3 align=\"left\"><font color='#DAA520'>Datapoint Description:</font></h3>\n\r\n\r\n| **Feature**              | **Description**    |\r\n|--------------------------|----------------------------------------------------------------------------------------------------------------------------------------------|\r\n| **id**                    | A unique identifier for each customer.                                                                                                                 |\r\n| **Age**                   | The age of the customer.                                                                                                                                 |\r\n| **Gender**                | The gender of the customer (e.g., Male, Female).                                                                                                      |\r\n| **Annual Income**         | The yearly income of the customer in monetary units.                                                                                                  |\r\n| **Marital Status**        | The marital status of the customer (e.g., Married, Divorced).                                                                                         |\r\n| **Number of Dependents**  | The number of dependents the customer is financially responsible for.                                                                                 |\r\n| **Education Level**       | The highest level of education attained by the customer.                                                                                             |\r\n| **Occupation**            | The customer's job or profession (e.g., Self-Employed, Unemployed, etc.).                                                                              |\r\n| **Health Score**          | A score representing the customer’s general health.                                                                                                  |\r\n| **Location**              | The geographic location of the customer (e.g., Urban, Suburban, Rural).                                                                                |\r\n| **Policy Type**           | The type of insurance policy held by the customer (e.g., Basic, Comprehensive, Premium).                                                              |\r\n| **Previous Claims**       | The number of claims the customer has filed in the past.                                                                                                |\r\n| **Vehicle Age**           | The age of the customer’s vehicle (in years).                                                                                                         |\r\n| **Credit Score**          | The customer's credit score, which reflects their financial reliability.                                                                                |\r\n| **Insurance Duration**    | The length of time (in years or months) the customer has been with the insurance company.                                                              |\r\n| **Policy Start Date**     | The date when the insurance policy was initiated.                                                                                                     |\r\n| **Customer Feedback**     | A qualitative assessment of customer satisfaction (e.g., Poor, Average, Good).                                                                          |\r\n| **Smoking Status**        | Whether the customer smokes or not.                                                                                                                     |\r\n| **Exercise Frequency**    | How often the customer exercises (e.g., Daily, Weekly, Rarely, Monthly).                                                                                |\r\n| **Property Type**         | The type of property the customer owns (e.g., House, Apartment, Condo).                                                           if you need anything else!\n","metadata":{}},{"cell_type":"markdown","source":"\n<div style=\"border-radius:5px; border:#6E2A04 solid; padding: 15px; background-color: #BED4D2; font-size:100%; text-align:left\">\n\n<h2 align=\"center\"><font color='#6E2A04'>Exploratory Data Analysis (EDA)</font></h2>\n","metadata":{}},{"cell_type":"markdown","source":"<h3 align=\"left\"><font color='#DAA520'>Target Destribution:</font></h3>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(5, 5))\n# Plotting\nsns.histplot(data=df_train, x='Premium Amount',bins=30, kde=True, alpha=0.5)\n\n# Adding a title\nplt.title('Target Distribution', fontsize=18, weight='bold')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:37:31.63737Z","iopub.execute_input":"2024-12-16T08:37:31.63761Z","iopub.status.idle":"2024-12-16T08:37:36.108754Z","shell.execute_reply.started":"2024-12-16T08:37:31.637585Z","shell.execute_reply":"2024-12-16T08:37:36.10792Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Skewness: %f\" % df_train['Premium Amount'].skew())\nprint(\"Kurtosis: %f\" % df_train['Premium Amount'].kurt())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:37:36.110058Z","iopub.execute_input":"2024-12-16T08:37:36.110775Z","iopub.status.idle":"2024-12-16T08:37:36.140221Z","shell.execute_reply.started":"2024-12-16T08:37:36.110732Z","shell.execute_reply":"2024-12-16T08:37:36.139298Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pt = PowerTransformer(method= 'box-cox')\ndf_train['Premium Amount']=pt.fit_transform(df_train[['Premium Amount']])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:37:36.141396Z","iopub.execute_input":"2024-12-16T08:37:36.141774Z","iopub.status.idle":"2024-12-16T08:37:41.246467Z","shell.execute_reply.started":"2024-12-16T08:37:36.141731Z","shell.execute_reply":"2024-12-16T08:37:41.245687Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(5, 5))\n# Plotting\nsns.histplot(data=df_train, x='Premium Amount',bins=30, kde=True, alpha=0.5)\n\n# Adding a title\nplt.title('Target Distribution', fontsize=18, weight='bold')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:37:41.247715Z","iopub.execute_input":"2024-12-16T08:37:41.248427Z","iopub.status.idle":"2024-12-16T08:37:45.628006Z","shell.execute_reply.started":"2024-12-16T08:37:41.248384Z","shell.execute_reply":"2024-12-16T08:37:45.62712Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Skewness: %f\" % df_train['Premium Amount'].skew())\nprint(\"Kurtosis: %f\" % df_train['Premium Amount'].kurt())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:37:45.629054Z","iopub.execute_input":"2024-12-16T08:37:45.629347Z","iopub.status.idle":"2024-12-16T08:37:45.657794Z","shell.execute_reply.started":"2024-12-16T08:37:45.629319Z","shell.execute_reply":"2024-12-16T08:37:45.656914Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h3 align=\"left\"><font color='#DAA520'>Numarical Variables:</font></h3>","metadata":{}},{"cell_type":"code","source":"# describe the Numarical columns in training data\ndf_train.describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:37:45.658949Z","iopub.execute_input":"2024-12-16T08:37:45.659579Z","iopub.status.idle":"2024-12-16T08:37:46.292351Z","shell.execute_reply.started":"2024-12-16T08:37:45.659545Z","shell.execute_reply":"2024-12-16T08:37:46.291454Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# describe the Numarical columns testing data   \ndf_test.describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:37:46.293699Z","iopub.execute_input":"2024-12-16T08:37:46.294112Z","iopub.status.idle":"2024-12-16T08:37:46.610178Z","shell.execute_reply.started":"2024-12-16T08:37:46.294067Z","shell.execute_reply":"2024-12-16T08:37:46.609269Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"numeric_vars = df_train.select_dtypes(\"number\")\n\ndef diagnostic(df, var):\n    fig = plt.figure(figsize = (10, 4))\n    plt.subplot(1,3,1)\n    df[var].hist(bins = 40)\n    plt.title(\"Distribution of {}\".format(var))\n    \n    plt.subplot(1,3,2)\n    stats.probplot(df[var], dist = \"norm\", plot = plt)\n    plt.ylabel(\"Quantiles\")\n    \n    plt.subplot(1,3,3)\n    sns.boxplot(y = df[var])\n    plt.title(\"Boxplot\")\n    plt.show()\n    \nfor var in numeric_vars:\n    diagnostic(df_train, var)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:37:46.611324Z","iopub.execute_input":"2024-12-16T08:37:46.611611Z","iopub.status.idle":"2024-12-16T08:38:22.494707Z","shell.execute_reply.started":"2024-12-16T08:37:46.611583Z","shell.execute_reply":"2024-12-16T08:38:22.493699Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h3 align=\"left\"><font color='#DAA520'>Categorical Variables:</font></h3>","metadata":{}},{"cell_type":"code","source":"# describe the Categorical columns\n\ndf_train.describe(include=['O'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:38:22.495732Z","iopub.execute_input":"2024-12-16T08:38:22.495988Z","iopub.status.idle":"2024-12-16T08:38:24.191027Z","shell.execute_reply.started":"2024-12-16T08:38:22.495963Z","shell.execute_reply":"2024-12-16T08:38:24.190129Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_test.describe(include=['O'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:38:24.192032Z","iopub.execute_input":"2024-12-16T08:38:24.192321Z","iopub.status.idle":"2024-12-16T08:38:25.300561Z","shell.execute_reply.started":"2024-12-16T08:38:24.192293Z","shell.execute_reply":"2024-12-16T08:38:25.299692Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_cat = df_train.select_dtypes(include=['object'])\n\n# it is good to look at the list of distinct values and check for ordinal variables\nuniques = []\nfor f in df_cat.columns:\n    item = {'feature':f}\n    item['No. of Unique'] = df_cat[f].nunique()\n    item['values'] = df_cat[f].unique().tolist()\n    value_counts = df_cat[f].value_counts(normalize=True) * 100\n    item['percentages'] = [round(value_counts.get(val, 0),0) for val in item['values']]\n    uniques.append(item)\n\ndf_uniques = pd.DataFrame(uniques)\ndf_uniques = df_uniques.set_index('feature')\ndf_uniques","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:38:25.301828Z","iopub.execute_input":"2024-12-16T08:38:25.302202Z","iopub.status.idle":"2024-12-16T08:38:28.369945Z","shell.execute_reply.started":"2024-12-16T08:38:25.302162Z","shell.execute_reply":"2024-12-16T08:38:28.369079Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Group by 'Customer Feedback' and calculate both mean and count for 'Premium Amount'\nFeedback_stats = df_train.groupby('Customer Feedback')['Premium Amount'].agg(mean='mean')\n\nFeedback_stats.sort_values(['mean'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:38:28.371114Z","iopub.execute_input":"2024-12-16T08:38:28.371496Z","iopub.status.idle":"2024-12-16T08:38:28.471116Z","shell.execute_reply.started":"2024-12-16T08:38:28.371455Z","shell.execute_reply":"2024-12-16T08:38:28.470352Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Group by 'Customer Feedback' and calculate both mean and count for 'Premium Amount'\nLocation_stats = df_train.groupby('Location')['Premium Amount'].agg(mean='mean')\n\nLocation_stats.sort_values(['mean'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:38:28.472028Z","iopub.execute_input":"2024-12-16T08:38:28.472313Z","iopub.status.idle":"2024-12-16T08:38:28.545644Z","shell.execute_reply.started":"2024-12-16T08:38:28.472287Z","shell.execute_reply":"2024-12-16T08:38:28.544749Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(10, 5))\n\n# Create a histogram with KDE overlay\nsns.histplot(data=df_train, x='Premium Amount', hue='Gender', \n             bins=30, kde=True, alpha=0.5)\nplt.title('Premium Amount With Gender', fontsize=18, weight='bold')\nplt.legend()  \nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:38:28.546776Z","iopub.execute_input":"2024-12-16T08:38:28.547428Z","iopub.status.idle":"2024-12-16T08:38:34.410844Z","shell.execute_reply.started":"2024-12-16T08:38:28.547379Z","shell.execute_reply":"2024-12-16T08:38:34.409953Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sns.violinplot(x = \"Premium Amount\",\n            y = \"Education Level\",\n            kind = \"box\",\n            height = 6,\n            aspect = 1.8,\n            color = \"#FBC02D\",\n            data = df_train).set(title = \"Premium Amount by Education Level\");","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:38:34.414493Z","iopub.execute_input":"2024-12-16T08:38:34.414774Z","iopub.status.idle":"2024-12-16T08:38:37.114011Z","shell.execute_reply.started":"2024-12-16T08:38:34.414747Z","shell.execute_reply":"2024-12-16T08:38:37.113082Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sns.violinplot(x = \"Premium Amount\",\n            y = 'Policy Type',\n            kind = \"box\",\n            height = 6,\n            aspect = 1.8,\n            color = \"#FBC02D\",\n            data = df_train).set(title = \"Premium Amount by Education Level\");","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:38:37.115084Z","iopub.execute_input":"2024-12-16T08:38:37.115378Z","iopub.status.idle":"2024-12-16T08:38:39.849011Z","shell.execute_reply.started":"2024-12-16T08:38:37.115348Z","shell.execute_reply":"2024-12-16T08:38:39.84817Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# drop id\ndf_train = df_train.drop(columns=['id'])\ndf_test = df_test.drop(columns=['id'])\n\n# Convert Policy Start Date to datetime\ndf_train['Policy Start Date'] = pd.to_datetime(df_train['Policy Start Date'])\ndf_test['Policy Start Date'] = pd.to_datetime(df_test['Policy Start Date'])","metadata":{"execution":{"iopub.status.busy":"2024-12-16T08:38:39.850155Z","iopub.execute_input":"2024-12-16T08:38:39.850562Z","iopub.status.idle":"2024-12-16T08:38:40.498318Z","shell.execute_reply.started":"2024-12-16T08:38:39.85052Z","shell.execute_reply":"2024-12-16T08:38:40.497575Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h3 align=\"left\"><font color='#DAA520'>Correlation Matrix:</font></h3>","metadata":{}},{"cell_type":"code","source":"!pip install dython","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:38:40.499332Z","iopub.execute_input":"2024-12-16T08:38:40.499599Z","iopub.status.idle":"2024-12-16T08:38:49.541559Z","shell.execute_reply.started":"2024-12-16T08:38:40.499572Z","shell.execute_reply":"2024-12-16T08:38:49.540655Z"},"jupyter":{"source_hidden":true,"outputs_hidden":true},"collapsed":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from dython.nominal import associations\n\nassociations_df = associations(df_train[:10000], nominal_columns='all', plot=False)\ncorr_matrix = associations_df['corr']\nplt.figure(figsize=(20, 8))\nplt.gcf().set_facecolor('#FFFDD0') \nsns.heatmap(corr_matrix, annot=True, fmt='.2f', cmap='copper', linewidths=0.5)\nplt.title('Correlation Matrix including Categorical Features')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:38:49.542988Z","iopub.execute_input":"2024-12-16T08:38:49.543326Z","iopub.status.idle":"2024-12-16T08:39:11.821165Z","shell.execute_reply.started":"2024-12-16T08:38:49.543292Z","shell.execute_reply":"2024-12-16T08:39:11.820289Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"corr = corr_matrix['Premium Amount'].sort_values(ascending = False)\npd.DataFrame(corr).style.background_gradient(cmap='copper')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:39:11.822169Z","iopub.execute_input":"2024-12-16T08:39:11.822466Z","iopub.status.idle":"2024-12-16T08:39:11.869578Z","shell.execute_reply.started":"2024-12-16T08:39:11.822437Z","shell.execute_reply":"2024-12-16T08:39:11.868795Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h3 align=\"left\"><font color='#DAA520'>Missing Values:</font></h3>","metadata":{}},{"cell_type":"code","source":"# get the number of missing data points per column\nmissing_val_count_by_column = df_train.isnull().sum()\n\n# how many total missing values do we have from the entire dataset?\n# percent of data that is missing\npercent_missing = (missing_val_count_by_column.sum()/np.product(df_train.shape)) * 100\nprint(\"Missing cells %: {:,.4f}\".format(percent_missing))\n\n#missing data per column\ntotal = df_train.isnull().sum().sort_values(ascending=False)\npercent = (missing_val_count_by_column/df_train.isnull().count())*100\nmissing_data = pd.concat([total, percent], axis=1, keys=['Total', 'Percent'])\nmissing_data.head(20).style.background_gradient(cmap='copper')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:39:11.870451Z","iopub.execute_input":"2024-12-16T08:39:11.87069Z","iopub.status.idle":"2024-12-16T08:39:13.373249Z","shell.execute_reply.started":"2024-12-16T08:39:11.870666Z","shell.execute_reply":"2024-12-16T08:39:13.37226Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h3 align=\"left\"><font color='#DAA520'>Add More Features:</font></h3>","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\n# Define a function to apply the transformations\ndef feature_engineering(df):\n    # Age Group\n    df['Age Group'] = pd.cut(df['Age'], bins=[0, 18, 35, 50, 100], labels=['Under 18', 'Young Adult', 'Middle-Aged', 'Senior'])\n    \n    # Dependent Ratio\n    df['Dependent Ratio'] = df['Number of Dependents'] / (df['Annual Income'] / 1000)\n    \n    # Claim-to-Income Ratio\n    df['Claim-to-Income Ratio'] = df['Previous Claims'] / df['Annual Income']\n    \n\n    df['Health-Income Interaction'] = df['Health Score'] * df['Annual Income']\n    \n    # Date Features\n    df['Policy Year'] = df['Policy Start Date'].dt.year\n    df['Policy Day'] = df['Policy Start Date'].dt.day\n    df['Policy Month'] = df['Policy Start Date'].dt.month  # Ensure this is calculated\n    \n    # Cyclical Features\n    df['Policy Month Sin'] = np.sin(2 * np.pi * df['Policy Month'] / 12)\n    df['Policy Month Cos'] = np.cos(2 * np.pi * df['Policy Month'] / 12)\n    df['Policy Day Sin'] = np.sin(2 * np.pi * df['Policy Day'] / 31)\n    df['Policy Day Cos'] = np.cos(2 * np.pi * df['Policy Day'] / 31)\n    df['Policy Year Sin'] = np.sin(2 * np.pi * df['Policy Year'] / df['Policy Year'].max())\n    df['Policy Year Cos'] = np.cos(2 * np.pi * df['Policy Year'] / df['Policy Year'].max())\n    \n    return df\n\n# Apply the function to both datasets\ndf_train = feature_engineering(df_train)\ndf_test = feature_engineering(df_test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:39:13.374503Z","iopub.execute_input":"2024-12-16T08:39:13.374865Z","iopub.status.idle":"2024-12-16T08:39:13.933022Z","shell.execute_reply.started":"2024-12-16T08:39:13.374826Z","shell.execute_reply":"2024-12-16T08:39:13.93213Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train.drop(columns=['Policy Start Date'],axis=1, inplace=True)\ndf_test.drop(columns=['Policy Start Date'],axis=1, inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:39:13.933957Z","iopub.execute_input":"2024-12-16T08:39:13.934193Z","iopub.status.idle":"2024-12-16T08:39:14.218375Z","shell.execute_reply.started":"2024-12-16T08:39:13.934169Z","shell.execute_reply":"2024-12-16T08:39:14.217688Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n<div style=\"border-radius:5px; border:#6E2A04 solid; padding: 15px; background-color: #BED4D2; font-size:100%; text-align:left\">\n\n<h2 align=\"center\"><font color='#6E2A04'>Model Building</font></h2>\n","metadata":{}},{"cell_type":"code","source":"# Assuming `train` is your dataset, and 'Premium Amount' is the target variable\ny = df_train['Premium Amount']\nX = df_train.drop(['Premium Amount'], axis=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:39:14.219379Z","iopub.execute_input":"2024-12-16T08:39:14.219651Z","iopub.status.idle":"2024-12-16T08:39:14.335331Z","shell.execute_reply.started":"2024-12-16T08:39:14.219625Z","shell.execute_reply":"2024-12-16T08:39:14.334629Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Imputers for missing values\nnumeric_imputer = SimpleImputer(strategy='median')\ncategorical_imputer = SimpleImputer(strategy='most_frequent')\n\n# Preprocessing pipeline for numeric data\nnumeric_transformer = Pipeline(steps=[\n    ('imputer', numeric_imputer)\n])\n\n# Preprocessing pipeline for categorical data\ncategorical_transformer = Pipeline(steps=[\n    ('imputer', categorical_imputer)\n])\n\n\ncat_col = X.select_dtypes(include=['object']).columns\nnum_col = X.select_dtypes(include=['number']).columns\n# Combine numeric and categorical pipelines\npreprocessor = ColumnTransformer(\n    transformers=[\n        ('num', numeric_transformer, num_col),\n        ('cat', categorical_transformer, cat_col)\n    ]\n)\n\n# Fit and transform the training data\nX_train = preprocessor.fit_transform(X)\n\n# Transform the test data\nX_test = preprocessor.transform(df_test)\n\n\n# Convert transformed data back to DataFrame (optional, for interpretability)\nfeature_names = (\n    list(num_col) +  # Ensure this is a list\n    list(cat_col)\n)\ndf_train_processed = pd.DataFrame(X_train, columns=(feature_names))\ndf_test_processed = pd.DataFrame(X_test, columns=feature_names)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:39:14.336317Z","iopub.execute_input":"2024-12-16T08:39:14.33657Z","iopub.status.idle":"2024-12-16T08:39:22.489799Z","shell.execute_reply.started":"2024-12-16T08:39:14.336545Z","shell.execute_reply":"2024-12-16T08:39:22.489054Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_train, X_val, y_train, y_val = train_test_split(df_train_processed, y, test_size=0.2, random_state=123)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:39:22.490888Z","iopub.execute_input":"2024-12-16T08:39:22.491257Z","iopub.status.idle":"2024-12-16T08:39:26.587991Z","shell.execute_reply.started":"2024-12-16T08:39:22.491203Z","shell.execute_reply":"2024-12-16T08:39:26.587265Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def root_mean_squared_log_error(y_true, y_pred):\n    \"\"\"\n    Compute the Root Mean Squared Logarithmic Error (RMSLE).\n    \"\"\"\n    return np.sqrt(mean_squared_log_error(y_true, y_pred))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:39:26.589087Z","iopub.execute_input":"2024-12-16T08:39:26.589466Z","iopub.status.idle":"2024-12-16T08:39:26.594562Z","shell.execute_reply.started":"2024-12-16T08:39:26.589427Z","shell.execute_reply":"2024-12-16T08:39:26.593677Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def objective(trial):\n    # Suggest hyperparameters (remove colsample_bytree)\n    n_estimators = trial.suggest_int('n_estimators', 400, 1000)\n    learning_rate = trial.suggest_float('learning_rate', 0.01, 0.3)\n    reg_lambda = trial.suggest_float('reg_lambda', 0.0, 4.0)\n    reg_alpha = trial.suggest_float('reg_alpha', 0.0, 4.0)\n    max_depth = trial.suggest_int('max_depth', 2, 10)\n    gamma = trial.suggest_float('gamma', 0.0, 0.5)\n\n    \n    # Initialize the CatBoost model with GPU\n    model = CatBoostRegressor(\n        iterations=n_estimators,\n        learning_rate=learning_rate,\n        depth=max_depth,\n        l2_leaf_reg=reg_lambda,\n        random_seed=123,\n        task_type=\"GPU\",  # Enable GPU usage\n        devices='0',      # Specify the GPU device\n        verbose=0,\n        cat_features= list(cat_col),\n    )\n    \n    # Train model\n    model.fit(X_train, y_train)\n    \n    # Predict and calculate the error\n    y_pred = model.predict(X_val)\n    score = mean_squared_error(y_val, y_pred)\n    \n    return score\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:39:26.595756Z","iopub.execute_input":"2024-12-16T08:39:26.596401Z","iopub.status.idle":"2024-12-16T08:39:26.607566Z","shell.execute_reply.started":"2024-12-16T08:39:26.59636Z","shell.execute_reply":"2024-12-16T08:39:26.606879Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create an Optuna study\nstudy = optuna.create_study(direction='minimize', sampler=optuna.samplers.RandomSampler(seed=123))\noptuna.logging.set_verbosity(optuna.logging.WARNING)\n\n# Function to log the best trial\ndef log_best_trial(study, trial):\n    if study.best_trial == trial:\n        print(f\"New best trial: {trial.number} with value: {trial.value} and params: {trial.params}\")\n\n# Run the optimization\nstudy.optimize(objective, n_trials=50, callbacks=[log_best_trial])\n\n# Get the best parameters\nbest_params = study.best_params\nbest_score = study.best_value\nprint(f\"Best Hyperparameters: {best_params}\")\nprint(f\"Best MSE: {best_score:.6f}\")\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T08:39:26.608479Z","iopub.execute_input":"2024-12-16T08:39:26.608714Z","iopub.status.idle":"2024-12-16T09:16:26.480663Z","shell.execute_reply.started":"2024-12-16T08:39:26.608689Z","shell.execute_reply":"2024-12-16T09:16:26.479655Z"},"collapsed":true,"jupyter":{"outputs_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Get the best hyperparameters\nn_estimators = best_params['n_estimators']\nlearning_rate = best_params['learning_rate']\nmax_depth = best_params['max_depth']\nreg_lambda = best_params['reg_lambda']\ngamma = best_params['gamma']\n\n# Train the final CatBoost model with the best parameters\nbest_catboost =  CatBoostRegressor(\n        iterations=n_estimators,\n        learning_rate=learning_rate,\n        depth=max_depth,\n        l2_leaf_reg=reg_lambda,\n        random_seed=123,\n        task_type=\"GPU\",  # Enable GPU usage\n        devices='0',      # Specify the GPU device\n        verbose=0,\n        early_stopping_rounds=20, \n        cat_features= list(cat_col)\n\n    )\n    ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T09:16:26.481829Z","iopub.execute_input":"2024-12-16T09:16:26.482137Z","iopub.status.idle":"2024-12-16T09:16:26.487975Z","shell.execute_reply.started":"2024-12-16T09:16:26.482104Z","shell.execute_reply":"2024-12-16T09:16:26.486891Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Fit the model\neval_set = (X_val, y_val)\nbest_catboost.fit(X_train, y_train, eval_set=eval_set)\n\n# Make predictions\ny_pred = best_catboost.predict(X_val)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T09:16:26.489121Z","iopub.execute_input":"2024-12-16T09:16:26.489498Z","iopub.status.idle":"2024-12-16T09:17:08.601884Z","shell.execute_reply.started":"2024-12-16T09:16:26.48945Z","shell.execute_reply":"2024-12-16T09:17:08.601152Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Feature importance\nfeature_importances = best_catboost.get_feature_importance()\nfeature_names = df_train_processed.columns\n\n# Plot the feature importances\nplt.figure(figsize=(10, 6))\nplt.barh(feature_names, feature_importances)\nplt.xlabel('Feature Importance')\nplt.title('CatBoost Feature Importances')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T09:17:08.603053Z","iopub.execute_input":"2024-12-16T09:17:08.603706Z","iopub.status.idle":"2024-12-16T09:17:08.980059Z","shell.execute_reply.started":"2024-12-16T09:17:08.603664Z","shell.execute_reply":"2024-12-16T09:17:08.979279Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 3. Residual Analysis\nresiduals = y_val - y_pred\n\n# Residuals vs Predicted Values\nplt.figure(figsize=(12, 6))\nsns.scatterplot(x=y_pred, y=residuals, alpha=0.6)\nplt.axhline(y=0, color='red', linestyle='--', linewidth=1.5)\nplt.title(\"Residuals vs Predicted Values\", fontsize=18, fontweight='bold')\nplt.xlabel(\"Predicted Values\", fontsize=12)\nplt.ylabel(\"Residuals\", fontsize=12)\nplt.tight_layout()\nplt.show()\n\n# Residual Distribution\nplt.figure(figsize=(10, 6))\nsns.histplot(residuals, bins=30, kde=True)\nplt.axvline(x=0, color='red', linestyle='--', linewidth=1.5)\nplt.title(\"Distribution of Residuals\", fontsize=16, fontweight='bold')\nplt.xlabel(\"Residuals\", fontsize=12)\nplt.ylabel(\"Frequency\", fontsize=12)\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T09:17:08.981283Z","iopub.execute_input":"2024-12-16T09:17:08.98162Z","iopub.status.idle":"2024-12-16T09:17:11.031128Z","shell.execute_reply.started":"2024-12-16T09:17:08.981577Z","shell.execute_reply":"2024-12-16T09:17:11.030287Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n<div style=\"border-radius:5px; border:#6E2A04 solid; padding: 15px; background-color: #BED4D2; font-size:100%; text-align:left\">\n\n<h2 align=\"center\"><font color='#6E2A04'>Predict & Submission</font></h2>\n","metadata":{}},{"cell_type":"code","source":"results = best_catboost.predict(df_test_processed)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T09:17:11.032372Z","iopub.execute_input":"2024-12-16T09:17:11.033076Z","iopub.status.idle":"2024-12-16T09:17:13.229075Z","shell.execute_reply.started":"2024-12-16T09:17:11.03303Z","shell.execute_reply":"2024-12-16T09:17:13.228352Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission = pd.read_csv(\"/kaggle/input/playground-series-s4e12/sample_submission.csv\")\nsubmission['Premium Amount'] = results\nsubmission['Premium Amount'] = pt.inverse_transform(submission[['Premium Amount']])\nsubmission.to_csv('submission.csv', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-16T09:20:14.346139Z","iopub.execute_input":"2024-12-16T09:20:14.347061Z","iopub.status.idle":"2024-12-16T09:20:15.925646Z","shell.execute_reply.started":"2024-12-16T09:20:14.347022Z","shell.execute_reply":"2024-12-16T09:20:15.924776Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}