{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🚀Regression with an Insurance Dataset - Locket Start🚀\nWelcome to the Locket Start EDA of our insurance dataset.  \nIn this Notebook, we explore the data, understand the distribution, missing values, and important features, and then build a prediction model using CatBoost. \n\n- **ver. 3**: Added feature importance visualization and discussion.\n- **ver. 2**: Implemented training on GPU for faster computation.\n- **ver. 1**: Basic inference using CatBoost.","metadata":{}},{"cell_type":"markdown","source":"# 1. 🔥Import Libraries\nImport all necessary libraries at the beginning for clarity.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport pandas as pd\nimport sys\nimport os\n\npd.set_option('display.max_columns', 500)\npd.set_option('display.max_rows', 5000)\nsns.set_style('whitegrid')\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:38:36.443069Z","iopub.execute_input":"2024-12-08T13:38:36.443403Z","iopub.status.idle":"2024-12-08T13:38:36.448587Z","shell.execute_reply.started":"2024-12-08T13:38:36.443374Z","shell.execute_reply":"2024-12-08T13:38:36.447739Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 2. 👀Load and Preview the Data\nLoad the dataset and perform an initial inspection.","metadata":{}},{"cell_type":"code","source":"input = '/kaggle/input/playground-series-s4e12'\ntrain = pd.read_csv(os.path.join(input, 'train.csv'), index_col='id')\ntest = pd.read_csv(os.path.join(input, 'test.csv'), index_col='id')\ntrain.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:38:36.449946Z","iopub.execute_input":"2024-12-08T13:38:36.450225Z","iopub.status.idle":"2024-12-08T13:38:42.183538Z","shell.execute_reply.started":"2024-12-08T13:38:36.4502Z","shell.execute_reply":"2024-12-08T13:38:42.182719Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.1. Data Overview\nLet's take a quick look at the data to understand its structure.","metadata":{}},{"cell_type":"code","source":"print(f\"{'-'*10}train{'-'*10}\")\nprint(train.info())\nprint(f\"{'-'*10}test{'-'*10}\")\nprint(test.info())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:38:42.184672Z","iopub.execute_input":"2024-12-08T13:38:42.184947Z","iopub.status.idle":"2024-12-08T13:38:43.089646Z","shell.execute_reply.started":"2024-12-08T13:38:42.184922Z","shell.execute_reply":"2024-12-08T13:38:43.088773Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.2. Data Description\n| Name                | Dtype    | Missing Values (Train/Test)     | Description                                          |\n|---------------------|----------|----------------------------------|------------------------------------------------------|\n| Age                 | float64  | Train: 18705 / Test: 12489       | Age of the policyholder                              |\n| Gender              | object   | Train: 0 / Test: 0               | Gender of the policyholder                           |\n| Annual Income       | float64  | Train: 44949 / Test: 29860       | Yearly income                                        |\n| Marital Status      | object   | Train: 18529 / Test: 12336       | Marital status of the policyholder                   |\n| Number of Dependents| float64  | Train: 109672 / Test: 73130      | Number of dependents                                 |\n| Education Level     | object   | Train: 0 / Test: 0               | Highest education level achieved                     |\n| Occupation          | object   | Train: 358075 / Test: 239125     | Occupation of the policyholder                       |\n| Health Score        | float64  | Train: 74076 / Test: 49449       | Health score                                         |\n| Location            | object   | Train: 0 / Test: 0               | Geographical location of the policyholder            |\n| Policy Type         | object   | Train: 0 / Test: 0               | Type of insurance policy                             |\n| Previous Claims     | float64  | Train: 364029 / Test: 242802     | Number of previous claims                            |\n| Vehicle Age         | float64  | Train: 6 / Test: 3               | Age of the insured vehicle                           |\n| Credit Score        | float64  | Train: 137882 / Test: 91451      | Credit score of the policyholder                     |\n| Insurance Duration  | float64  | Train: 1 / Test: 2               | Duration of the insurance policy                     |\n| Policy Start Date   | object   | Train: 0 / Test: 0               | Start date of the insurance policy                   |\n| Customer Feedback   | object   | Train: 77824 / Test: 52276       | Feedback provided by the customer                    |\n| Smoking Status      | object   | Train: 0 / Test: 0               | Indicates whether the customer is a smoker           |\n| Exercise Frequency  | object   | Train: 0 / Test: 0               | Frequency of exercise by the policyholder            |\n| Property Type       | object   | Train: 0 / Test: 0               | Type of property owned                               |\n| Premium Amount      | float64  | Train: 0 / Test: N/A             | Insurance premium amount (Target)                       |\n\n","metadata":{}},{"cell_type":"markdown","source":"# 4. 📈Exploratory Data Analysis (EDA)","metadata":{}},{"cell_type":"markdown","source":"## 4.1. Target Variable Distribution\n","metadata":{}},{"cell_type":"code","source":"target = 'Premium Amount'\nplt.figure(figsize=(10,6))\nsns.histplot(train[target], kde=True, palette='viridis')\nplt.title(f\"{target} Distribution (Train)\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:38:43.091231Z","iopub.execute_input":"2024-12-08T13:38:43.091529Z","iopub.status.idle":"2024-12-08T13:38:48.181734Z","shell.execute_reply.started":"2024-12-08T13:38:43.0915Z","shell.execute_reply":"2024-12-08T13:38:48.180903Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.2.Numeric Features EDA\n\nIn this section, we will explore the numeric features. \nWe will check their distributions, summary statistics, and relationships with the target. \n\n**Steps**\n1. Examine the distribution of numeric features\n2. Check correlation matrix","metadata":{}},{"cell_type":"code","source":"# Separate numeric and categorical columns\nnumeric_cols = train.select_dtypes(include=[np.number]).columns.tolist()\nnumeric_cols = [col for col in numeric_cols if col != target]  # excluding target\n\n# ----- Numeric Features EDA -----\nprint(\"Numeric Features EDA\")\nprint(\"Checking distributions of numeric features... \")\nnum_bin = 10\nfor col in numeric_cols:\n    print(f\"{col} Analytics\")\n    print(f\"isnull:\\n   Train: {train[col].isnull().sum()}\\n    Test: {test[col].isnull().sum()}\")\n    train[f\"{col}_bins\"] = pd.qcut(train[col], q=num_bin, duplicates='drop')\n    fig, ax = plt.subplots(1, 3, figsize=(12, 4))\n    sns.violinplot(x=f\"{col}_bins\", y=target, data=train, palette='viridis', ax=ax[0])\n    sns.histplot(train[col], kde=True, palette='viridis', ax=ax[1])\n    sns.histplot(test[col], kde=True, palette='viridis', ax=ax[2])\n    fig.suptitle(f\"{col} Distribution\")\n    ax[0].tick_params(axis='x', rotation=45)\n    ax[0].set_title(f\"{col} vs {target}\")\n    ax[1].set_title(\"Train Data\")\n    ax[2].set_title(\"Test Data\")\n    plt.show()\n    train.drop(columns=f\"{col}_bins\", inplace=True)\n\nprint(\"Checking correlation matrix... \")\nplt.figure(figsize=(10,6))\ncorr = train[numeric_cols + [target]].corr()\nsns.heatmap(corr, annot=True, cmap=\"coolwarm\")\nplt.title(\"Correlation Matrix\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:38:48.182953Z","iopub.execute_input":"2024-12-08T13:38:48.183335Z","iopub.status.idle":"2024-12-08T13:40:13.889809Z","shell.execute_reply.started":"2024-12-08T13:38:48.183284Z","shell.execute_reply":"2024-12-08T13:40:13.888885Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4.3. Categorical Features EDA\n\nIn this section, we will analyze the categorical features. \nWe will look at their frequency counts, unique values, and how they relate to the target. \n\n**Steps**\n1. Count values for each categorical feature\n2. Check the number of unique categories\n3. Explore the relationship with the target","metadata":{}},{"cell_type":"code","source":"# ----- Categorical Features EDA -----\nprint(\"Categorical Features EDA\")\ncategorical_cols = train.select_dtypes(exclude=[np.number]).columns.tolist()\nprint(\"Counting values for categorical features... \")\nfor col in categorical_cols:\n    top_train_x = train[col].value_counts().sort_values(ascending=False).head(20).index\n    top_test_x = test[col].value_counts().sort_values(ascending=False).head(20).index\n    \n    top_train = train[train[col].isin(top_train_x)]\n    top_test = test[test[col].isin(top_test_x)]\n    \n    print(f\"{col} Analytics\")\n    print(f\"value_counts:\\n{train[col].value_counts()}\")\n    print(f\"isnull:\\n{train[col].isnull().sum()}\")\n    \n    fig, ax = plt.subplots(1, 3, figsize=(16, 4))\n    sns.violinplot(x=col, y=target, data=top_train, ax=ax[0], palette='viridis')\n    sns.countplot(x=col, data=top_train, order=top_train[col].value_counts().index, ax=ax[1], palette='viridis')\n    sns.countplot(x=col, data=top_test, order=top_test[col].value_counts().index, ax=ax[2], palette='viridis')\n    fig.suptitle(f\"{col} Value Counts\")\n    ax[0].set_title(f\"{col} vs {target}\")\n    ax[1].set_title(\"Train Data\")\n    ax[2].set_title(\"Test Data\")\n    for i in range(3):\n        ax[i].tick_params(axis='x', rotation=45)\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:40:13.891053Z","iopub.execute_input":"2024-12-08T13:40:13.891338Z","iopub.status.idle":"2024-12-08T13:41:00.160853Z","shell.execute_reply.started":"2024-12-08T13:40:13.891311Z","shell.execute_reply":"2024-12-08T13:41:00.160037Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 5. 🔨Data Preprocessing\nBefore training the CatBoost model, let's preprocess the data.\n\n**step**\n- Convert datetime columns to a proper datetime format and extract features such as year and month.\n- Handle missing values in both numeric and categorical features. \n- CatBoost can handle categorical features directly, so we don't need one-hot encoding, but we must identify categorical features. \n- Split the training data into training and validation sets.\n\n## 5.1. Convert datetime\nConvert Policy Start Date to datetime and extract year, month, day features.","metadata":{}},{"cell_type":"code","source":"for df in [train, test]:\n    df['Policy Start Date'] = pd.to_datetime(df['Policy Start Date'], errors='coerce')\n    df['Policy_Start_Year'] = df['Policy Start Date'].dt.year\n    df['Policy_Start_Month'] = df['Policy Start Date'].dt.month\n    df['Policy_Start_Day'] = df['Policy Start Date'].dt.day","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:41:00.162112Z","iopub.execute_input":"2024-12-08T13:41:00.162739Z","iopub.status.idle":"2024-12-08T13:41:00.919334Z","shell.execute_reply.started":"2024-12-08T13:41:00.162688Z","shell.execute_reply":"2024-12-08T13:41:00.918447Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5.2. Handle missing values\nEveryone has his or her own way of filling in missing values, **so I encourage you to try different!**","metadata":{}},{"cell_type":"code","source":"numeric_cols = train.select_dtypes(include=[np.number]).columns.drop(target, errors='ignore').tolist()\ncat_cols = train.select_dtypes(exclude=[np.number, 'datetime64[ns]']).columns.tolist()\n\n# Impute missing values in numeric columns with median.\nfor col in numeric_cols:\n    median_value = train[col].median()\n    train[col] = train[col].fillna(median_value)\n    test[col] = test[col].fillna(median_value)\n\n# Impute missing values in categorical columns with \"Unknown\". \nfor col in cat_cols:\n    train[col] = train[col].astype(str).fillna(\"Unknown\")\n    test[col] = test[col].astype(str).fillna(\"Unknown\")\n\n# Drop the original datetime column if not needed.\ntrain = train.drop('Policy Start Date', axis=1)\ntest = test.drop('Policy Start Date', axis=1)\n\n# Update numeric_cols and cat_cols if needed.\nnumeric_cols = train.select_dtypes(include=[np.number]).columns.drop(target, errors='ignore').tolist()\ncat_cols = train.select_dtypes(exclude=[np.number]).columns.drop(target, errors='ignore').tolist()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:41:00.920545Z","iopub.execute_input":"2024-12-08T13:41:00.920836Z","iopub.status.idle":"2024-12-08T13:41:03.962071Z","shell.execute_reply.started":"2024-12-08T13:41:00.920793Z","shell.execute_reply":"2024-12-08T13:41:03.961094Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5.3. Split the data\nSplit the training data into training and validation sets.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX = train.drop(columns=target)\ny = train[target]\n\nX_train, X_val, y_train, y_val = train_test_split(\n    X, y, test_size=0.2, random_state=42\n)\n\n# Identify categorical feature indices for CatBoost\ncat_features_indices = [X_train.columns.get_loc(c) for c in cat_cols if c in X_train.columns]\n\nprint(\"Data Preprocessing Done!\")\nprint(\"Train shape:\", X_train.shape)\nprint(\"Validation shape:\", X_val.shape)\nprint(\"Test shape:\", test.shape)\nprint(\"Categorical feature indices:\", cat_features_indices)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:41:03.963376Z","iopub.execute_input":"2024-12-08T13:41:03.963727Z","iopub.status.idle":"2024-12-08T13:41:05.052272Z","shell.execute_reply.started":"2024-12-08T13:41:03.963689Z","shell.execute_reply":"2024-12-08T13:41:05.051166Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 6. 🐈Model Training\n## 6.1. Catboost model training by cross-validation\nTrain Catboost with cross-validation. **Be creative in how you set hyperparameters. (Grid Search, Random Search and Optuna, etc...)**","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import KFold\nfrom sklearn.metrics import mean_squared_log_error\nfrom catboost import CatBoostRegressor\n\n# CatBoost parameters\nparams = {\n    'iterations': 1000,\n    'learning_rate': 0.1,\n    'depth': 6,\n    'loss_function': 'RMSE',\n    'random_seed': 42,\n    'task_type': 'GPU',\n    'verbose': 0\n}\n\nn_splits = 5\nkf = KFold(n_splits=n_splits, shuffle=True, random_state=42)\nfold_scores = []\nmodels = []\n\nfor i, (train_idx, val_idx) in enumerate(kf.split(X_train)):\n    X_train_fold, X_val_fold = X_train.iloc[train_idx], X_train.iloc[val_idx]\n    y_train_fold, y_val_fold = y_train.iloc[train_idx], y_train.iloc[val_idx]\n\n    # Initialize CatBoost model\n    model = CatBoostRegressor(**params, cat_features=cat_features_indices)\n    \n    # Fit model on this fold\n    model.fit(X_train_fold, np.log1p(y_train_fold), eval_set=(X_val_fold, np.log1p(y_val_fold)), early_stopping_rounds=50, use_best_model=True)\n    \n    # Predict on validation fold\n    y_pred_fold = np.expm1(model.predict(X_val_fold))\n    \n    # Calculate RMSLE for this fold\n    fold_rmsle = np.sqrt(mean_squared_log_error(y_val_fold, y_pred_fold))\n    fold_scores.append(fold_rmsle)\n    models.append(model)\n    \n    print(f\"Fold {i+1} RMSLE: {fold_rmsle}\")\n\nmean_rmsle = np.mean(fold_scores)\nprint(f\"Mean RMSLE: {mean_rmsle:.5f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:41:05.055284Z","iopub.execute_input":"2024-12-08T13:41:05.055553Z","iopub.status.idle":"2024-12-08T13:43:32.509564Z","shell.execute_reply.started":"2024-12-08T13:41:05.055528Z","shell.execute_reply":"2024-12-08T13:43:32.508643Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6.2. Valid and Predict\nPredict on validation and test data using each model and average the predictions.","metadata":{}},{"cell_type":"code","source":"val_preds_list = []\nfor i, model in enumerate(models):\n    preds = model.predict(X_val)\n    val_preds_list.append(np.expm1(preds))\n\nval_preds_array = np.column_stack(val_preds_list)\nval_final_preds = np.mean(val_preds_array, axis=1)\n\nprint(\"Averaged Validation Predictions Shape:\", val_final_preds.shape)\nprint(\"Averaged Validation Predictions RMSLE:\", np.sqrt(mean_squared_log_error(y_val, val_final_preds)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:43:32.510749Z","iopub.execute_input":"2024-12-08T13:43:32.51113Z","iopub.status.idle":"2024-12-08T13:43:34.554474Z","shell.execute_reply.started":"2024-12-08T13:43:32.511091Z","shell.execute_reply":"2024-12-08T13:43:34.553563Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6.3. Feature Importance Visualization\nGet feature importance from the CatBoost model","metadata":{}},{"cell_type":"code","source":"feature_importances_list = []\n\nfor i, m in enumerate(models):\n    fi = m.get_feature_importance()\n    feature_importances_list.append(fi)\n\nfeature_importances_array = np.array(feature_importances_list)\navg_feature_importances = feature_importances_array.mean(axis=0)\n\nfi_df = pd.DataFrame({\n    'feature': X_train.columns,\n    'importance': avg_feature_importances\n}).sort_values('importance', ascending=False)\n\n# Plot the averaged feature importance\nplt.figure(figsize=(10,6))\nsns.barplot(x='importance', y='feature', data=fi_df, palette='viridis')\nplt.title(\"Averaged Feature Importance Across Models\")\nplt.xlabel(\"Importance\")\nplt.ylabel(\"Feature\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:43:48.594119Z","iopub.execute_input":"2024-12-08T13:43:48.594459Z","iopub.status.idle":"2024-12-08T13:43:49.104231Z","shell.execute_reply.started":"2024-12-08T13:43:48.594429Z","shell.execute_reply":"2024-12-08T13:43:49.103429Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The importance of `Annual Income`, `Credit Score`, and `Health Score` is high, while the importance of categorical features such as `Smoking Type` is low.","metadata":{}},{"cell_type":"code","source":"test_preds_list = []\nfor i, model in enumerate(models):\n    preds = model.predict(test)\n    test_preds_list.append(np.expm1(preds))\n\ntest_preds_array = np.column_stack(test_preds_list)\ntest_final_preds = np.mean(test_preds_array, axis=1)\nprint(\"Averaged Test Predictions Shape:\", test_final_preds.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:43:34.55555Z","iopub.execute_input":"2024-12-08T13:43:34.555809Z","iopub.status.idle":"2024-12-08T13:43:41.156447Z","shell.execute_reply.started":"2024-12-08T13:43:34.555783Z","shell.execute_reply":"2024-12-08T13:43:41.155456Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 7. 📄Submit\nAssume we have test_final_preds as the final predictions for test data.","metadata":{}},{"cell_type":"code","source":"submission = pd.DataFrame({\n    'id': test.index,            # Use the index or a column from test_df as 'id'\n    'Premium Amount': test_final_preds  # The predicted premium amounts\n})\n\nsubmission.to_csv('submission.csv', index=False)\nprint(\"Submission file created: submission.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-08T13:43:41.15751Z","iopub.execute_input":"2024-12-08T13:43:41.157776Z","iopub.status.idle":"2024-12-08T13:43:42.552344Z","shell.execute_reply.started":"2024-12-08T13:43:41.157749Z","shell.execute_reply":"2024-12-08T13:43:42.551477Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 8. 🙌Summary\n\nIn this notebook, we performed the following steps:\n\n1. **Data Loading and EDA**:  \n   We loaded the training and test datasets, explored the distribution of features, checked missing values, and visualized basic relationships.\n\n2. **Data Preprocessing**:  \n   We converted datetime columns, handled missing values, and prepared numeric and categorical features. We used CatBoost’s ability to handle categorical features directly. \n\n3. **Model Training with CatBoost**:  \n   We trained a CatBoost Regressor with cross-validation using KFold. We used RMSLE as the main evaluation metric and improved the model by transforming the target with `log1p()`/`expm1()` for better performance.\n\n4. **Ensembling and Averaging Predictions**:  \n   We stored multiple models in a `models` list and averaged their predictions for validation and test sets to achieve a more robust prediction.\n\n5. **GPU Utilization**:  \n   We demonstrated how to enable GPU training in CatBoost for faster model training.\n\n6. **Submission File Creation**:  \n   Finally, we generated the `submission.csv` file by combining the test set predictions and `id`. This file can be uploaded as the final solution.\n\n**🚀Next Steps**:  \nWe can further tune the model parameters, try different feature engineering techniques, or experiment with other ensemble methods to improve performance.","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}