{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30823,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Uncovering Insights: Regression Analysis & Predictions on Insurance Data","metadata":{}},{"cell_type":"markdown","source":"This challenge focused on predicting insurance premiums based on various customer attributes. The competition aimed to evaluate the models using the Root Mean Squared Logarithmic Error (RMSLE), encouraging accurate predictions for both high and low premium values. The final submission file required predictions of the continuous target variable, Premium Amount, for each ID in the test set.","metadata":{}},{"cell_type":"markdown","source":"# Dataset Overview","metadata":{}},{"cell_type":"markdown","source":"* This is the Regression with Insurance Kaggle Competition\n* The Training Data consists of columns such as 1200000 rows and 21 columns\n* In Training Data the misisng values present in columns such as 18705 missing values present in the\tAge column, 44949  missing values present in the\tAnnual Income column, 18529\t missing values present in the Marital Status column, 109672  missing values present in the\tNumber of Dependents column, 358075  missing values present in the\tOccupation column, 74076  missing values present in the\tHealth Score column, 364029  missing values present in the\tPrevious Claims column, 6  missing values present in the\tVehicle Age column, 137882  missing values present in the\tCredit Score column, 1  missing values present in the\tInsurance Duration column & 77824  missing values present in the\tCustomer Feedback column\n* The Training Data consists of columns such as id, Age, Gender, Annual Income, Marital Status, Number of Dependents, Education Level, Occupation, Health Score, Location, Policy Type, Previous Claims, Vehicle Age, Credit Score, Insurance Duration, Policy Start Date, Customer Feedback, Smoking Status, Exercise Frequency, Property Type, Premium Amount\n* There is no duplicate present in the Training Data\n*  while the Test Data consists of 800000 rows and 20 columns\n*  The columns of Test Data includes id, Age, Gender, Annual Income, Marital Status, Number of Dependents, Education Level, Occupation, Health Score, Location, Policy Type, Previous Claims, Vehicle Age, Credit Score, Insurance Duration, Policy Start Date, Customer Feedback, Smoking Status, Exercise Frequency, Property Type\n*  There is also no duplicate present in Test Data\n* In Test Data there are 12489  missing values present in the Age column, 29860  missing values present in the\tAnnual Income column, 12336  missing values present in the\tMarital Status column, 73130  missing values present in the\tNumber of Dependents column, 239125\t missing values present in the Occupation column, 49449  missing values present in the Health Score column, 242802  missing values present in the\tPrevious Claims column, 3  missing values present in the Vehicle Age column, 91451  missing values present in the\tCredit Score column, 2  missing values present in the\tInsurance Duration column & 52276  missing values present in the\tCustomer Feedback column\n* The sample submission consists of 800000 rows and 2 columns and there is no duplicate and no misisng value present in sample submission\n* The sample submisison consists of columns such as id, Premium Amount\n\n## Competition Files\n\n| **File Name**              | **Description**                                                                 |\n|----------------------------|---------------------------------------------------------------------------------|\n| **train.csv**               |The training dataset; contains Premium Amount is the continuous target|\n| **test.csv**                |The test dataset; objective is to predict target Premium Amount for each row                |\n| **sample_submission.csv**   | A sample submission file; shows the correct format for predictions         |\n\n","metadata":{}},{"cell_type":"markdown","source":"## About Training Data Columns\n\n| **Column Name**           | **Description**                                |\n|---------------------------|------------------------------------------------|\n| **id**                    | Unique identifier for each customer            |\n| **Age**                   | Age of the customer                            |\n| **Gender**                | Gender of the customer (Male/Female)           |\n| **Annual Income**         | The annual income of the customer              |\n| **Marital Status**        | Marital status of the customer (Single/Married)|\n| **Number of Dependents**  | The number of dependents of the customer       |\n| **Education Level**       | The highest education level of the customer    |\n| **Occupation**            | The occupation of the customer                 |\n| **Health Score**          | A score representing the health condition of the customer |\n| **Location**              | The geographical location of the customer      |\n| **Policy Type**           | Type of insurance policy held by the customer |\n| **Previous Claims**       | Number of previous claims made by the customer |\n| **Vehicle Age**           | Age of the vehicle insured by the customer    |\n| **Credit Score**          | The customer's credit score                   |\n| **Insurance Duration**    | Duration for which the insurance policy is valid (in years) |\n| **Policy Start Date**     | The start date of the insurance policy        |\n| **Customer Feedback**     | Feedback rating given by the customer on the insurance policy |\n| **Smoking Status**        | Whether the customer is a smoker (Yes/No)      |\n| **Exercise Frequency**    | The frequency of exercise (e.g., Daily, Weekly) |\n| **Property Type**         | Type of property owned by the customer (e.g., House, Apartment) |\n","metadata":{}},{"cell_type":"markdown","source":"# About Author\nHi Kagglers! I'm Amin Abbasi, a passionate Data Scientist with keen interest in exploring and applying diverse data science techniques. As dedicated to derive meaningful insights and making impactful decisions through data, I actively engage in ML projects and contribute to real word problems and Kaggle by sharing detailed analysis and actionable insights. I'm excited to share my latest project Exploring and prediction the Insurance based on various factors. In this notebook, I to create various visualization plots using subplots to gain deep insights. Following this is I first Impute the missing values present in this Dataset as there are a lot of missing values present in this dataset and then Train and evaluate the CatBoost & LGBM Model using the Best prameters and Ensemble them then make predictions by implementing the Kfold Cross validation to improve results further on it\n\n| Name               | Email                                               | LinkedIn                                                  | GitHub                                           | Kaggle                                        |\n|--------------------|-----------------------------------------------------|-----------------------------------------------------------|--------------------------------------------------|-----------------------------------------------|\n| **Amin Abbasi**  | mabba@uic.edu | <a href=\"https://www.linkedin.com/in/maabba/\" style=\"text-decoration: none; font-size: 16px;\"><img src=\"https://img.shields.io/badge/LinkedIn-%2300A4CC.svg?style=for-the-badge&logo=LinkedIn&logoColor=white\" alt=\"LinkedIn Badge\"></a> | <a href=\"https://github.com/AminAb30\" style=\"text-decoration: none; font-size: 16px;\"><img src=\"https://img.shields.io/badge/GitHub-%23FF6F61.svg?style=for-the-badge&logo=GitHub&logoColor=white\" alt=\"GitHub Badge\"></a> | <a href=\"https://www.kaggle.com/sirvana\" style=\"text-decoration: none; font-size: 16px;\"><img src=\"https://img.shields.io/badge/Kaggle-%238a2be2.svg?style=for-the-badge&logo=Kaggle&logoColor=white\" alt=\"Kaggle Badge\"></a> |\n\n","metadata":{}},{"cell_type":"markdown","source":"# Import Libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom sklearn.preprocessing import OneHotEncoder, OrdinalEncoder\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.linear_model import LinearRegression\nfrom xgboost import XGBRegressor\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.metrics import mean_squared_log_error","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-05T23:43:39.381153Z","iopub.execute_input":"2025-01-05T23:43:39.381966Z","iopub.status.idle":"2025-01-05T23:43:41.597341Z","shell.execute_reply.started":"2025-01-05T23:43:39.381913Z","shell.execute_reply":"2025-01-05T23:43:41.596367Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Loading Dataset","metadata":{}},{"cell_type":"code","source":"data = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv', index_col='id')\ntest_df = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv', index_col='id')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-05T23:43:46.175783Z","iopub.execute_input":"2025-01-05T23:43:46.176309Z","iopub.status.idle":"2025-01-05T23:43:57.445138Z","shell.execute_reply.started":"2025-01-05T23:43:46.176261Z","shell.execute_reply":"2025-01-05T23:43:57.444039Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Explore Train and Test Datasets","metadata":{}},{"cell_type":"markdown","source":"## Explore Training Data Valiables","metadata":{}},{"cell_type":"markdown","source":"Make it more attractive with HDML codes","metadata":{}},{"cell_type":"code","source":"# Function to style tables\ndef style_table(df):\n    styled_df = df.style.set_table_styles([ \n        {\"selector\": \"th\", \"props\": [(\"color\", \"white\"), (\"background-color\", \"#B37F07\")]}  # Deep blue background for headers\n    ]).set_properties(**{\"text-align\": \"center\"}).hide(axis=\"index\")\n    return styled_df.to_html()\n\n# Function to generate random shades of color\ndef generate_random_color():\n    color = \"#{:02x}{:02x}{:02x}\".format(\n        random.randint(50, 150),\n        random.randint(50, 150),\n        random.randint(150, 255)\n    )\n    return color","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"print( df.shape[0])\n\nprint( train_df.shape[0])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-05T23:46:01.259576Z","iopub.execute_input":"2025-01-05T23:46:01.259945Z","iopub.status.idle":"2025-01-05T23:46:01.265021Z","shell.execute_reply.started":"2025-01-05T23:46:01.259909Z","shell.execute_reply":"2025-01-05T23:46:01.263833Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Feature Engeering","metadata":{}},{"cell_type":"code","source":"df['Policy Start Date'] = pd.DatetimeIndex(df['Policy Start Date'])\ndf['year'] = df['Policy Start Date'].dt.year\ndf['month'] = df['Policy Start Date'].dt.month\ndf['date'] = df['Policy Start Date'].dt.day\ndf['day'] = df['Policy Start Date'].dt.weekday","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:45:31.085236Z","iopub.execute_input":"2024-12-31T01:45:31.085591Z","iopub.status.idle":"2024-12-31T01:45:32.237044Z","shell.execute_reply.started":"2024-12-31T01:45:31.085558Z","shell.execute_reply":"2024-12-31T01:45:32.236378Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Checking NULL values","metadata":{}},{"cell_type":"code","source":"# Create a heatmap to visualize null values\nplt.figure(figsize=(12, 8))  # Adjust the figure size as needed\nsns.heatmap(df.isnull(), cbar=False, cmap='viridis', yticklabels=False)\n\nplt.title('Heatmap of Missing Values')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:45:32.238254Z","iopub.execute_input":"2024-12-31T01:45:32.2385Z","iopub.status.idle":"2024-12-31T01:46:09.920476Z","shell.execute_reply.started":"2024-12-31T01:45:32.23848Z","shell.execute_reply":"2024-12-31T01:46:09.919613Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"missing = df.isna().sum().reset_index()\nmissing.columns = ['features','missing_count']\nmissing['percentage'] = missing['missing_count']/df.shape[0]*100\n(missing[missing['missing_count']>0]\n .sort_values(by='missing_count',ascending=False)\n .reset_index(drop=True)\n .style.background_gradient())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:46:09.92143Z","iopub.execute_input":"2024-12-31T01:46:09.921751Z","iopub.status.idle":"2024-12-31T01:46:10.783397Z","shell.execute_reply.started":"2024-12-31T01:46:09.921718Z","shell.execute_reply":"2024-12-31T01:46:10.782569Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Cheking Outliers","metadata":{}},{"cell_type":"code","source":"# Extract numerical columns except 'Premium Amount'\ndf_numerical = df.select_dtypes(exclude=['object'])\n\n# Initialize variables for subplot layout\ncolumns = df_numerical.columns.drop('Premium Amount')  # Exclude 'Premium Amount'\nnum_plots = len(columns)\nrows = (num_plots + 2) // 3  # Calculate number of rows needed (3 plots per row)\n\n# Create subplots\nfig, axes = plt.subplots(rows, 3, figsize=(18, rows * 5))  # Adjust figure size for better readability\naxes = axes.flatten()  # Flatten axes array for easier indexing\n\n# Plot each column\nfor i, column in enumerate(columns):\n    ax = axes[i]\n    ax.scatter(df_numerical[column], df_numerical['Premium Amount'], alpha=0.7)\n    ax.set_title(f'{column} vs SalePrice')\n    ax.set_xlabel(column)\n    ax.set_ylabel('Premium Amount')\n    ax.grid(True)\n\n# Hide any unused subplot axes\nfor j in range(i + 1, len(axes)):\n    fig.delaxes(axes[j])\n\n# Adjust layout to prevent overlap\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:46:10.784465Z","iopub.execute_input":"2024-12-31T01:46:10.784778Z","iopub.status.idle":"2024-12-31T01:47:08.141387Z","shell.execute_reply.started":"2024-12-31T01:46:10.784758Z","shell.execute_reply":"2024-12-31T01:47:08.140504Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Data Preprocessing","metadata":{}},{"cell_type":"markdown","source":"## Handling NULL values.","metadata":{}},{"cell_type":"code","source":"# Identify columns with NULL values\ncolumns_with_null = df.columns[df.isnull().any()].tolist()\n\n# Display the result\nprint(\"Columns with NULL values:\")\nprint(columns_with_null)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:47:08.14219Z","iopub.execute_input":"2024-12-31T01:47:08.142499Z","iopub.status.idle":"2024-12-31T01:47:08.909087Z","shell.execute_reply.started":"2024-12-31T01:47:08.142472Z","shell.execute_reply":"2024-12-31T01:47:08.908341Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Numerical Columns**: Age, Annual Income, Health Score, Previous Claims, Vehicle Age, Credit Score, Insurance Duration.\n\n**Categorical Columns**: Marital Status, Number of Dependents, Occupation, Customer Feedback","metadata":{}},{"cell_type":"code","source":"# Mean Imputation:\ndf['Age'].fillna(df['Age'].mean(), inplace=True)\ndf['Annual Income'].fillna(df['Annual Income'].mean(), inplace=True)\ndf['Health Score'].fillna(df['Health Score'].mean(), inplace=True)\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:47:08.911314Z","iopub.execute_input":"2024-12-31T01:47:08.911529Z","iopub.status.idle":"2024-12-31T01:47:08.958847Z","shell.execute_reply.started":"2024-12-31T01:47:08.91151Z","shell.execute_reply":"2024-12-31T01:47:08.958131Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Median Imputation\ndf['Credit Score'].fillna(df['Credit Score'].median(), inplace=True)\ndf['Vehicle Age'].fillna(df['Vehicle Age'].median(), inplace=True)\ndf['Number of Dependents'].fillna(0, inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:47:08.960125Z","iopub.execute_input":"2024-12-31T01:47:08.960341Z","iopub.status.idle":"2024-12-31T01:47:09.064539Z","shell.execute_reply.started":"2024-12-31T01:47:08.960323Z","shell.execute_reply":"2024-12-31T01:47:09.06379Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Domain Specific Constants.\ndf['Previous Claims'].fillna(0, inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:47:09.065453Z","iopub.execute_input":"2024-12-31T01:47:09.065739Z","iopub.status.idle":"2024-12-31T01:47:09.080915Z","shell.execute_reply.started":"2024-12-31T01:47:09.065709Z","shell.execute_reply":"2024-12-31T01:47:09.080131Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Categorical Data MODE imputation.\ndf['Marital Status'].fillna(df['Marital Status'].mode()[0], inplace=True)\ndf['Occupation'].fillna(df['Occupation'].mode()[0], inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:47:09.081751Z","iopub.execute_input":"2024-12-31T01:47:09.081952Z","iopub.status.idle":"2024-12-31T01:47:09.51078Z","shell.execute_reply.started":"2024-12-31T01:47:09.081934Z","shell.execute_reply":"2024-12-31T01:47:09.509824Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Filling with `Unknown` value.\ndf['Customer Feedback'].fillna('Unknown', inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:47:09.511797Z","iopub.execute_input":"2024-12-31T01:47:09.51214Z","iopub.status.idle":"2024-12-31T01:47:09.615012Z","shell.execute_reply.started":"2024-12-31T01:47:09.512103Z","shell.execute_reply":"2024-12-31T01:47:09.614145Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# May have some logical sequence.\ndf['Marital Status'].fillna(method='ffill', inplace=True)\ndf['Occupation'].fillna(method='bfill', inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:47:09.615847Z","iopub.execute_input":"2024-12-31T01:47:09.6161Z","iopub.status.idle":"2024-12-31T01:47:09.967305Z","shell.execute_reply.started":"2024-12-31T01:47:09.616078Z","shell.execute_reply":"2024-12-31T01:47:09.966398Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# If there is no feedback then the customers are most likely satisfied\ndf['Customer Feedback'].fillna('Good', inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:47:09.96828Z","iopub.execute_input":"2024-12-31T01:47:09.968602Z","iopub.status.idle":"2024-12-31T01:47:10.04904Z","shell.execute_reply.started":"2024-12-31T01:47:09.968569Z","shell.execute_reply":"2024-12-31T01:47:10.048101Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Encoding the categorical columns.","metadata":{}},{"cell_type":"code","source":"# Get categorical columns\ncategorical_columns = df.select_dtypes(include=['object']).columns\n\n# Print the categorical columns\nprint(\"Categorical Columns:\")\nprint(categorical_columns)\n\n\n# Use pandas get_dummies for One-Hot Encoding\ndf = pd.get_dummies(df, columns=['Gender', 'Marital Status', 'Education Level', \n                                 'Occupation', 'Location', 'Policy Type', \n                                 'Smoking Status', 'Property Type', 'Exercise Frequency',\n                                 'Customer Feedback'], drop_first=True)\n\nprint(df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:47:10.050028Z","iopub.execute_input":"2024-12-31T01:47:10.050359Z","iopub.status.idle":"2024-12-31T01:47:12.736446Z","shell.execute_reply.started":"2024-12-31T01:47:10.050332Z","shell.execute_reply":"2024-12-31T01:47:12.735507Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Preparing Dataset for Training and Testing","metadata":{}},{"cell_type":"code","source":"training_data = df[0:len(train_df)]\ntesting_data = df[len(train_df):]\ntesting_data = testing_data.drop(columns='Premium Amount')\n\nX = training_data.drop(columns='Premium Amount')\ny = training_data['Premium Amount']\n# Don't get afraid, as the evaluation will be done based on Root Mean Squared Logarithmic Error (RMSLE)\n# That's why using this log1p.\ny_log1p = np.log1p(y)\n\nX_train, X_test, y_train, y_test = train_test_split(X, y_log1p, test_size = 0.2)\ny_train = np.reshape(y_train,(-1, 1))\ny_test = np.reshape(y_test,(-1, 1))\nX_train.shape, y_train.shape\n\nX['Insurance Duration'].fillna(X['Insurance Duration'].median(), inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:47:12.737519Z","iopub.execute_input":"2024-12-31T01:47:12.737837Z","iopub.status.idle":"2024-12-31T01:47:13.108574Z","shell.execute_reply.started":"2024-12-31T01:47:12.737806Z","shell.execute_reply":"2024-12-31T01:47:13.107854Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Model Initialization","metadata":{}},{"cell_type":"code","source":"# Calculate RMSLE\ndef rmsle(y_true,y_pred):\n    return np.sqrt(mean_squared_log_error(y_true,y_pred))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:47:13.109292Z","iopub.execute_input":"2024-12-31T01:47:13.10974Z","iopub.status.idle":"2024-12-31T01:47:13.113331Z","shell.execute_reply.started":"2024-12-31T01:47:13.109713Z","shell.execute_reply":"2024-12-31T01:47:13.112651Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Model 1: (CatBoostRegressor)","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import KFold\nfrom catboost import CatBoostRegressor \nfrom sklearn.metrics import mean_squared_log_error\n\ndef train_model():\n    kf=KFold(n_splits=5,shuffle=True,random_state=42)\n    oof=np.zeros(len(X))\n    models=[]\n    for fold,(train_idx,valid_idx) in enumerate(kf.split(X)):\n        print(f\"fold{fold+1}\")\n        x_train,x_valid=X.iloc[train_idx],X.iloc[valid_idx]\n        y_train,y_valid=y_log1p.iloc[train_idx],y_log1p.iloc[valid_idx]\n        model=CatBoostRegressor(\n            iterations=2000,\n            learning_rate=0.05,\n            depth=6,\n            eval_metric=\"RMSE\",\n            random_seed=42,\n            verbose=200,\n            task_type=\"GPU\",\n            l2_leaf_reg=0.7\n        )\n        model.fit(x_train,y_train,eval_set=(x_valid,y_valid),early_stopping_rounds=300,)\n        models.append(model)\n        oof[valid_idx]=np.maximum(0,model.predict(x_valid))\n        fold_rmsle=rmsle(np.expm1(y_valid),np.expm1(oof[valid_idx]))\n        print(f'fold{fold+1}RMSLE:{fold_rmsle}')\n    return models,oof","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:47:13.11423Z","iopub.execute_input":"2024-12-31T01:47:13.114462Z","iopub.status.idle":"2024-12-31T01:47:13.320161Z","shell.execute_reply.started":"2024-12-31T01:47:13.114433Z","shell.execute_reply":"2024-12-31T01:47:13.31948Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"models,oof=train_model()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:47:13.32094Z","iopub.execute_input":"2024-12-31T01:47:13.321191Z","iopub.status.idle":"2024-12-31T01:48:29.623746Z","shell.execute_reply.started":"2024-12-31T01:47:13.321171Z","shell.execute_reply":"2024-12-31T01:48:29.622778Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"sample=pd.read_csv('/kaggle/input/playground-series-s4e12/sample_submission.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:48:29.624677Z","iopub.execute_input":"2024-12-31T01:48:29.624922Z","iopub.status.idle":"2024-12-31T01:48:29.899352Z","shell.execute_reply.started":"2024-12-31T01:48:29.624902Z","shell.execute_reply":"2024-12-31T01:48:29.898443Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"predictions=np.zeros(len(testing_data))\nfor model in models:\n    predictions+=np.maximum(0,np.expm1(model.predict(testing_data)))/len(models)\nsample[\"Premium Amount\"]=predictions \nsample.to_csv(\"output.csv\",index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T01:52:40.615396Z","iopub.execute_input":"2024-12-31T01:52:40.615751Z","iopub.status.idle":"2024-12-31T01:52:46.35178Z","shell.execute_reply.started":"2024-12-31T01:52:40.615721Z","shell.execute_reply":"2024-12-31T01:52:46.351095Z"}},"outputs":[],"execution_count":null}]}