{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Premium Prediction in the Insurance Industry\n\nPremium prediction in the insurance industry refers to the process of **estimating the amount of insurance premium** a customer should pay. This amount is typically determined based on various **personal and policy-related features**. \n\n### Key Factors Influencing Premium Prediction:\nThe premium is usually based on the perceived **risk** associated with the customer, which may depend on the following factors:\n- **Age**: Older or younger individuals may have different risk profiles.\n- **Income**: Higher income individuals might have different types of coverage, influencing their premiums.\n- **Vehicle Type**: The age, make, and model of the vehicle can affect the premium calculation.\n- **Claim History**: Previous claims made by the customer can influence future premium rates.\n- **Other Factors**: Additional data such as location, driving history, and the number of dependents may also play a role.\n\nBy analyzing these factors, insurers can estimate the appropriate premium for each individual, ensuring that the risk is properly assessed and that customers are charged accordingly.\n","metadata":{}},{"cell_type":"markdown","source":"## Step 1: Load Data","metadata":{}},{"cell_type":"code","source":"\n# Importing libraries\nimport pandas as pd\n\n# Load train, test, and sample submission datasets\ntrain_path = '/kaggle/input/playground-series-s4e12/train.csv'\ntest_path = '/kaggle/input/playground-series-s4e12/test.csv'\ntrain_data = pd.read_csv(train_path)\ntest_data = pd.read_csv(test_path)\n\n# Display the first few rows of the train dataset\nprint(\"Train Data:\")\nprint(train_data.head(2))\n\nprint(\"\\nTest Data:\")\nprint(test_data.head(2))\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-14T11:34:37.440774Z","iopub.execute_input":"2024-12-14T11:34:37.441382Z","iopub.status.idle":"2024-12-14T11:34:45.963757Z","shell.execute_reply.started":"2024-12-14T11:34:37.441341Z","shell.execute_reply":"2024-12-14T11:34:45.962185Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 2: Explore the Dataset\n1. Check Data Types: Understand the structure of features.\n2. Missing Values: Identify columns with missing values.\n3. Summary Statistics: Inspect distributions, outliers, and ranges.\n4. Target Variable: Analyze Premium Amount in the training data.","metadata":{}},{"cell_type":"code","source":"#1.Check Data Types: Understand the structure of features.\n# Check data types and summary\nprint(train_data.info())\nprint(train_data.describe())\n\n# Check missing values\n# Display missing values for train data\ntrain_missing = train_data.isnull().sum()\ntrain_missing = train_missing[train_missing > 0]\nprint(\"\\n\\033[1;34mTrain Data Missing Values:\\033[0m\")\nprint(train_missing)\n\n# Display missing values for test data\ntest_missing = test_data.isnull().sum()\ntest_missing = test_missing[test_missing > 0]\nprint(\"\\n\\033[1;34mTest Data Missing Values:\\033[0m\") \nprint(test_missing)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:55:45.157836Z","iopub.execute_input":"2024-12-14T16:55:45.15822Z","iopub.status.idle":"2024-12-14T16:55:47.562605Z","shell.execute_reply.started":"2024-12-14T16:55:45.158188Z","shell.execute_reply":"2024-12-14T16:55:47.561425Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 3: Inspect target variable","metadata":{}},{"cell_type":"code","source":"# Inspect target variable\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings\nwarnings.filterwarnings('ignore')\n\nplt.figure(figsize=(8, 6))\nsns.histplot(train_data['Premium Amount'], kde=True, color='blue')\nplt.title(\"Premium Amount Distribution\")\nplt.show()\nprint(\"Central tendency, spread, and shape of the Premium Amount as per the graph:\")\nprint(\"1. Central Tendency: The data is left-skewed, which indicates that the mean is likely to be lower than the median. This is because the left tail of the distribution pulls the mean down.\")\nprint(\"2. Spread: The spread of the data appears to be wide, with a significant number of data points clustered on the right side, and a long tail extending to the left.\")\nprint(\"3. Shape: The distribution is asymmetrical, with a peak on the right side and a long left tail. This confirms the left skewness, meaning that while most of the premium amounts are concentrated on the higher end, there are some significantly lower values that stretch the distribution to the left.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:56:25.6257Z","iopub.execute_input":"2024-12-14T16:56:25.626729Z","iopub.status.idle":"2024-12-14T16:56:31.548296Z","shell.execute_reply.started":"2024-12-14T16:56:25.626686Z","shell.execute_reply":"2024-12-14T16:56:31.547256Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 4:Visualize distributions of numerical features","metadata":{}},{"cell_type":"code","source":"#2. Distribution and Summary of Numerical Features\nnumerical_features = ['Age', 'Annual Income', 'Health Score', 'Credit Score', 'Insurance Duration', 'Vehicle Age']\n\nfor feature in numerical_features:\n    plt.figure(figsize=(8, 4))\n    sns.histplot(train_data[feature], kde=True, color='green', bins=30)\n    plt.title(f'Distribution of {feature}')\n    plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:56:55.321313Z","iopub.execute_input":"2024-12-14T16:56:55.321722Z","iopub.status.idle":"2024-12-14T16:57:26.890613Z","shell.execute_reply.started":"2024-12-14T16:56:55.321673Z","shell.execute_reply":"2024-12-14T16:57:26.889633Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 5:Visualize distributions of categorigal features along with Permium","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# List of categorical features\ncategorical_features = ['Gender', 'Marital Status', 'Education Level', 'Occupation', \n                        'Location', 'Policy Type', 'Smoking Status', 'Exercise Frequency', 'Property Type']\n\nfor feature in categorical_features:\n    # Aggregate Premium Amount by each category\n    category_premium_mean = train_data.groupby(feature)['Premium Amount'].mean().reset_index()\n    # Create a figure\n    plt.figure(figsize=(12, 6))\n    \n    # Bar plot for counts\n    plt.subplot(1, 2, 1)\n    sns.countplot(data=train_data, x=feature, order=train_data[feature].value_counts().index, palette='Set2')\n    plt.title(f'Distribution of {feature}', fontsize=14)\n    plt.xlabel(feature, fontsize=12)\n    plt.ylabel('Count', fontsize=12)\n    plt.xticks(rotation=45, fontsize=10)\n    \n    # Line chart for Premium Amount\n    plt.subplot(1, 2, 2)\n    plt.plot(category_premium_mean[feature], category_premium_mean['Premium Amount'], marker='o', color='b', linestyle='-')\n    plt.title(f'Premium Amount by {feature}', fontsize=14)\n    plt.xlabel(feature, fontsize=12)\n    plt.ylabel('Average Premium Amount', fontsize=12)\n    plt.xticks(rotation=45, fontsize=10)\n    \n    # Adjust layout\n    plt.tight_layout()\n    plt.show()\n    print(\"category_premium_mean \",category_premium_mean )\n    \n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:58:35.337664Z","iopub.execute_input":"2024-12-14T16:58:35.338059Z","iopub.status.idle":"2024-12-14T16:58:44.622717Z","shell.execute_reply.started":"2024-12-14T16:58:35.338023Z","shell.execute_reply":"2024-12-14T16:58:44.621548Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train_data=train_data.copy()\ndf_test_data=test_data.copy()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T11:38:16.858053Z","iopub.execute_input":"2024-12-14T11:38:16.858506Z","iopub.status.idle":"2024-12-14T11:38:17.061862Z","shell.execute_reply.started":"2024-12-14T11:38:16.858466Z","shell.execute_reply":"2024-12-14T11:38:17.060796Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":" ## Step 6:Data Preprocessing","metadata":{}},{"cell_type":"code","source":"#Preprocessing\n#Encode Categorical Variables: Convert columns like Gender, Marital Status, and Occupation into numerical values using one-hot or label encoding.\n#Handle Missing Values: Impute or drop missing data.\n#Date Handling: Parse Policy Start Date to extract year, month, etc.\n#Normalize/Standardize: Scale numerical features like Annual Income, Health Score, and Credit Score.\nfrom sklearn.preprocessing import LabelEncoder, StandardScaler\nfrom sklearn.model_selection import train_test_split\n\n# Handle categorical variables using LabelEncoder or get_dummies\ncat_columns = ['Gender', 'Marital Status', 'Education Level', 'Occupation', 'Location', 'Policy Type', \n               'Customer Feedback','Smoking Status', 'Exercise Frequency', 'Property Type']\nfor col in cat_columns:\n    le = LabelEncoder()\n    df_train_data[col] = le.fit_transform(df_train_data[col])\n    df_test_data[col] = le.transform(df_test_data[col])  # Use the same encoding on test data\n\n# Parse dates and extract useful features\ndf_train_data['Policy Start Date'] = pd.to_datetime(df_train_data['Policy Start Date'])\ndf_test_data['Policy Start Date'] = pd.to_datetime(df_test_data['Policy Start Date'])\ndf_train_data['Policy Start Month'] = df_train_data['Policy Start Date'].dt.month\ndf_test_data['Policy Start Month'] = df_test_data['Policy Start Date'].dt.month\n\n# Drop unused columns\ndf_train_data = df_train_data.drop(['id', 'Policy Start Date'], axis=1)\ndf_test_data = df_test_data.drop(['id', 'Policy Start Date'], axis=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:59:34.825816Z","iopub.execute_input":"2024-12-14T16:59:34.826186Z","iopub.status.idle":"2024-12-14T16:59:36.963617Z","shell.execute_reply.started":"2024-12-14T16:59:34.826155Z","shell.execute_reply":"2024-12-14T16:59:36.962147Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 7:Correlation Analysis","metadata":{}},{"cell_type":"code","source":"#4.Correlation Analysis\n# Correlation heatmap\nplt.figure(figsize=(12, 8))\ncorr = df_train_data.corr()\nsns.heatmap(corr, annot=True, fmt='.2f', cmap='coolwarm')\nplt.title(\"Correlation Matrix\")\nplt.show()\nprint(\"The correlation heatmap does not reveal any strong relationships between the variables.\")\nprint(\"Most correlation values are close to 0, indicating weak or no linear dependencies.\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T17:00:22.420781Z","iopub.execute_input":"2024-12-14T17:00:22.42115Z","iopub.status.idle":"2024-12-14T17:00:25.224006Z","shell.execute_reply.started":"2024-12-14T17:00:22.421116Z","shell.execute_reply":"2024-12-14T17:00:25.222922Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_cleaned=df_train_data.dropna()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T17:00:28.606571Z","iopub.execute_input":"2024-12-14T17:00:28.606966Z","iopub.status.idle":"2024-12-14T17:00:28.696001Z","shell.execute_reply.started":"2024-12-14T17:00:28.606931Z","shell.execute_reply":"2024-12-14T17:00:28.694857Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 7:Remove outliers","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\n# Assuming you have already loaded your data into 'df_cleaned'\n# df_cleaned = pd.read_csv('your_data.csv')\n\n# List of columns to check for outliers\ncolumns_to_check = [\n    'Age', 'Annual Income', 'Health Score', 'Number of Dependents', \n    'Previous Claims', 'Credit Score', 'Premium Amount'\n]\n\n# Function to remove outliers based on IQR\ndef remove_outliers_iqr(df, columns):\n    for column in columns:\n        # Calculate the IQR\n        Q1 = df[column].quantile(0.25)\n        Q3 = df[column].quantile(0.75)\n        IQR = Q3 - Q1\n        \n        # Define the bounds for outliers\n        lower_bound = Q1 - 1.5 * IQR\n        upper_bound = Q3 + 1.5 * IQR\n        \n        # Filter out outliers\n        df = df[(df[column] >= lower_bound) & (df[column] <= upper_bound)]\n        \n    return df\n\n# Remove outliers from the specified columns\ndf_cleaned_no_outliers = remove_outliers_iqr(df_cleaned, columns_to_check)\n\n# Check the shape of the new dataframe\nprint(f\"Original shape: {df_cleaned.shape}\")\nprint(f\"Shape after removing outliers: {df_cleaned_no_outliers.shape}\")\n\n# Optionally, you can inspect the rows that were removed (outliers)\n# outliers = df_cleaned[~df_cleaned.index.isin(df_cleaned_no_outliers.index)]\n# print(outliers)\n\n# Now you can proceed with further analysis or plotting on the cleaned data\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T17:00:56.425436Z","iopub.execute_input":"2024-12-14T17:00:56.42584Z","iopub.status.idle":"2024-12-14T17:00:56.854192Z","shell.execute_reply.started":"2024-12-14T17:00:56.425805Z","shell.execute_reply":"2024-12-14T17:00:56.853123Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 9:Feature Engineering ","metadata":{}},{"cell_type":"code","source":"# Example interaction terms\nprint('Income_Dependants: Combines Annual Income and Number of Dependents to reflect financial responsibility per household.')\n\ndf_cleaned_no_outliers['Income_Dependants'] = df_cleaned_no_outliers['Annual Income'] * df_cleaned_no_outliers['Number of Dependents']\nprint(\"Age_Vehicle_Age: Multiplies Age and Vehicle Age to capture the relationship between a person's age and their vehicle's age.\")\ndf_cleaned_no_outliers['Age_Vehicle_Age'] = df_cleaned_no_outliers['Age'] * df_cleaned_no_outliers['Vehicle Age']\nprint(\"Income_Age: Combines Annual Income and Age to represent income potential relative to age.\")\ndf_cleaned_no_outliers['Income_Age'] = df_cleaned_no_outliers['Annual Income'] * df_cleaned_no_outliers['Age']\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T17:02:39.541832Z","iopub.execute_input":"2024-12-14T17:02:39.542246Z","iopub.status.idle":"2024-12-14T17:02:39.558005Z","shell.execute_reply.started":"2024-12-14T17:02:39.542211Z","shell.execute_reply":"2024-12-14T17:02:39.556969Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 9:Correaltion Analysis after removing outlier & Feature Engineering ","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Calculate the correlation matrix\ncorrelation_matrix = df_cleaned_no_outliers.corr()\n\n# Check the correlation matrix shape to ensure it's correct\nprint(correlation_matrix.shape)\n# Increase plot size to fit all columns\nplt.figure(figsize=(20, 18))  # Adjust the size (width x height)\n\n# Plot the heatmap\nsns.heatmap(correlation_matrix, annot=True, cmap='coolwarm', fmt='.2f', vmin=-1, vmax=1)\n\n# Rotate x and y axis labels for better visibility\nplt.xticks(rotation=90)  # Rotate column labels\nplt.yticks(rotation=0)   # Rotate row labels to make them readable\n\n# Show the plot\nplt.title(\"Correlation Matrix\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T17:03:17.857078Z","iopub.execute_input":"2024-12-14T17:03:17.857469Z","iopub.status.idle":"2024-12-14T17:03:20.779413Z","shell.execute_reply.started":"2024-12-14T17:03:17.857418Z","shell.execute_reply":"2024-12-14T17:03:20.778148Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 10:Predicting insurance premiums using LightGBM with RMSE evaluation metric.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import mean_squared_error\nfrom lightgbm import LGBMRegressor\n\n\n# Step 3: Split the data into features (X) and target (y)\nX = df_cleaned_no_outliers.drop(columns=['Premium Amount'])  # Drop target and unwanted features\ny = df_cleaned_no_outliers['Premium Amount']\n\n# Step 4: Train-test split\nX_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42)\n\n# Step 5: Initialize the LightGBM Regressor\nlgb_model = LGBMRegressor(\n    objective='regression',\n    metric='rmse',\n    boosting_type='gbdt',\n    num_leaves=31,\n    learning_rate=0.05,\n    feature_fraction=0.9,\n    n_estimators=1000,  # Set the max number of trees\n    early_stopping_rounds=50,  # Enable early stopping\n    verbose=100  # Set verbosity to print the progress every 100 iterations\n)\n\n# Step 6: Train the model\nlgb_model.fit(X_train, y_train, \n              eval_set=[(X_val, y_val)], \n              eval_metric='rmse')\n\n## Step 7: Predict on the validation set\ny_pred = lgb_model.predict(X_val)\n\n# Step 8: Evaluate the model using RMSE\nrmse = mean_squared_error(y_val, y_pred, squared=False)\nprint(f'Root Mean Squared Error (RMSE): {rmse}')\nprint(f\"this case, an RMSE of approximately {rmse} suggests that, on average, the model's predictions are off by around {rmse} units from the actual premium amount.\")\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T17:05:11.754338Z","iopub.execute_input":"2024-12-14T17:05:11.754745Z","iopub.status.idle":"2024-12-14T17:05:24.042863Z","shell.execute_reply.started":"2024-12-14T17:05:11.75471Z","shell.execute_reply":"2024-12-14T17:05:24.041782Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 11:Visualizing model performance: Actual vs Predicted Premium and the distribution of residuals.","metadata":{}},{"cell_type":"code","source":"# Check predicted vs actual values\nimport matplotlib.pyplot as plt\n\nplt.scatter(y_val, y_pred, alpha=0.5)\nplt.xlabel('Actual Premium Amount')\nplt.ylabel('Predicted Premium Amount')\nplt.title('Actual vs Predicted Premium')\nplt.show()\n\n# Calculate the residuals\nresiduals = y_val - y_pred\nplt.figure(figsize=(10, 6))\nplt.hist(residuals, bins=50, color='blue', alpha=0.7)\nplt.title('Residuals Distribution')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T17:05:58.406201Z","iopub.execute_input":"2024-12-14T17:05:58.406593Z","iopub.status.idle":"2024-12-14T17:05:59.291178Z","shell.execute_reply.started":"2024-12-14T17:05:58.406556Z","shell.execute_reply":"2024-12-14T17:05:59.290022Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 12:Splitting test data into chunks for efficient prediction and generating submission file.\n","metadata":{}},{"cell_type":"code","source":"chunk_size = 10000  # Adjust based on memory availability\n\n# Prepare the submission DataFrame\nsubmission = pd.DataFrame()\n\nfor i in range(0, len(test_data), chunk_size):\n    chunk = test_data.iloc[i:i+chunk_size]\n    chunk_encoded = pd.get_dummies(chunk, drop_first=True)\n    chunk_encoded = chunk_encoded.reindex(columns=X_train.columns, fill_value=0)\n    chunk_pred = lgb_model.predict(chunk_encoded)\n    \n    # Store predictions with IDs\n    submission = pd.concat([submission, pd.DataFrame({\n        'id': chunk['id'],  # Ensure 'id' column exists in test_data\n        'premium': chunk_pred\n    })], ignore_index=True)\n\nsubmission.to_csv('submission.csv', index=False)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T17:06:31.293766Z","iopub.execute_input":"2024-12-14T17:06:31.294533Z","iopub.status.idle":"2024-12-14T17:06:51.236236Z","shell.execute_reply.started":"2024-12-14T17:06:31.294496Z","shell.execute_reply":"2024-12-14T17:06:51.235175Z"}},"outputs":[],"execution_count":null}]}