{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"},{"sourceId":9178166,"sourceType":"datasetVersion","datasetId":5547076}],"dockerImageVersionId":30786,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false},"papermill":{"default_parameters":{},"duration":4733.8758,"end_time":"2024-10-12T17:15:48.677904","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-10-12T15:56:54.802104","version":"2.6.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"15cc171e","cell_type":"markdown","source":"<p style=\"font-family: 'Brush Script MT', cursive; font-size: 80px; text-align: center; color: #4682B4;\">Regression Analysis on Insurance Dataset</p>\n","metadata":{"papermill":{"duration":0.011758,"end_time":"2024-10-12T15:56:57.608201","exception":false,"start_time":"2024-10-12T15:56:57.596443","status":"completed"},"tags":[]}},{"id":"a631c2e9","cell_type":"markdown","source":"  [![GitHub](https://img.shields.io/badge/GitHub-Profile-blue?style=for-the-badge&logo=github)](https://github.com/iammuhammadfurqan)\n\n  [![Kaggle](https://img.shields.io/badge/Kaggle-Profile-blue?style=for-the-badge&logo=kaggle)](https://www.kaggle.com/muhammadfurqan0)\n\n  [![LinkedIn](https://img.shields.io/badge/LinkedIn-Profile-blue?style=for-the-badge&logo=linkedin)](https://www.linkedin.com/in/iammuhammadfurqan/)\n\n  [![Gmail](https://img.shields.io/badge/Gmail-Contact%20Me-red?style=for-the-badge&logo=gmail)](mailto:sheikhfurqan048@gmail.com)","metadata":{"papermill":{"duration":0.010551,"end_time":"2024-10-12T15:56:57.629678","exception":false,"start_time":"2024-10-12T15:56:57.619127","status":"completed"},"tags":[]}},{"id":"a36072c4","cell_type":"markdown","source":"**Data Description: Insurance Policy and Customer Dataset**  \r\n---\r\n\r\nThis dataset provides information about insurance policies and customer demographics, including personal, financial, and lifestyle details. It is suitable for analyzing factors that influence insurance premiums, customer behavior, and policy trends. The dataset comprises the following columns:\r\n\r\n- **id**: Unique identifier for each customer or policy.  \r\n- **Age**: The age of the customer in years.  \r\n- **Gender**: The gender of the customer (e.g., Male, Female).  \r\n- **Annual Income**: The annual income of the customer in currency units.  \r\n- **Marital Status**: The marital status of the customer (e.g., Single, Married, Divorced).  \r\n- **Number of Dependents**: The number of dependents (e.g., children, dependents) supported by the customer.  \r\n- **Education Level**: The highest level of education attained by the customer (e.g., High School, Bachelor’s, Master’s).  \r\n- **Occupation**: The customer's occupation or professional role.  \r\n- **Health Score**: A numeric representation of the customer’s overall health condition (e.g., on a scale of 1-10).  \r\n- **Location**: The geographic location or residence of the customer.  \r\n- **Policy Type**: The type of insurance policy purchased (e.g., Comprehensive, Third-party).  \r\n- **Previous Claims**: The number of insurance claims made by the customer in the past.  \r\n- **Vehicle Age**: The age of the customer’s insured vehicle in years.  \r\n- **Credit Score**: The customer’s credit score, reflecting their financial reliability.  \r\n- **Insurance Duration**: The length of time the policy is active, in months or years.  \r\n- **Policy Start Date**: The starting date of the insurance policy.  \r\n- **Customer Feedback**: Feedback or reviews provided by the customer about the service.  \r\n- **Smoking Status**: Indicates whether the customer is a smoker (Yes/No).  \r\n- **Exercise Frequency**: The frequency with which the customer engages in physical exercise (e.g., Daily, Weekly).  \r\n- **Property Type**: The type of property owned by the customer (e.g., Apartment, House, Commercial).  \r\n- **Premium Amount**: The premium amount paid by the customer for the insurance policy.  \r\n\r\nThis dataset is a valuable resource for insurers, analysts, and researchers interested in understanding customer profiles, predicting insurance risks, and optimizing policy pricing.","metadata":{"papermill":{"duration":0.010441,"end_time":"2024-10-12T15:56:57.651112","exception":false,"start_time":"2024-10-12T15:56:57.640671","status":"completed"},"tags":[]}},{"id":"9cec0dd6","cell_type":"markdown","source":"<p style=\"font-family: 'Brush Script MT', cursive; font-size: 80px; text-align: center; color: #4682B4;\">Importing the Libraries</p>\n","metadata":{"papermill":{"duration":0.010435,"end_time":"2024-10-12T15:56:57.672331","exception":false,"start_time":"2024-10-12T15:56:57.661896","status":"completed"},"tags":[]}},{"id":"0da4751d","cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport plotly.graph_objects as go\nimport plotly.express as px\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import StandardScaler, OneHotEncoder\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.preprocessing import FunctionTransformer\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.metrics import mean_squared_error, r2_score\nfrom sklearn.experimental import enable_iterative_imputer  # Enable the iterative imputer\nfrom sklearn.impute import IterativeImputer","metadata":{"papermill":{"duration":3.152487,"end_time":"2024-10-12T15:57:00.835468","exception":false,"start_time":"2024-10-12T15:56:57.682981","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:17:53.106335Z","iopub.execute_input":"2024-12-24T08:17:53.106724Z","iopub.status.idle":"2024-12-24T08:17:53.114237Z","shell.execute_reply.started":"2024-12-24T08:17:53.10669Z","shell.execute_reply":"2024-12-24T08:17:53.112954Z"}},"outputs":[],"execution_count":null},{"id":"e37f904f","cell_type":"code","source":"def load_dataset(file_path, **kwargs):\n    \"\"\"\n    Load a dataset from the specified file path.\n    Args:\n        file_path (str): Path to the dataset.\n        **kwargs: Additional arguments for pd.read_csv (e.g., delimiter, encoding).\n    Returns:\n        pd.DataFrame: Loaded dataset.\n    \"\"\"\n    try:\n        df = pd.read_csv(file_path, **kwargs)\n        print(f\"Dataset loaded successfully from {file_path}\")\n        return df\n    except Exception as e:\n        print(f\"Error loading dataset: {e}\")\n        return None\n\n# Example usage\ntrain_path = \"/kaggle/input/playground-series-s4e12/train.csv\"\ntest_path = \"/kaggle/input/playground-series-s4e12/test.csv\"\n\ndf = load_dataset(train_path)\ndf_test = load_dataset(test_path)\n","metadata":{"papermill":{"duration":1.350233,"end_time":"2024-10-12T15:57:02.197686","exception":false,"start_time":"2024-10-12T15:57:00.847453","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:17:53.116456Z","iopub.execute_input":"2024-12-24T08:17:53.11683Z","iopub.status.idle":"2024-12-24T08:18:00.505737Z","shell.execute_reply.started":"2024-12-24T08:17:53.116795Z","shell.execute_reply":"2024-12-24T08:18:00.504457Z"}},"outputs":[],"execution_count":null},{"id":"e2d481ab","cell_type":"markdown","source":"<p style=\"font-family: 'Brush Script MT', cursive; font-size: 80px; text-align: center; color: #4682B4;\">Sneak Preview of Data</p>\n","metadata":{"papermill":{"duration":0.010884,"end_time":"2024-10-12T15:57:02.220166","exception":false,"start_time":"2024-10-12T15:57:02.209282","status":"completed"},"tags":[]}},{"id":"faeef065","cell_type":"code","source":"df.head()","metadata":{"papermill":{"duration":0.043767,"end_time":"2024-10-12T15:57:02.275082","exception":false,"start_time":"2024-10-12T15:57:02.231315","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:00.507534Z","iopub.execute_input":"2024-12-24T08:18:00.5079Z","iopub.status.idle":"2024-12-24T08:18:00.531888Z","shell.execute_reply.started":"2024-12-24T08:18:00.507868Z","shell.execute_reply":"2024-12-24T08:18:00.530752Z"}},"outputs":[],"execution_count":null},{"id":"942b2f48","cell_type":"code","source":"df.tail()","metadata":{"papermill":{"duration":0.030096,"end_time":"2024-10-12T15:57:02.316615","exception":false,"start_time":"2024-10-12T15:57:02.286519","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:00.533285Z","iopub.execute_input":"2024-12-24T08:18:00.53364Z","iopub.status.idle":"2024-12-24T08:18:00.562318Z","shell.execute_reply.started":"2024-12-24T08:18:00.533608Z","shell.execute_reply":"2024-12-24T08:18:00.56128Z"}},"outputs":[],"execution_count":null},{"id":"f8b0e793","cell_type":"markdown","source":"<p style=\"font-family: 'Brush Script MT', cursive; font-size: 80px; text-align: center; color: #4682B4;\">Little Bit EDA</p>\n","metadata":{"papermill":{"duration":0.01148,"end_time":"2024-10-12T15:57:02.33991","exception":false,"start_time":"2024-10-12T15:57:02.32843","status":"completed"},"tags":[]}},{"id":"e004e447","cell_type":"code","source":"#Check the shape of data\nprint(f'The Training Dataset has {df.shape[0]} rows and {df.shape[1]} columns.')\nprint(f'The Training Dataset has {df_test.shape[0]} rows and {df_test.shape[1]} columns.')","metadata":{"papermill":{"duration":0.021197,"end_time":"2024-10-12T15:57:02.372624","exception":false,"start_time":"2024-10-12T15:57:02.351427","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:00.564444Z","iopub.execute_input":"2024-12-24T08:18:00.564765Z","iopub.status.idle":"2024-12-24T08:18:00.576662Z","shell.execute_reply.started":"2024-12-24T08:18:00.564731Z","shell.execute_reply":"2024-12-24T08:18:00.575631Z"}},"outputs":[],"execution_count":null},{"id":"b794e4d5","cell_type":"code","source":"df.columns","metadata":{"papermill":{"duration":0.021641,"end_time":"2024-10-12T15:57:02.405642","exception":false,"start_time":"2024-10-12T15:57:02.384001","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:00.578017Z","iopub.execute_input":"2024-12-24T08:18:00.578401Z","iopub.status.idle":"2024-12-24T08:18:00.593368Z","shell.execute_reply.started":"2024-12-24T08:18:00.578339Z","shell.execute_reply":"2024-12-24T08:18:00.592154Z"}},"outputs":[],"execution_count":null},{"id":"c878eddf","cell_type":"code","source":"df.info()","metadata":{"papermill":{"duration":0.117864,"end_time":"2024-10-12T15:57:02.535131","exception":false,"start_time":"2024-10-12T15:57:02.417267","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:00.595108Z","iopub.execute_input":"2024-12-24T08:18:00.595695Z","iopub.status.idle":"2024-12-24T08:18:01.241566Z","shell.execute_reply.started":"2024-12-24T08:18:00.595646Z","shell.execute_reply":"2024-12-24T08:18:01.240456Z"}},"outputs":[],"execution_count":null},{"id":"8d4ff734-5968-42c1-982e-af1f98a19bac","cell_type":"code","source":"# Assuming your data is in a DataFrame named `df`\ndf['Policy Start Date'] = pd.to_datetime(df['Policy Start Date'], errors='coerce')\n\n# Verifying the conversion\nprint(df['Policy Start Date'].head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:01.242917Z","iopub.execute_input":"2024-12-24T08:18:01.243218Z","iopub.status.idle":"2024-12-24T08:18:01.652734Z","shell.execute_reply.started":"2024-12-24T08:18:01.243189Z","shell.execute_reply":"2024-12-24T08:18:01.651709Z"}},"outputs":[],"execution_count":null},{"id":"e6b0284e","cell_type":"code","source":"df.describe().T","metadata":{"papermill":{"duration":0.061415,"end_time":"2024-10-12T15:57:02.608392","exception":false,"start_time":"2024-10-12T15:57:02.546977","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:01.654255Z","iopub.execute_input":"2024-12-24T08:18:01.655443Z","iopub.status.idle":"2024-12-24T08:18:02.41665Z","shell.execute_reply.started":"2024-12-24T08:18:01.655393Z","shell.execute_reply":"2024-12-24T08:18:02.415443Z"}},"outputs":[],"execution_count":null},{"id":"0859fdd9-269c-407e-b760-e232e13b9de3","cell_type":"markdown","source":"### Observations:\n\n1. **ID**: Unique identifier for each record, ranges from 0 to 1,199,999.\n2. **Age**: Median age is 41 years, with values ranging from 18 to 64. Indicates a working-age population.\n3. **Annual Income**: Highly varied, with a median of $23,911, and outliers up to $149,997.\n4. **Number of Dependents**: Most have 1–3 dependents, with a median of 2.\n5. **Health Score**: Median is 24.58, skewed positively with a high max of 58.98.\n6. **Previous Claims**: Majority have 0–1 claims, with up to 9 claims observed.\n7. **Vehicle Age**: Median vehicle age is 10 years, with a range of 0–19 years.\n8. **Credit Score**: Median score is 595, with a wide range indicating varied financial stability.\n9. **Insurance Duration**: Policies range from 1–9 years, with a median duration of 5 years.\n10. **Policy Start Date**: Spans 2019–2024, indicating policy trends over years.\n11. **Premium Amount**: Median premium is $872, with significant variation (range: $20–$4,999).","metadata":{}},{"id":"6d96a140","cell_type":"code","source":"# Summary statistics for categorical variables\ndf.describe(include='object')","metadata":{"papermill":{"duration":0.252907,"end_time":"2024-10-12T15:57:02.873332","exception":false,"start_time":"2024-10-12T15:57:02.620425","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:02.417982Z","iopub.execute_input":"2024-12-24T08:18:02.418477Z","iopub.status.idle":"2024-12-24T08:18:04.33816Z","shell.execute_reply.started":"2024-12-24T08:18:02.418416Z","shell.execute_reply":"2024-12-24T08:18:04.337063Z"}},"outputs":[],"execution_count":null},{"id":"980a49f8-8a95-4da7-93ad-b0bda4c36558","cell_type":"markdown","source":"### Observations:\n\n1. **Gender**: Equal representation with males being slightly dominant (top: Male, freq: 602,571).  \n2. **Marital Status**: Majority are Single (395,391), followed by Married and Divorced.  \n3. **Education Level**: Diverse distribution, Master's degree holders dominate (303,818).  \n4. **Occupation**: Most are Employed (282,750), followed by Unemployed and Self-employed.  \n5. **Location**: Suburban residents are the majority (401,542).  \n6. **Policy Type**: Premium policies are the most common (401,846).  \n7. **Customer Feedback**: Average feedback is most frequent (377,905 responses).  \n8. **Smoking Status**: Majority are smokers (601,873).  \n9. **Exercise Frequency**: Weekly exercise is the top frequency (306,179).  \n10. **Property Type**: Most participants own a house (400,349).","metadata":{}},{"id":"3267f518-6568-4b8b-8328-94387526fd09","cell_type":"code","source":"def calculate_missing_percentage(data):\n    \"\"\"\n    Calculate the percentage of missing values for each column in the dataset.\n    \"\"\"\n    missing_percentage = (data.isnull().sum() / len(data)) * 100\n    return missing_percentage.round(2)\n\ndef plot_missing_percentage(missing_percentage, title=\"Percentage of Missing Values by Column\"):\n    \"\"\"\n    Plot the percentage of missing values for each column as a horizontal bar chart.\n    \"\"\"\n    missing_percentage.plot(kind='barh', figsize=(10, 6), color='skyblue', edgecolor=\"black\", linewidth=1.0)\n    plt.title(title)\n    plt.xlabel('Percentage')\n    plt.ylabel('Columns')\n    plt.show()\n\ndef plot_missing_heatmap(data, title=\"Heatmap of Missing Values\"):\n    \"\"\"\n    Create a heatmap visualization of missing values in the dataset.\n    \"\"\"\n    plt.figure(figsize=(12, 8))\n    sns.heatmap(data.isnull(), cbar=False, cmap='viridis')\n    plt.title(title)\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:04.341369Z","iopub.execute_input":"2024-12-24T08:18:04.341713Z","iopub.status.idle":"2024-12-24T08:18:04.349086Z","shell.execute_reply.started":"2024-12-24T08:18:04.341681Z","shell.execute_reply":"2024-12-24T08:18:04.34765Z"}},"outputs":[],"execution_count":null},{"id":"05b407bb-a271-4055-95d5-23521542f09e","cell_type":"code","source":"# Example usage:\nprint(\"Missing Percentage in Training Data\")\nprint(\"------------------------------------\")\ntrain_missing_percentage = calculate_missing_percentage(df)\nprint(train_missing_percentage)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:04.350943Z","iopub.execute_input":"2024-12-24T08:18:04.3513Z","iopub.status.idle":"2024-12-24T08:18:04.947217Z","shell.execute_reply.started":"2024-12-24T08:18:04.351241Z","shell.execute_reply":"2024-12-24T08:18:04.946188Z"}},"outputs":[],"execution_count":null},{"id":"f3905a3e-0b2b-4782-8593-ad61ada072e5","cell_type":"code","source":"print(\"\\nMissing Percentage in Testing Data\")\nprint(\"------------------------------------\")\ntest_missing_percentage = calculate_missing_percentage(df_test)\nprint(test_missing_percentage)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:04.948559Z","iopub.execute_input":"2024-12-24T08:18:04.948867Z","iopub.status.idle":"2024-12-24T08:18:05.377728Z","shell.execute_reply.started":"2024-12-24T08:18:04.948838Z","shell.execute_reply":"2024-12-24T08:18:05.376678Z"}},"outputs":[],"execution_count":null},{"id":"e6f99679-0961-4ad2-b732-71a2d4e503da","cell_type":"code","source":"# Plot for training data\nplot_missing_percentage(train_missing_percentage, title=\"Percentage of Missing Values in Training Data\")\nplot_missing_heatmap(df, title=\"Heatmap of Missing Values in Training Data\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:05.379182Z","iopub.execute_input":"2024-12-24T08:18:05.379611Z","iopub.status.idle":"2024-12-24T08:18:30.347024Z","shell.execute_reply.started":"2024-12-24T08:18:05.379567Z","shell.execute_reply":"2024-12-24T08:18:30.345964Z"}},"outputs":[],"execution_count":null},{"id":"eaf1cb0d-c1cc-4192-a56e-17c4f94345be","cell_type":"code","source":"# Plot for test data\nplot_missing_percentage(test_missing_percentage, title=\"Percentage of Missing Values in Test Data\")\nplot_missing_heatmap(df_test, title=\"Heatmap of Missing Values in Test Data\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:30.348397Z","iopub.execute_input":"2024-12-24T08:18:30.348708Z","iopub.status.idle":"2024-12-24T08:18:46.378909Z","shell.execute_reply.started":"2024-12-24T08:18:30.348678Z","shell.execute_reply":"2024-12-24T08:18:46.37767Z"}},"outputs":[],"execution_count":null},{"id":"49c65487-5c1d-4fed-8b96-c2c71d3841b7","cell_type":"markdown","source":"### Observations on Missing Data:\n\n#### Training Data:\n- **High Missing Values**:  \n  - Occupation (29.84%)  \n  - Previous Claims (30.34%)  \n  - Credit Score (11.49%)  \n  - Number of Dependents (9.14%)  \n\n- **Moderate Missing Values**:  \n  - Health Score (6.17%)  \n  - Customer Feedback (6.49%)  \n\n- **Low Missing Values**:  \n  - Age (1.56%)  \n  - Marital Status (1.54%)  \n  - Annual Income (3.75%)  \n\n#### Testing Data:\n- Missing percentages are similar to the training data.  \n  - Highest: Occupation (29.89%), Previous Claims (30.35%).  \n  - Other variables follow the same pattern.  \n","metadata":{}},{"id":"f2e77671","cell_type":"code","source":"# Check for duplicate rows\nduplicates = df.duplicated()\n\n# Print the number of duplicate rows\nprint(f\"Number of duplicate rows: {duplicates.sum()}\")\n\ntest_duplicates = df_test.duplicated()\n\nprint(f\"Number of duplicate rows: {test_duplicates.sum()}\")\n","metadata":{"papermill":{"duration":0.292404,"end_time":"2024-10-12T15:57:09.023453","exception":false,"start_time":"2024-10-12T15:57:08.731049","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:46.380663Z","iopub.execute_input":"2024-12-24T08:18:46.381084Z","iopub.status.idle":"2024-12-24T08:18:48.866819Z","shell.execute_reply.started":"2024-12-24T08:18:46.38104Z","shell.execute_reply":"2024-12-24T08:18:48.865516Z"}},"outputs":[],"execution_count":null},{"id":"62ca1e4d-bbb8-4980-a5d6-91bbb8ba09c3","cell_type":"code","source":"df.columns","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:48.868614Z","iopub.execute_input":"2024-12-24T08:18:48.868992Z","iopub.status.idle":"2024-12-24T08:18:48.875646Z","shell.execute_reply.started":"2024-12-24T08:18:48.868958Z","shell.execute_reply":"2024-12-24T08:18:48.874532Z"}},"outputs":[],"execution_count":null},{"id":"3e0fc5dc-1c5c-49e2-96d4-045d53aba09f","cell_type":"code","source":"import pandas as pd\n\ndef process_features(df, id_column, date_column, target_column=None):\n    \"\"\"\n    Processes numerical and categorical features in the dataset and extracts additional features from a date column.\n    \n    Parameters:\n    - df (pd.DataFrame): The input dataset.\n    - id_column (str): The name of the unique identifier to exclude from features.\n    - date_column (str): The name of the date column to process.\n    - target_column (str, optional): The name of the target variable to exclude from features. Default is None.\n    \n    Returns:\n    - numerical_features (list): List of numerical feature names.\n    - categorical_features (list): List of categorical feature names.\n    \"\"\"\n    # Remove target and unique identifier from numerical features\n    numerical_features = df.select_dtypes(include=['int64', 'float64']).columns.tolist()\n    if target_column and target_column in numerical_features:\n        numerical_features.remove(target_column)\n    if id_column in numerical_features:\n        numerical_features.remove(id_column)\n\n    # Handle date feature separately\n    if date_column in df.columns:\n        df[date_column] = pd.to_datetime(df[date_column])\n        df['Policy Year'] = df[date_column].dt.year\n        df['Policy Month'] = df[date_column].dt.month\n        df['Policy Day'] = df[date_column].dt.day\n\n        # Add new date-derived features to numerical features\n        numerical_features += ['Policy Year', 'Policy Month', 'Policy Day']\n\n    # Define categorical features\n    categorical_features = df.select_dtypes(include=['object']).columns.tolist()\n\n    return numerical_features, categorical_features\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:48.877503Z","iopub.execute_input":"2024-12-24T08:18:48.877968Z","iopub.status.idle":"2024-12-24T08:18:48.889409Z","shell.execute_reply.started":"2024-12-24T08:18:48.877895Z","shell.execute_reply":"2024-12-24T08:18:48.888287Z"}},"outputs":[],"execution_count":null},{"id":"748f1ca7-e0f9-4ecd-9cfb-daa6bf428277","cell_type":"code","source":"numerical_features, categorical_features = process_features(\n    df=df, \n    target_column='Premium Amount', \n    id_column='id', \n    date_column='Policy Start Date'\n)\n\nprint(\"Numerical Features:\", numerical_features)\nprint(\"Categorical Features:\", categorical_features)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:48.891053Z","iopub.execute_input":"2024-12-24T08:18:48.891783Z","iopub.status.idle":"2024-12-24T08:18:49.626791Z","shell.execute_reply.started":"2024-12-24T08:18:48.891735Z","shell.execute_reply":"2024-12-24T08:18:49.625594Z"}},"outputs":[],"execution_count":null},{"id":"c1459a91-c64a-4276-9d25-141f681a80d2","cell_type":"code","source":"numerical_features, categorical_features = process_features(\n    df=df_test, \n    id_column='id', \n    date_column='Policy Start Date'\n)\n\nprint(\"Numerical Features:\", numerical_features)\nprint(\"Categorical Features:\", categorical_features)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:49.628213Z","iopub.execute_input":"2024-12-24T08:18:49.628571Z","iopub.status.idle":"2024-12-24T08:18:50.357136Z","shell.execute_reply.started":"2024-12-24T08:18:49.628539Z","shell.execute_reply":"2024-12-24T08:18:50.355945Z"}},"outputs":[],"execution_count":null},{"id":"a0460273","cell_type":"code","source":"import plotly.graph_objects as go\n\ndef plot_interactive_pie_chart(df, column):\n    \"\"\"\n    Plot an interactive pie chart for the distribution of a categorical column using Plotly.\n    Args:\n        df (pd.DataFrame): DataFrame containing the data.\n        column (str): The column name to plot.\n    \"\"\"\n    counts = df[column].value_counts()\n    labels = counts.index\n    sizes = counts.values\n\n    # Create the pie chart\n    fig = go.Figure(\n        data=[\n            go.Pie(\n                labels=labels,\n                values=sizes,\n                hole=0.4,  # Creates a donut chart\n                textinfo=\"percent+label\",  # Show percentage and labels\n                marker=dict(\n                    line=dict(color=\"black\", width=2),  # Add a border\n                    colors=px.colors.qualitative.Set3[:len(labels)],  # Attractive colors\n                ),\n                pull=[0.1 if i == sizes.argmax() else 0 for i in range(len(sizes))],  # Highlight the largest slice\n            )\n        ]\n    )\n\n    # Update layout with animations and title\n    fig.update_layout(\n        title={\n            \"text\": f\"Distribution of {column}\",\n            \"y\": 0.9,\n            \"x\": 0.5,\n            \"xanchor\": \"center\",\n            \"yanchor\": \"top\",\n        },\n        template=\"presentation\",\n    )\n\n    # Show the chart\n    fig.show()\n\n# Example usage for all categorical features\nfor column in categorical_features:\n    plot_interactive_pie_chart(df, column)\n","metadata":{"papermill":{"duration":0.041562,"end_time":"2024-10-12T15:57:09.080969","exception":false,"start_time":"2024-10-12T15:57:09.039407","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:50.358892Z","iopub.execute_input":"2024-12-24T08:18:50.359239Z","iopub.status.idle":"2024-12-24T08:18:52.194353Z","shell.execute_reply.started":"2024-12-24T08:18:50.359205Z","shell.execute_reply":"2024-12-24T08:18:52.193251Z"}},"outputs":[],"execution_count":null},{"id":"aed9a97c","cell_type":"code","source":"\n# Plot histograms for numeric columns\ndf[numerical_features].hist(figsize=(16, 12), bins=20, color='skyblue', edgecolor='black')\nplt.suptitle(\"Histograms of Numeric Columns\", fontsize=16)\nplt.tight_layout(rect=[0, 0.03, 1, 0.95])  # Adjust layout for the title\nplt.show()\n","metadata":{"papermill":{"duration":0.790316,"end_time":"2024-10-12T15:57:09.887518","exception":false,"start_time":"2024-10-12T15:57:09.097202","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:52.196175Z","iopub.execute_input":"2024-12-24T08:18:52.196544Z","iopub.status.idle":"2024-12-24T08:18:55.275633Z","shell.execute_reply.started":"2024-12-24T08:18:52.196502Z","shell.execute_reply":"2024-12-24T08:18:55.27449Z"}},"outputs":[],"execution_count":null},{"id":"ab0378a0","cell_type":"code","source":"# Subplots for each numerical feature\nnum_cols = len(numerical_features)\nfig, axes = plt.subplots(nrows=(num_cols // 3 + 1), ncols=3, figsize=(16, 4 * (num_cols // 3 + 1)))\n\n# Flatten axes for easy indexing\naxes = axes.flatten()\n\nfor i, col in enumerate(numerical_features):\n    sns.boxplot(data=df[col], ax=axes[i], palette=\"Set3\")\n    axes[i].set_title(f'Box Plot of {col}')\n    axes[i].tick_params(axis='x', rotation=45)\n\n# Hide unused subplots\nfor j in range(i + 1, len(axes)):\n    fig.delaxes(axes[j])\n\nplt.tight_layout()\nplt.show()\n","metadata":{"papermill":{"duration":0.288148,"end_time":"2024-10-12T15:57:10.193088","exception":false,"start_time":"2024-10-12T15:57:09.90494","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:55.277366Z","iopub.execute_input":"2024-12-24T08:18:55.277704Z","iopub.status.idle":"2024-12-24T08:18:57.779107Z","shell.execute_reply.started":"2024-12-24T08:18:55.27767Z","shell.execute_reply":"2024-12-24T08:18:57.77796Z"}},"outputs":[],"execution_count":null},{"id":"a356a6fa-d6b9-420f-9051-3a56bbbbbd0f","cell_type":"code","source":"# Correlation heatmap for numerical features\nplt.figure(figsize=(10, 8))\ncorrelation_matrix = df[numerical_features].corr()\nsns.heatmap(correlation_matrix, annot=True, cmap='coolwarm', fmt='.2f')\nplt.title(\"Correlation Heatmap\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-24T08:18:57.780553Z","iopub.execute_input":"2024-12-24T08:18:57.780984Z","iopub.status.idle":"2024-12-24T08:18:58.922374Z","shell.execute_reply.started":"2024-12-24T08:18:57.78094Z","shell.execute_reply":"2024-12-24T08:18:58.921325Z"}},"outputs":[],"execution_count":null},{"id":"f7d0ab84-236d-4ccf-8e7a-d43a4f4c0c7a","cell_type":"markdown","source":"### Observations from Correlation Heatmap:\n\n1. **Weak Correlations**:\n   - Most features show minimal correlation with each other, with values close to 0, indicating independence.\n   - Example: Variables like `Age`, `Number of Dependents`, and `Vehicle Age` exhibit negligible correlation with others.\n\n2. **Negative Correlations**:\n   - `Annual Income` and `Credit Score` have a moderate negative correlation (-0.20), suggesting higher income might be associated with lower credit scores.\n\n3. **Policy Year and Month**:\n   - A negative correlation (-0.26) is observed between `Policy Year` and `Policy Month`, which may hint at temporal or seasonal trends.\n\n4. **Independent Variables**:\n   - Features like `Health Score`, `Previous Claims`, and `Insurance Duration` seem largely uncorrelated with others.\n\nThis suggests a dataset with diverse, independent features useful for predictive modeling.","metadata":{}}]}