{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30839,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:46:54.318304Z","iopub.execute_input":"2025-01-22T06:46:54.318796Z","iopub.status.idle":"2025-01-22T06:46:54.33936Z","shell.execute_reply.started":"2025-01-22T06:46:54.318761Z","shell.execute_reply":"2025-01-22T06:46:54.338114Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sns                       \nimport matplotlib.pyplot as plt             \n%matplotlib inline\nsns.set(color_codes=True)\nfrom sklearn.preprocessing import LabelEncoder, OneHotEncoder","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:46:54.341085Z","iopub.execute_input":"2025-01-22T06:46:54.341498Z","iopub.status.idle":"2025-01-22T06:46:54.350863Z","shell.execute_reply.started":"2025-01-22T06:46:54.341455Z","shell.execute_reply":"2025-01-22T06:46:54.349855Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\ndf_test = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:46:54.372638Z","iopub.execute_input":"2025-01-22T06:46:54.372997Z","iopub.status.idle":"2025-01-22T06:47:01.980659Z","shell.execute_reply.started":"2025-01-22T06:46:54.372968Z","shell.execute_reply":"2025-01-22T06:47:01.979495Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **Let us first see the train data set.**","metadata":{}},{"cell_type":"code","source":"df_train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:01.982101Z","iopub.execute_input":"2025-01-22T06:47:01.982454Z","iopub.status.idle":"2025-01-22T06:47:02.011618Z","shell.execute_reply.started":"2025-01-22T06:47:01.98241Z","shell.execute_reply":"2025-01-22T06:47:02.010344Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train.describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:02.014116Z","iopub.execute_input":"2025-01-22T06:47:02.014426Z","iopub.status.idle":"2025-01-22T06:47:02.680302Z","shell.execute_reply.started":"2025-01-22T06:47:02.014399Z","shell.execute_reply":"2025-01-22T06:47:02.679421Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train.dtypes","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:02.681951Z","iopub.execute_input":"2025-01-22T06:47:02.682252Z","iopub.status.idle":"2025-01-22T06:47:02.690196Z","shell.execute_reply.started":"2025-01-22T06:47:02.682224Z","shell.execute_reply":"2025-01-22T06:47:02.689127Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **Checking the number of null values in a table column is crucial for several reasons:**\n\n**Data Quality Assessment**: Null values represent missing or unknown information. Identifying their frequency helps understand the completeness and quality of your data.\n\n**Data Cleaning**: Null values can impact data analysis and reporting. Knowing their extent allows you to:\nImpute missing values: Replace nulls with reasonable estimates (e.g., mean, median, or using machine learning techniques).\n\n**Remove rows with null values**: If the number of nulls is significant and imputation isn't feasible.\n\n**Analysis Accuracy**: Null values can skew statistical calculations and produce misleading results. Cleaning the data ensures accurate analysis.\n\n**Data Integrity**: Some columns might have constraints (e.g., NOT NULL) that disallow null values. Checking helps enforce these constraints.\n\n**Database Design**: The presence of many null values might indicate design flaws or the need for adjustments to data collection processes.","metadata":{}},{"cell_type":"code","source":"print(df_train.isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:02.692021Z","iopub.execute_input":"2025-01-22T06:47:02.692646Z","iopub.status.idle":"2025-01-22T06:47:03.460381Z","shell.execute_reply.started":"2025-01-22T06:47:02.6926Z","shell.execute_reply":"2025-01-22T06:47:03.459114Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"By **understanding the categories and ranges of the data**, we can preprocess it effectively, engineer relevant features, select appropriate models, and ultimately build more accurate and robust machine learning models.\n\nIt helps in various Perspective: \n\n**1. Data Preprocessing:**\n\n**a. Categorical Columns:**\n\n**Encoding**: Many machine learning algorithms require numerical input. You'll need to convert categorical values into numerical representations.\n- One-Hot Encoding: Creates binary columns for each category. Understanding the categories helps determine the number of new columns needed.\n- Label Encoding: Assigns a unique integer to each category. Useful when there's an inherent order in the categories.\n\n**Handling Rare Categories**: Identifying rare categories can help decide whether to group them into an \"other\" category to prevent overfitting.\n\n\n**b. Numerical Columns:**\n\n**Scaling**: Some algorithms are sensitive to feature scales.\n\n- Normalization (Min-Max Scaling): Scales features to a specific range (e.g., 0 to 1). Knowing the range helps determine the scaling factor.\n- Standardization (Z-score Normalization): Transforms features to have zero mean and unit variance. Understanding the range can help identify potential outliers.\n\n**Outlier Detection**: The range can help identify potential outliers that might skew model performance.\n\n**2. Feature Engineering:**\n\n- Creating New Features: Understanding the categories and ranges can inspire the creation of new features that might improve model performance.\n  \n- Categorical Interactions: Creating new features by combining categories from different columns.\nBinning Numerical Features: Dividing numerical features into bins can capture non-linear relationships.\n\n**3. Model Selection:**\n\n- Algorithm Choice: Some algorithms are more suitable for categorical data, while others are better for numerical data. Understanding the data characteristics helps choose the appropriate model.\n- Hyperparameter Tuning: The range of numerical features can influence the choice of hyperparameters for certain algorithms.\n\n**4. Data Visualization:**\n\n- Understanding Data Distribution: The range of numerical features helps choose appropriate scales for visualizations (e.g., histograms, box plots).\n- Identifying Relationships: Visualizing categorical features can reveal patterns and relationships with other variables.","metadata":{}},{"cell_type":"code","source":"for i,col in enumerate(['Gender','Marital Status','Education Level','Occupation','Location','Policy Type','Policy Start Date','Customer Feedback','Smoking Status','Exercise Frequency','Property Type']):\n    print(col, ':', df_train[col].unique())    ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:03.461537Z","iopub.execute_input":"2025-01-22T06:47:03.462043Z","iopub.status.idle":"2025-01-22T06:47:04.422527Z","shell.execute_reply.started":"2025-01-22T06:47:03.461983Z","shell.execute_reply":"2025-01-22T06:47:04.421456Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for i,col in enumerate(['Age','Annual Income','Number of Dependents','Health Score','Previous Claims','Vehicle Age','Credit Score','Insurance Duration','Premium Amount']):\n    \n    start_range = df_train[col].min()\n    end_range = df_train[col].max()\n    print(f\"Column: {col}\")\n    print(f\"  Starting Range: {start_range}\")\n    print(f\"  Ending Range: {end_range}\")\n    print()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:04.423576Z","iopub.execute_input":"2025-01-22T06:47:04.423997Z","iopub.status.idle":"2025-01-22T06:47:04.478256Z","shell.execute_reply.started":"2025-01-22T06:47:04.42396Z","shell.execute_reply":"2025-01-22T06:47:04.477186Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**The target column (Premium Amount) is the central focus of any supervised machine learning problem. It's the variable we are trying to predict or understand.**","metadata":{}},{"cell_type":"code","source":"\nnum_buckets = 500\nplt.hist(df_train['Premium Amount'], bins=num_buckets, edgecolor='black')\n\nplt.xlabel('Value of Premium Amount')\nplt.ylabel('Number of Entries')\nplt.title('Histogram of Premium Amount')\n\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:04.479485Z","iopub.execute_input":"2025-01-22T06:47:04.479842Z","iopub.status.idle":"2025-01-22T06:47:05.555335Z","shell.execute_reply.started":"2025-01-22T06:47:04.479815Z","shell.execute_reply":"2025-01-22T06:47:05.554142Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**And now some experiments to understand the relationships between the features and see their distributions**","metadata":{}},{"cell_type":"markdown","source":"First I do not want nan in the categorical columns since I can treat them as \"Data Not Available\".","metadata":{}},{"cell_type":"code","source":"\ncolumns_to_replace = ['Gender','Marital Status','Education Level','Occupation','Location','Policy Type','Policy Start Date','Customer Feedback','Smoking Status','Exercise Frequency','Property Type']\n \n# Replace NaN values with \"Data Not Available\"\ndf_train[columns_to_replace] = df_train[columns_to_replace].fillna(\"Data Not Available\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:05.557866Z","iopub.execute_input":"2025-01-22T06:47:05.55819Z","iopub.status.idle":"2025-01-22T06:47:06.997637Z","shell.execute_reply.started":"2025-01-22T06:47:05.558163Z","shell.execute_reply":"2025-01-22T06:47:06.996516Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for i,col in enumerate(['Gender','Marital Status','Education Level','Occupation','Location','Policy Type','Customer Feedback','Smoking Status','Exercise Frequency','Property Type']):\n    print(col, ':', df_train[col].unique()) ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:06.998881Z","iopub.execute_input":"2025-01-22T06:47:06.999159Z","iopub.status.idle":"2025-01-22T06:47:07.681719Z","shell.execute_reply.started":"2025-01-22T06:47:06.999136Z","shell.execute_reply":"2025-01-22T06:47:07.680659Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now the categorical columns data distribution:","metadata":{}},{"cell_type":"code","source":"cat_cols = ['Gender','Marital Status','Education Level','Occupation','Location','Policy Type','Customer Feedback','Smoking Status','Exercise Frequency','Property Type']\n\n\nfig, axes = plt.subplots(5, 2, figsize=(10, 15)) \n\nfor i, col in enumerate(cat_cols):\n    row = i // 2  \n    col_idx = i % 2  \n    \n    counts = df_train[col].value_counts()\n  \n    axes[row, col_idx].pie(counts, labels=counts.index, autopct='%1.1f%%', startangle=140)\n    axes[row, col_idx].set_title(f\"Distribution of {col}\")\n\nplt.tight_layout()  \nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:07.682804Z","iopub.execute_input":"2025-01-22T06:47:07.683221Z","iopub.status.idle":"2025-01-22T06:47:09.661407Z","shell.execute_reply.started":"2025-01-22T06:47:07.683183Z","shell.execute_reply":"2025-01-22T06:47:09.660088Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now let's understand the numerical columns:","metadata":{}},{"cell_type":"code","source":"num_cols = ['Age','Annual Income','Number of Dependents','Health Score','Previous Claims','Vehicle Age','Credit Score','Insurance Duration']\n\nfig, axes = plt.subplots(4, 2, figsize=(10, 12)) \n\n\ncolors = ['steelblue', 'forestgreen', 'indianred', 'goldenrod', 'darkorchid', 'teal', 'lightblue', 'olive'] \n\nfor i, col in enumerate(num_cols):\n    row = i // 2  \n    col_idx = i % 2  \n    \n    axes[row, col_idx].hist(df_train[col], bins=20, color=colors[i])  \n    axes[row, col_idx].set_title(f\"Distribution of {col}\")\n\nplt.tight_layout()  \nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:09.662519Z","iopub.execute_input":"2025-01-22T06:47:09.662814Z","iopub.status.idle":"2025-01-22T06:47:12.228472Z","shell.execute_reply.started":"2025-01-22T06:47:09.66279Z","shell.execute_reply":"2025-01-22T06:47:12.227136Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"le = LabelEncoder()\n# Convert categorical columns to numerical using LabelEncoder\nfor col in cat_cols:\n    df_train[col] = le.fit_transform(df_train[col])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:12.229863Z","iopub.execute_input":"2025-01-22T06:47:12.230193Z","iopub.status.idle":"2025-01-22T06:47:14.712497Z","shell.execute_reply.started":"2025-01-22T06:47:12.230163Z","shell.execute_reply":"2025-01-22T06:47:14.711092Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for i,col in enumerate(['Gender','Marital Status','Education Level','Occupation','Location','Policy Type','Customer Feedback','Smoking Status','Exercise Frequency','Property Type']):\n    print(col, ':', df_train[col].unique()) ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:14.71361Z","iopub.execute_input":"2025-01-22T06:47:14.713951Z","iopub.status.idle":"2025-01-22T06:47:14.796853Z","shell.execute_reply.started":"2025-01-22T06:47:14.713922Z","shell.execute_reply":"2025-01-22T06:47:14.795696Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train.dtypes","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:14.797817Z","iopub.execute_input":"2025-01-22T06:47:14.798091Z","iopub.status.idle":"2025-01-22T06:47:14.806072Z","shell.execute_reply.started":"2025-01-22T06:47:14.79807Z","shell.execute_reply":"2025-01-22T06:47:14.804867Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_1 = df_train.drop('Policy Start Date', axis=1)\ncorr_mat = df_1.corr()\n\nplt.figure(figsize=(22,7))\nmask = np.zeros_like(corr_mat)\nmask[np.triu_indices_from(mask)] = True\nsns.heatmap(corr_mat, \n            mask=mask,\n            annot=corr_mat.round(2), \n            cmap='coolwarm',  \n            vmin=-1, vmax=1,  \n            center=0,        \n            linewidths=.5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:14.807202Z","iopub.execute_input":"2025-01-22T06:47:14.807643Z","iopub.status.idle":"2025-01-22T06:47:17.820574Z","shell.execute_reply.started":"2025-01-22T06:47:14.807572Z","shell.execute_reply":"2025-01-22T06:47:17.819492Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Positive Correlation:**\n\n- Definition: When two variables move in the same direction.   \n- Interpretation: As one variable increases, the other variable also tends to increase.\n(Previous Claims, Annual Income), (Health Score, Annual Income), (Credit Score,Previous Claims), (Premium Amount, Previous Claims)\n   \n\n\n**Negative Correlation:**\n\n- Definition: When two variables move in opposite directions.   \n- Interpretation: As one variable increases, the other variable tends to decrease.\n(Credit Score, Annual Income), (Credit Score, Premium Amount)\n","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10, 6)) \nplt.scatter(df_train['Credit Score'],df_train['Previous Claims'],  c=df_train['Premium Amount'], cmap='viridis') \nplt.xlabel('Credit Score')\nplt.ylabel('Previous Claims')\n\nplt.title('Scatter Plot of Credit Score vs. Previous Claims Colored by Target')\nplt.colorbar() \nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:17.821794Z","iopub.execute_input":"2025-01-22T06:47:17.822185Z","iopub.status.idle":"2025-01-22T06:47:32.904871Z","shell.execute_reply.started":"2025-01-22T06:47:17.822146Z","shell.execute_reply":"2025-01-22T06:47:32.903669Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(10, 6)) \nplt.scatter(df_train['Annual Income'], df_train['Previous Claims'], c=df_train['Credit Score'], cmap='viridis') \nplt.xlabel('Annual Income')\nplt.ylabel('Previous Claims')\nplt.title('Scatter Plot of Annual Income vs. Previous Claims Colored by Credit Score')\nplt.colorbar() \nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-22T06:47:32.906179Z","iopub.execute_input":"2025-01-22T06:47:32.906524Z"}},"outputs":[],"execution_count":null}]}