{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"text-align:center; font-size:48px; color:red;\">🏦 Regression with Insurance Data 💵</h1>\n\n<center>\n    <img src=\"https://www.insuranceproaz.com/wp-content/uploads/2018/08/main-banner.jpg\">\n</center>\n\n***\n\n<div style=\"border: 4px solid blue; padding:4px 2px;\">\n    <p style=\"text-align: center; font-family:monospace; font-size:20px; font-height: 1.7em;\">Thank you for visiting my notebook. If you like the notebook, please don't forget to upvote it. Thank you!</p>\n    <p style=\"text-align: center; font-family:monospace; font-size:20px; font-height: 1.7em;\">It takes a lot of effort to produce these kernels, so if you are forking it, please make an effort to upvote it as well.</p>\n</div>\n\n***\n\n<h1 style=\"font-size:36px;\"> 📖 Introduction 📖</h1>\n\n<div class=\"alert alert-block alert-success\" style=\"padding: 2em; background-color: #fff; border-radius: 5em; border-width:4px; font-size:20px; font-family: verdana; line-height: 1.7em;\">\n    \nWelcome to Episode 12 of the Playground Series on Kaggle! In this notebook, we are excited to explore the challenge of predicting insurance premiums using a comprehensive insurance dataset. This episode is designed to engage data enthusiasts at all levels, providing an opportunity to apply regression techniques to a real-world problem.\n\nIn this episode, we will focus on a dataset that captures various factors influencing insurance costs, including demographic information, policy details, and historical claims data. Our main objective is to build a robust regression model that accurately estimates insurance premiums based on these features.\n\nThroughout this notebook, we will cover essential steps in the data analysis process, including:\n<ul>\n<li><strong>Data Cleaning:</strong> Preparing the dataset for analysis by handling missing values and outliers.</li>\n<li><strong>Exploratory Data Analysis (EDA):</strong> Visualizing and analyzing the data to uncover patterns and relationships that may inform our modeling approach.</li>\n<li><strong>Feature Engineering:</strong> Creating new features that can enhance the performance of our regression models.</li>\n<li><strong>Modeling:</strong> Implementing various regression algorithms and evaluating their performance to find the best fit for our data.</li>\n</ul>\nWe encourage you to follow along, experiment with different approaches, and share your insights and results. The Kaggle community thrives on collaboration and knowledge sharing, so your contributions are invaluable.\n\nLet’s dive into the Regression with an Insurance Dataset and see what insights we can uncover together!\n</div>\n\n<h1 style=\"font-size:36px\">🔬 Feature Description 🔬</h1>\n\n<div class=\"alert alert-block alert-info\" style=\"padding: 2em; background-color: #fff; border-radius: 5em; border-width:4px;\">\n    <ol style=\"font-size:20px; font-family:verdana; line-height:1.7em;\">\n        <li><strong>Age:</strong> The age of the individual, typically in years. It could be used to predict insurance risks and premiums, as age often correlates with health and driving habits.</li>\n        <li><strong>Annual Income:</strong> The yearly income of the individual. This could help predict the customer's ability to pay premiums and might correlate with risk assessment (e.g., higher income might correlate with lower claims).</li>\n        <li><strong>Number of Dependents:</strong> The count of individuals that the customer supports, such as children or elderly parents. More dependents may influence insurance choices and policy types.</li>\n        <li><strong>Health Score:</strong> A numerical value representing the overall health condition of the individual, which can be relevant for predicting medical or life insurance claims.</li>\n        <li><strong>Previous Claims:</strong> A count of previous insurance claims made by the individual. This feature can be crucial for assessing the likelihood of future claims.</li>\n        <li><strong>Vehicle Age:</strong> The age of the vehicle the customer owns. Older vehicles could indicate higher maintenance costs and might affect auto insurance policies and premiums.</li>\n        <li><strong>Credit Score:</strong> A numerical score representing the customer's creditworthiness. Insurance companies may use this to predict the risk of fraud or the likelihood of timely payments.</li>\n        <li><strong>Insurance Duration:</strong> The number of years the individual has been with the current insurance provider. Longer durations may indicate a loyal customer, while a shorter duration might reflect new customers or those at risk of switching.</li>\n        <li><strong>Day:</strong> The day of the month when a transaction or insurance policy is processed, which could be part of time series features or used for detecting seasonal trends in claims.</li>\n        <li><strong>Month:</strong> The month of the year, which may help identify seasonality in claims, such as an increase in accidents during winter months or claims related to weather.</li>\n        <li><strong>Year:</strong> The year of the transaction or policy inception, used for identifying long-term trends in data, like the growth of claims over time.</li>\n        <li><strong>Gender:</strong> The gender of the customer. Gender may be used in some models to analyze risk, although this could be sensitive in certain contexts.</li>\n        <li><strong>Marital Status:</strong> Whether the individual is single, married, or divorced. Marital status may correlate with lifestyle factors, such as driving habits or home insurance choices.</li>\n        <li><strong>Education Level:</strong> The highest level of education attained by the customer. This could be a factor in understanding socio-economic status and might influence lifestyle-related claims.</li>\n        <li><strong>Occupation:</strong> The profession of the individual. Certain occupations may have higher or lower insurance risks (e.g., a professional driver versus a desk job).</li>\n        <li><strong>Location:</strong> The geographic location of the individual, which can affect risk assessments due to environmental factors (e.g., higher accident rates in urban areas or higher claims in disaster-prone regions).</li>\n        <li><strong>Policy Type:</strong> The type of insurance policy the customer has, such as auto, life, health, or property insurance. Different policy types come with different risk profiles and premiums.</li>\n        <li><strong>Customer Feedback:</strong> A rating or qualitative measure of customer satisfaction or feedback. This could be used to gauge customer retention or potential risk of cancellation.</li>\n        <li><strong>Smoking Status:</strong> Whether the individual smokes, which could have a significant impact on health insurance premiums and risk assessments.</li>\n        <li><strong>Exercise Frequency:</strong> How often the individual engages in physical activity. This can be used to assess health risks, as higher exercise frequency might correlate with better health outcomes.</li>\n        <li><strong>Type:</strong> The type of property owned by the customer, such as a house, condo, or apartment. This can influence property insurance rates, with larger or more valuable properties often requiring higher premiums.</li>\n    </ol>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<h1 style=\"font-size:36px;\">📚 Import Libraries 📚</h1>","metadata":{}},{"cell_type":"code","source":"%%capture\n\n!pip uninstall -y scikit-learn\n!pip install scikit-learn==1.5.2","metadata":{"execution":{"iopub.status.busy":"2025-01-11T07:09:32.808539Z","iopub.execute_input":"2025-01-11T07:09:32.808845Z","iopub.status.idle":"2025-01-11T07:09:41.289576Z","shell.execute_reply.started":"2025-01-11T07:09:32.808819Z","shell.execute_reply":"2025-01-11T07:09:41.28852Z"},"_kg_hide-output":false,"_kg_hide-input":false,"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport gc\nimport warnings\n\nimport sklearn\nsklearn.set_config(transform_output=\"pandas\")\nimport category_encoders as ce\nfrom sklearn.preprocessing import StandardScaler, FunctionTransformer, LabelEncoder, OneHotEncoder, MinMaxScaler\nfrom sklearn.experimental import enable_iterative_imputer\nfrom sklearn.impute import SimpleImputer, IterativeImputer\nfrom sklearn.pipeline import make_pipeline, Pipeline\nfrom sklearn.compose import ColumnTransformer, make_column_selector, make_column_transformer\nfrom sklearn.metrics import root_mean_squared_log_error, mean_squared_log_error, make_scorer\nfrom sklearn.model_selection import cross_val_score, KFold\nfrom sklearn.feature_selection import mutual_info_regression\n\nfrom sklearn.linear_model import LinearRegression\nfrom xgboost import XGBRegressor, DMatrix, XGBRFRegressor\nfrom catboost import CatBoostRegressor, Pool\nfrom lightgbm import LGBMRegressor, early_stopping\nfrom sklearn.ensemble import HistGradientBoostingRegressor, VotingRegressor\n\n# from autogluon.tabular import TabularDataset, TabularPredictor\n\nwarnings.filterwarnings(\"ignore\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:09:41.290709Z","iopub.execute_input":"2025-01-11T07:09:41.290969Z","iopub.status.idle":"2025-01-11T07:09:46.417404Z","shell.execute_reply.started":"2025-01-11T07:09:41.290949Z","shell.execute_reply":"2025-01-11T07:09:46.416711Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h1 style=\"font-size:36px;\">💾 Load Data 📀</h1>","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv', index_col='id', engine='pyarrow')\ntest = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv', index_col='id', engine='pyarrow')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:09:46.4186Z","iopub.execute_input":"2025-01-11T07:09:46.419158Z","iopub.status.idle":"2025-01-11T07:09:49.161846Z","shell.execute_reply.started":"2025-01-11T07:09:46.419134Z","shell.execute_reply":"2025-01-11T07:09:49.161176Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T12:23:13.682177Z","iopub.execute_input":"2024-12-30T12:23:13.682392Z","iopub.status.idle":"2024-12-30T12:23:13.711572Z","shell.execute_reply.started":"2024-12-30T12:23:13.682374Z","shell.execute_reply":"2024-12-30T12:23:13.710727Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T12:23:13.712337Z","iopub.execute_input":"2024-12-30T12:23:13.712661Z","iopub.status.idle":"2024-12-30T12:23:13.730281Z","shell.execute_reply.started":"2024-12-30T12:23:13.712631Z","shell.execute_reply":"2024-12-30T12:23:13.729331Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def date_separator(df):\n    df = df.copy()\n    df['Start Day'] = df['Policy Start Date'].dt.day\n    df['Start Month'] = df['Policy Start Date'].dt.month\n    df['Start Year'] = df['Policy Start Date'].dt.year\n    return df\n\ntrain = date_separator(train)\ntest = date_separator(test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:09:49.162957Z","iopub.execute_input":"2025-01-11T07:09:49.163237Z","iopub.status.idle":"2025-01-11T07:09:50.389641Z","shell.execute_reply.started":"2025-01-11T07:09:49.163215Z","shell.execute_reply":"2025-01-11T07:09:50.388711Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"target = 'Premium Amount'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:09:50.390609Z","iopub.execute_input":"2025-01-11T07:09:50.390877Z","iopub.status.idle":"2025-01-11T07:09:50.394452Z","shell.execute_reply.started":"2025-01-11T07:09:50.390855Z","shell.execute_reply":"2025-01-11T07:09:50.393704Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"numerical_features = train.drop(target, axis=1).select_dtypes(include=np.number).columns.values\nprint(numerical_features)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:09:50.396179Z","iopub.execute_input":"2025-01-11T07:09:50.396472Z","iopub.status.idle":"2025-01-11T07:09:50.612495Z","shell.execute_reply.started":"2025-01-11T07:09:50.396445Z","shell.execute_reply":"2025-01-11T07:09:50.611606Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"categorical_features = train.drop(target, axis=1).select_dtypes(include='object').columns.values\nprint(categorical_features)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:09:50.613675Z","iopub.execute_input":"2025-01-11T07:09:50.614018Z","iopub.status.idle":"2025-01-11T07:09:50.898838Z","shell.execute_reply.started":"2025-01-11T07:09:50.613983Z","shell.execute_reply":"2025-01-11T07:09:50.898108Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T12:23:15.38699Z","iopub.execute_input":"2024-12-30T12:23:15.387236Z","iopub.status.idle":"2024-12-30T12:23:15.894698Z","shell.execute_reply.started":"2024-12-30T12:23:15.387203Z","shell.execute_reply":"2024-12-30T12:23:15.89395Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.isna().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T12:23:15.895493Z","iopub.execute_input":"2024-12-30T12:23:15.895814Z","iopub.status.idle":"2024-12-30T12:23:16.381218Z","shell.execute_reply.started":"2024-12-30T12:23:15.895784Z","shell.execute_reply":"2024-12-30T12:23:16.380505Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.duplicated().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T12:23:16.382162Z","iopub.execute_input":"2024-12-30T12:23:16.382399Z","iopub.status.idle":"2024-12-30T12:23:17.598768Z","shell.execute_reply.started":"2024-12-30T12:23:16.382368Z","shell.execute_reply":"2024-12-30T12:23:17.597932Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h1 style=\"font-size:36px;\">📈 EDA 📊</h1>","metadata":{}},{"cell_type":"code","source":"train[numerical_features].astype(np.float_).describe().T","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T12:23:17.599494Z","iopub.execute_input":"2024-12-30T12:23:17.59975Z","iopub.status.idle":"2024-12-30T12:23:18.358629Z","shell.execute_reply.started":"2024-12-30T12:23:17.59973Z","shell.execute_reply":"2024-12-30T12:23:18.357833Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.describe(include='O').T","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T12:23:18.359402Z","iopub.execute_input":"2024-12-30T12:23:18.359699Z","iopub.status.idle":"2024-12-30T12:23:19.408033Z","shell.execute_reply.started":"2024-12-30T12:23:18.359676Z","shell.execute_reply":"2024-12-30T12:23:19.407239Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Distribution of Features\n\n**We will be using 10% of the training and test data to get an idea of how the features are distributed as our dataset is pretty big.**","metadata":{}},{"cell_type":"code","source":"train_sample = train.sample(frac=0.1)\ntest_sample = test.sample(frac=0.1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T12:25:00.13024Z","iopub.execute_input":"2024-12-30T12:25:00.130568Z","iopub.status.idle":"2024-12-30T12:25:00.296402Z","shell.execute_reply.started":"2024-12-30T12:25:00.13054Z","shell.execute_reply":"2024-12-30T12:25:00.295662Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def numerical_feature_distribution_plot(col, train, test):\n    fig, axes = plt.subplot_mosaic([['hist', 'hist'], ['box', 'violin']], figsize=(14,9))\n    \n    sns.histplot(train, x=col, ax=axes['hist'])\n    sns.boxplot(train, x=col, ax=axes['box'])\n    sns.violinplot(train, x=col, ax=axes['violin'])\n\n    if len(test)>0:\n        sns.histplot(test, x=col, ax=axes['hist'])\n        axes['hist'].legend(['train', 'test'])\n        \n        sns.boxplot(test, x=col, ax=axes['box'])\n        axes['box'].legend(['train', 'test'])\n\n        sns.violinplot(test, x=col, ax=axes['violin'])\n        axes['violin'].legend(['train', 'test'])\n        \n    axes['box'].set_title(f\"Boxplot of {col}\", size=20, weight='bold')\n    axes['hist'].set_title(f\"Histogram of {col}\", size=20, weight='bold')\n    axes['violin'].set_title(f\"Violin plot of {col}\", size=20, weight='bold')\n\n    fig.tight_layout()\n    fig.suptitle(f\"Distribution of {col}\", y=1.05, size=36, color='red')\n    fig.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T11:41:01.817423Z","iopub.execute_input":"2024-12-30T11:41:01.817653Z","iopub.status.idle":"2024-12-30T11:41:01.8241Z","shell.execute_reply.started":"2024-12-30T11:41:01.817617Z","shell.execute_reply":"2024-12-30T11:41:01.823205Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for col in numerical_features:\n    numerical_feature_distribution_plot(col, train_sample, test_sample)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T11:41:01.825017Z","iopub.execute_input":"2024-12-30T11:41:01.825249Z","iopub.status.idle":"2024-12-30T11:41:17.309056Z","shell.execute_reply.started":"2024-12-30T11:41:01.82523Z","shell.execute_reply":"2024-12-30T11:41:17.308204Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def categorical_feature_distribution_plot(col, train, test):\n    fig, axes = plt.subplots(1,2,figsize=(16,6))\n\n    sns.countplot(train, x=col, ax=axes[0], palette='plasma')\n    for container in axes[0].containers:\n        axes[0].bar_label(container)\n    axes[0].set_title(f\"Barplot of {col}\", size=20, weight='bold')\n    \n    axes[1] = train[col].value_counts().plot.pie(autopct=\"%.2f%%\", pctdistance=0.75, colors=sns.color_palette('Set2'))\n    axes[1].add_artist(plt.Circle((0,0), radius=0.5, facecolor='w'))\n    axes[1].set_title(f\"Pie Chart of {col}\", size=20, weight='bold')\n    axes[1].set_ylabel(\"\")\n\n    plt.tight_layout()\n    plt.suptitle(f\"Distribution of {col}\", y=1.1, size=36, color='red')\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T11:41:17.310088Z","iopub.execute_input":"2024-12-30T11:41:17.310397Z","iopub.status.idle":"2024-12-30T11:41:17.316286Z","shell.execute_reply.started":"2024-12-30T11:41:17.310373Z","shell.execute_reply":"2024-12-30T11:41:17.315433Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for col in categorical_features:\n    categorical_feature_distribution_plot(col, train_sample, test_sample)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T11:41:17.317175Z","iopub.execute_input":"2024-12-30T11:41:17.317446Z","iopub.status.idle":"2024-12-30T11:41:21.401187Z","shell.execute_reply.started":"2024-12-30T11:41:17.317426Z","shell.execute_reply":"2024-12-30T11:41:21.400225Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"numerical_feature_distribution_plot(target, train_sample, [])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T11:41:21.402137Z","iopub.execute_input":"2024-12-30T11:41:21.40243Z","iopub.status.idle":"2024-12-30T11:41:22.491685Z","shell.execute_reply.started":"2024-12-30T11:41:21.402397Z","shell.execute_reply":"2024-12-30T11:41:22.490853Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(16,10))\nfor i, col in enumerate(numerical_features):\n    plt.subplot(3,4,i+1)\n    sns.regplot(train_sample, x=col, y=target)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T11:41:22.492713Z","iopub.execute_input":"2024-12-30T11:41:22.49304Z","iopub.status.idle":"2024-12-30T11:42:17.798989Z","shell.execute_reply.started":"2024-12-30T11:41:22.493007Z","shell.execute_reply":"2024-12-30T11:42:17.798074Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_new = train_sample.copy()\nfor col in categorical_features:\n    train_new[col], _ = train_new[col].factorize()\n\ncor_mat = train_new.corr(method='pearson')\nmask = np.triu(cor_mat)\nplt.figure(figsize=(16,10))\nsns.heatmap(cor_mat, mask=mask, annot=True, cmap='winter', fmt=\".2f\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T11:42:17.79984Z","iopub.execute_input":"2024-12-30T11:42:17.800061Z","iopub.status.idle":"2024-12-30T11:42:18.811899Z","shell.execute_reply.started":"2024-12-30T11:42:17.800042Z","shell.execute_reply":"2024-12-30T11:42:18.810936Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h1 style=\"font-size:36px;\">🤖 Model Training ⚙️</h1>","metadata":{}},{"cell_type":"code","source":"def model_trainer(model, X, y, test, n_splits=5, random_state=42, model_name=None):\n    kfold = KFold(n_splits, random_state=random_state, shuffle=True)\n\n\n    print(\"=\"*80)\n    print(f\"Model: {model.__class__.__name__}\")\n    print(\"=\"*80)\n\n    scores = []\n    oof_preds = np.zeros(len(test))\n    oof_train_preds = np.zeros(len(y))\n\n    for i, (train_idx, valid_idx) in enumerate(kfold.split(X)):\n        X_train, y_train = X.iloc[train_idx], y[train_idx]\n        X_valid, y_valid = X.iloc[valid_idx], y[valid_idx]\n\n        if model_name == 'xgb':\n            model.fit(X_train, y_train, eval_set=[(X_valid, y_valid)], verbose=0)\n            booster = model.get_booster()\n            \n            y_pred = booster.predict(DMatrix(X_valid), iteration_range=(0, model.best_iteration+1))\n            test_pred = booster.predict(DMatrix(test), iteration_range=(0, model.best_iteration+1))\n            oof_train_preds[train_idx] = booster.predict(DMatrix(X_train), iteration_range=(0, model.best_iteration+1))\n\n        elif model_name == 'cat':\n            trainPool = Pool(X_train ,y_train)\n            testPool = Pool(test)\n            validPool = Pool(X_valid, y_valid)\n\n            model.fit(X=trainPool, eval_set=validPool, verbose=0, early_stopping_rounds=200)\n            y_pred = model.predict(validPool)\n            test_pred = model.predict(testPool)\n            oof_train_preds[train_idx] = model.predict(Pool(X_train))\n\n        elif model_name == 'lgb':\n            model.fit(X_train, y_train, eval_set=[(X_valid, y_valid)], eval_metric='rmse', callbacks=[early_stopping(200, verbose=0)])\n            y_pred = model.predict(X_valid, num_iteration=model.best_iteration_)\n            test_pred = model.predict(test, num_iteration=model.best_iteration_)\n            oof_train_preds[train_idx] = model.predict(X_train, num_iteration=model.best_iteration_)\n\n        else:\n            model.fit(X_train, y_train)\n            y_pred = model.predict(X_valid)\n            test_pred = model.predict(test)\n            oof_train_preds[train_idx] = model.predict(X_train)\n        \n        rmsle = root_mean_squared_log_error(np.expm1(y_valid), np.expm1(y_pred))\n        oof_preds += test_pred\n        print(f\"Fold {i+1} --> RMSLE: {rmsle:.5f}\")\n        scores.append(rmsle)\n\n    print(f\"\\nAverage Fold RMSLE: {np.mean(scores):.5f} \\xb1 {np.std(scores):.5f}\")\n    print(\"\\n\")\n    return oof_preds/n_splits, oof_train_preds","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:10:43.332588Z","iopub.execute_input":"2025-01-11T07:10:43.332893Z","iopub.status.idle":"2025-01-11T07:10:43.342454Z","shell.execute_reply.started":"2025-01-11T07:10:43.332869Z","shell.execute_reply":"2025-01-11T07:10:43.341582Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"preprocessing = ColumnTransformer([\n    ('num', make_pipeline(\n        SimpleImputer(strategy='mean', add_indicator=True), \n        FunctionTransformer(), StandardScaler()), \n     numerical_features),\n    ('cat', make_pipeline(SimpleImputer(strategy='most_frequent'), ce.cat_boost.CatBoostEncoder()), categorical_features)\n], remainder='drop')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:10:51.332774Z","iopub.execute_input":"2025-01-11T07:10:51.333075Z","iopub.status.idle":"2025-01-11T07:10:51.337294Z","shell.execute_reply.started":"2025-01-11T07:10:51.333045Z","shell.execute_reply":"2025-01-11T07:10:51.336459Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X = train.copy()\ny = X.pop(target)\ny = np.log1p(y)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:10:53.429128Z","iopub.execute_input":"2025-01-11T07:10:53.429451Z","iopub.status.idle":"2025-01-11T07:10:53.572779Z","shell.execute_reply.started":"2025-01-11T07:10:53.429426Z","shell.execute_reply":"2025-01-11T07:10:53.572095Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_processed = preprocessing.fit_transform(X, y)\ntestProcessed = preprocessing.transform(test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:10:55.127942Z","iopub.execute_input":"2025-01-11T07:10:55.128231Z","iopub.status.idle":"2025-01-11T07:11:05.457949Z","shell.execute_reply.started":"2025-01-11T07:10:55.128209Z","shell.execute_reply":"2025-01-11T07:11:05.456917Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Mutual Information Score","metadata":{}},{"cell_type":"code","source":"X_sample = train_sample.copy()\ny_sample = X_sample.pop(target)\ny_sample = np.log1p(y_sample)\nX_mi = preprocessing.transform(X_sample)\n\nmi_scores = mutual_info_regression(X_mi, y_sample)\nmi_scores_df = pd.DataFrame({\n    'Feature': X_mi.columns.values,\n    'Score': mi_scores\n}).sort_values(by='Score', ascending=False).reset_index(drop=True)\n\nmi_scores_df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T12:34:49.494955Z","iopub.execute_input":"2024-12-30T12:34:49.495165Z","iopub.status.idle":"2024-12-30T12:35:18.407964Z","shell.execute_reply.started":"2024-12-30T12:34:49.495147Z","shell.execute_reply":"2024-12-30T12:35:18.407201Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sns.barplot(mi_scores_df, x='Score', y='Feature')\nplt.title(\"Mutual Information Score\", size=36, color='red', y=1.05)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-30T12:35:18.409364Z","iopub.execute_input":"2024-12-30T12:35:18.409671Z","iopub.status.idle":"2024-12-30T12:35:18.820792Z","shell.execute_reply.started":"2024-12-30T12:35:18.409648Z","shell.execute_reply":"2024-12-30T12:35:18.819988Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Model Predictions","metadata":{}},{"cell_type":"code","source":"test_preds = pd.DataFrame()\ntrain_preds = pd.DataFrame()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:11:19.430398Z","iopub.execute_input":"2025-01-11T07:11:19.430691Z","iopub.status.idle":"2025-01-11T07:11:19.436293Z","shell.execute_reply.started":"2025-01-11T07:11:19.43067Z","shell.execute_reply":"2025-01-11T07:11:19.435526Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## XGBoost","metadata":{}},{"cell_type":"code","source":"%%time\n\nxgb_params = {\n    'objective': 'reg:squaredlogerror',\n    'eval_metric': 'rmse',\n    'n_estimators': 3000, # 166 \n    'learning_rate': 0.09928981614640733, \n    'subsample': 0.9981966139506873, \n    'colsample_bytree': 0.9700848676733003, \n    'max_depth': 31, \n    'lambda': 1.6381115886394215e-05, \n    'alpha': 0.8746443028529881, \n    'gamma': 0.11687689198413533, \n    'grow_policy': 'lossguide', \n    # 'n_jobs': 4,\n    'early_stopping_rounds': 200,\n    'device': 'cuda',\n    'tree_method': 'hist'\n}\n\nxgb = XGBRegressor(**xgb_params)\n\ntest_preds['xgb'], train_preds['xgb'] = model_trainer(xgb, X_processed, y, testProcessed, random_state=0, model_name='xgb')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:12:10.147839Z","iopub.execute_input":"2025-01-11T07:12:10.148135Z","iopub.status.idle":"2025-01-11T07:13:00.333721Z","shell.execute_reply.started":"2025-01-11T07:12:10.148113Z","shell.execute_reply":"2025-01-11T07:13:00.332905Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## CatBoost","metadata":{}},{"cell_type":"code","source":"%%time\n\ncat = CatBoostRegressor(\n    n_estimators=3000,\n    learning_rate=0.05, \n    # od_type='Iter',\n    # od_wait=20,\n    # random_state=42,\n    # use_best_model=True,\n    task_type='GPU', \n    verbose=False, \n    allow_writing_files=False, \n)\n\ntest_preds['cat'], train_preds['cat'] = model_trainer(cat, X_processed, y, testProcessed, random_state=0, model_name='cat')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:13:05.238649Z","iopub.execute_input":"2025-01-11T07:13:05.238947Z","iopub.status.idle":"2025-01-11T07:14:22.916153Z","shell.execute_reply.started":"2025-01-11T07:13:05.238922Z","shell.execute_reply":"2025-01-11T07:14:22.915273Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## LightGBM","metadata":{}},{"cell_type":"code","source":"%%time\n\nlgb = LGBMRegressor(\n    device='gpu', verbosity=-1, \n    n_estimators=3000, learning_rate=0.1\n)\n\ntest_preds['lgb'], train_preds['lgb'] = model_trainer(lgb, X_processed, y, testProcessed, random_state=0, model_name='lgb')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:14:22.917605Z","iopub.execute_input":"2025-01-11T07:14:22.91798Z","iopub.status.idle":"2025-01-11T07:15:45.270443Z","shell.execute_reply.started":"2025-01-11T07:14:22.91794Z","shell.execute_reply":"2025-01-11T07:15:45.269592Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## HistGradientBoosting","metadata":{}},{"cell_type":"code","source":"%%time\n\nhgb = HistGradientBoostingRegressor()\n\ntest_preds['hgb'], train_preds['hgb'] = model_trainer(hgb, X_processed, y, testProcessed, random_state=0)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:15:45.271675Z","iopub.execute_input":"2025-01-11T07:15:45.2719Z","iopub.status.idle":"2025-01-11T07:16:51.832927Z","shell.execute_reply.started":"2025-01-11T07:15:45.27188Z","shell.execute_reply":"2025-01-11T07:16:51.83198Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pred = np.mean(test_preds.to_numpy(), axis=1)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:19:28.979293Z","iopub.execute_input":"2025-01-11T07:19:28.979621Z","execution_failed":"2025-01-11T07:21:06.894Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<h1 style=\"font-size:36px;\">🏎️ Submission 🚔</h1>","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(\"/kaggle/input/playground-series-s4e12/sample_submission.csv\")\nsub[target] = np.expm1(pred)\nsub.to_csv(\"submission.csv\", index=False)\nsub.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-11T07:19:21.126744Z","iopub.execute_input":"2025-01-11T07:19:21.127082Z","iopub.status.idle":"2025-01-11T07:19:22.750832Z","shell.execute_reply.started":"2025-01-11T07:19:21.127045Z","shell.execute_reply":"2025-01-11T07:19:22.750105Z"}},"outputs":[],"execution_count":null}]}