{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"The provided solution walks through the following steps:\n\n1. Introduction & Setup\n2. Data Loading\n3. Initial Exploration and EDA\n4. Data Cleaning and Preprocessing\n5. Feature Engineering\n6. Train-Test Split (From Provided train/test data)\n7. Modeling with XGBoost\n8. Bayesian Optimization with Optuna (or a similar Bayesian optimization library)\n9. Model Evaluation using RMSLE\n10. Explainability with SHAP\n11. Conclusion\n","metadata":{}},{"cell_type":"markdown","source":"High-Level Notes\n\nRMSLE Metric:\n\nRMSLE = \n1\n�\n∑\n�\n=\n1\n�\n(\nlog\n⁡\n(\n�\n�\n�\n�\n�\n+\n1\n)\n−\nlog\n⁡\n(\n�\n�\n�\n�\n�\n�\n�\n+\n1\n)\n)\n2\nn\n1\n​\n ∑ \ni=1\nn\n​\n (log(pred \ni\n​\n +1)−log(actual \ni\n​\n +1)) \n2\n \n​\n \n\nWe will implement a custom metric function or use the built-in functions if available.\n\nSHAP:\nWe will use SHAP to understand global and local importance. Global importance comes from averaging SHAP values across the test set. Local importance can be demonstrated for a handful of individual samples.\n\nBayesian Optimization:\nWe will use Optuna to optimize XGBoost hyperparameters.","metadata":{}},{"cell_type":"code","source":"# ========================================\n# 1. Introduction & Setup\n# ========================================\n\n# This notebook demonstrates an end-to-end workflow: \n# EDA, data cleaning, feature engineering, \n# modeling with XGBoost, hyperparameter tuning using Optuna, \n# and explainability using SHAP.\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# For modeling\nfrom sklearn.model_selection import train_test_split, KFold\nfrom sklearn.metrics import mean_squared_log_error\nfrom sklearn.feature_extraction.text import TfidfVectorizer\n\n# XGBoost\nimport xgboost as xgb\n\n# For Bayesian Optimization\nimport optuna\n\n# For SHAP explainability\nimport shap\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\npd.set_option('display.max_columns', 100)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:39:19.722947Z","iopub.execute_input":"2024-12-14T16:39:19.724023Z","iopub.status.idle":"2024-12-14T16:39:19.732253Z","shell.execute_reply.started":"2024-12-14T16:39:19.72397Z","shell.execute_reply":"2024-12-14T16:39:19.731145Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ========================================\n# 2. Data Loading\n# ========================================\n\ntrain = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\ntest = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')\n\n# Check the shape\nprint(\"Train shape:\", train.shape)\nprint(\"Test shape:\", test.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:39:19.734483Z","iopub.execute_input":"2024-12-14T16:39:19.73496Z","iopub.status.idle":"2024-12-14T16:39:26.744312Z","shell.execute_reply.started":"2024-12-14T16:39:19.734912Z","shell.execute_reply":"2024-12-14T16:39:26.743232Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Identify target variable and ID column\nTARGET = 'Premium Amount'\nID_COL = 'id'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:39:26.745744Z","iopub.execute_input":"2024-12-14T16:39:26.746201Z","iopub.status.idle":"2024-12-14T16:39:26.75145Z","shell.execute_reply.started":"2024-12-14T16:39:26.746154Z","shell.execute_reply":"2024-12-14T16:39:26.750312Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ========================================\n# 3. Initial Exploration and EDA\n# ========================================\n\n# Quick look at the data\ndisplay(train.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:39:26.75421Z","iopub.execute_input":"2024-12-14T16:39:26.754834Z","iopub.status.idle":"2024-12-14T16:39:26.783132Z","shell.execute_reply.started":"2024-12-14T16:39:26.754798Z","shell.execute_reply":"2024-12-14T16:39:26.78176Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"display(train.describe(include='all'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:39:26.784367Z","iopub.execute_input":"2024-12-14T16:39:26.784714Z","iopub.status.idle":"2024-12-14T16:39:29.59321Z","shell.execute_reply.started":"2024-12-14T16:39:26.784669Z","shell.execute_reply":"2024-12-14T16:39:29.591926Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check data types\nprint(train.info())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:39:29.594501Z","iopub.execute_input":"2024-12-14T16:39:29.594827Z","iopub.status.idle":"2024-12-14T16:39:30.235089Z","shell.execute_reply.started":"2024-12-14T16:39:29.594794Z","shell.execute_reply":"2024-12-14T16:39:30.233787Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Distribution of target\nsns.histplot(train[TARGET], kde=True)\nplt.title(\"Distribution of Premium Amount\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:39:30.236485Z","iopub.execute_input":"2024-12-14T16:39:30.23695Z","iopub.status.idle":"2024-12-14T16:39:36.132979Z","shell.execute_reply.started":"2024-12-14T16:39:30.236902Z","shell.execute_reply":"2024-12-14T16:39:36.131684Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check missing values\nmissing_values = train.isnull().sum().sort_values(ascending=False)\nprint(\"Missing values in train:\\n\", missing_values)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:39:36.134307Z","iopub.execute_input":"2024-12-14T16:39:36.13463Z","iopub.status.idle":"2024-12-14T16:39:36.769478Z","shell.execute_reply.started":"2024-12-14T16:39:36.134597Z","shell.execute_reply":"2024-12-14T16:39:36.768285Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Categorical Features distribution\ncat_features = ['Gender','Marital Status','Education Level','Occupation',\n                'Location','Policy Type','Smoking Status','Exercise Frequency',\n                'Property Type']\n\nfor c in cat_features:\n    plt.figure(figsize=(6,4))\n    sns.countplot(x=c, data=train)\n    plt.title(f\"Distribution of {c}\")\n    plt.xticks(rotation=45)\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:39:36.773839Z","iopub.execute_input":"2024-12-14T16:39:36.774618Z","iopub.status.idle":"2024-12-14T16:39:44.243345Z","shell.execute_reply.started":"2024-12-14T16:39:36.774567Z","shell.execute_reply":"2024-12-14T16:39:44.241977Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Relationship between numeric features and target\nnum_features = ['Age','Annual Income','Number of Dependents','Health Score',\n                'Previous Claims','Vehicle Age','Credit Score','Insurance Duration']\nfig, axes = plt.subplots(len(num_features), 1, figsize=(6, 40))\n\nfor i, col in enumerate(num_features):\n    sns.scatterplot(x=train[col], y=train[TARGET], ax=axes[i])\n    axes[i].set_title(f\"{col} vs {TARGET}\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:39:44.244728Z","iopub.execute_input":"2024-12-14T16:39:44.245111Z","iopub.status.idle":"2024-12-14T16:40:04.242944Z","shell.execute_reply.started":"2024-12-14T16:39:44.245075Z","shell.execute_reply":"2024-12-14T16:40:04.24172Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check correlation among numerical features\ncorr = train[num_features+[TARGET]].corr()\nplt.figure(figsize=(10,8))\nsns.heatmap(corr, annot=True, cmap='coolwarm')\nplt.title(\"Correlation Heatmap\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:04.244355Z","iopub.execute_input":"2024-12-14T16:40:04.244685Z","iopub.status.idle":"2024-12-14T16:40:05.106291Z","shell.execute_reply.started":"2024-12-14T16:40:04.244651Z","shell.execute_reply":"2024-12-14T16:40:05.10509Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ========================================\n# 4. Data Cleaning and Preprocessing\n# ========================================\n\n# Example of cleaning steps:\n# - Fix incorrect data types\n# Convert Policy Start Date to datetime if present and not in correct format\ntrain['Policy Start Date'] = pd.to_datetime(train['Policy Start Date'], errors='coerce')\ntest['Policy Start Date'] = pd.to_datetime(test['Policy Start Date'], errors='coerce')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:05.107944Z","iopub.execute_input":"2024-12-14T16:40:05.108261Z","iopub.status.idle":"2024-12-14T16:40:05.750672Z","shell.execute_reply.started":"2024-12-14T16:40:05.10823Z","shell.execute_reply":"2024-12-14T16:40:05.749771Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"columns_with_missing = ['Previous Claims', 'Occupation', 'Credit Score', \n                        'Number of Dependents', 'Customer Feedback',\n                        'Health Score', 'Annual Income', 'Age', \n                        'Marital Status', 'Vehicle Age', 'Insurance Duration'\n                       ]\n\n#for col in columns_with_missing:\n#    # Create a new binary column indicating missingness\n#    train[col+'_missing'] = train[col].isnull().astype(int)\n#    test[col+'_missing'] = test[col].isnull().astype(int)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:05.751718Z","iopub.execute_input":"2024-12-14T16:40:05.752017Z","iopub.status.idle":"2024-12-14T16:40:06.112533Z","shell.execute_reply.started":"2024-12-14T16:40:05.751988Z","shell.execute_reply":"2024-12-14T16:40:06.111544Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# We still need to provide some values for the model (XGBoost can't handle NaNs)\n# but we can now use a simple fill (median) since we have flags\n#float_columns_with_missing = ['Age']\n#                              'Annual Income', 'Number of Dependents', \n#                              'Health Score', 'Previous Claims', \n#                              'Vehicle Age', 'Credit Score', 'Insurance Duration'\n#                             ]\n\n#for col in float_columns_with_missing:\n#    median_val = train[col].median()\n#    train[col].fillna(median_val, inplace=True)\n#    test[col].fillna(median_val, inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:06.113785Z","iopub.execute_input":"2024-12-14T16:40:06.114139Z","iopub.status.idle":"2024-12-14T16:40:06.11874Z","shell.execute_reply.started":"2024-12-14T16:40:06.114105Z","shell.execute_reply":"2024-12-14T16:40:06.117685Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(train.isnull().sum())           ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:06.120448Z","iopub.execute_input":"2024-12-14T16:40:06.120813Z","iopub.status.idle":"2024-12-14T16:40:06.718703Z","shell.execute_reply.started":"2024-12-14T16:40:06.12074Z","shell.execute_reply":"2024-12-14T16:40:06.717486Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Address skewed features: 'Annual Income', 'Premium Amount', 'Health Score' may be skewed\n# Apply log transform to reduce skewness if needed (only to non-negative features)\n# We will be careful with transforming the target for RMSLE evaluation.\n\n#train['Annual Income'] = np.log1p(train['Annual Income'])\n#test['Annual Income'] = np.log1p(test['Annual Income'])\n\n#train['Health Score'] = np.log1p(train['Health Score'])\n#test['Health Score'] = np.log1p(test['Health Score'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:06.720309Z","iopub.execute_input":"2024-12-14T16:40:06.720606Z","iopub.status.idle":"2024-12-14T16:40:06.818923Z","shell.execute_reply.started":"2024-12-14T16:40:06.720577Z","shell.execute_reply":"2024-12-14T16:40:06.817846Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Consider transforming the target:\n# RMSLE is usually applied to positive targets. We can still predict on normal scale, \n# but optimizing a model on a log-transformed target often helps.\n# Handle target outliers and transform target (#1)\nupper_bound = train[TARGET].quantile(0.99)\ntrain[TARGET] = np.where(train[TARGET] > upper_bound, upper_bound, train[TARGET])\ntrain[TARGET] = np.log1p(train[TARGET])  # log-transform the target for modeling","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:06.820334Z","iopub.execute_input":"2024-12-14T16:40:06.82075Z","iopub.status.idle":"2024-12-14T16:40:06.879441Z","shell.execute_reply.started":"2024-12-14T16:40:06.820703Z","shell.execute_reply":"2024-12-14T16:40:06.878325Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ========================================\n# 5. Feature Engineering\n# ========================================\n\n# Example feature engineering:\n# - Extract date features from 'Policy Start Date'\ntrain['Policy_Year'] = train['Policy Start Date'].dt.year\ntrain['Policy_Month'] = train['Policy Start Date'].dt.month\ntrain['Policy_Day'] = train['Policy Start Date'].dt.day\n\ntest['Policy_Year'] = test['Policy Start Date'].dt.year\ntest['Policy_Month'] = test['Policy Start Date'].dt.month\ntest['Policy_Day'] = test['Policy Start Date'].dt.day","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:06.880899Z","iopub.execute_input":"2024-12-14T16:40:06.88135Z","iopub.status.idle":"2024-12-14T16:40:07.182965Z","shell.execute_reply.started":"2024-12-14T16:40:06.881292Z","shell.execute_reply":"2024-12-14T16:40:07.181953Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Drop original date column if it's no longer needed\ntrain.drop(['Policy Start Date'], axis=1, inplace=True)\ntest.drop(['Policy Start Date'], axis=1, inplace=True)\n\n# Text feature: 'Customer Feedback'\n# For example, we could do a length count or number of words as a simple feature\ntrain['Feedback_Length'] = train['Customer Feedback'].astype(str).apply(lambda x: len(x))\ntest['Feedback_Length'] = test['Customer Feedback'].astype(str).apply(lambda x: len(x))\n\ntrain['Feedback_WordCount'] = train['Customer Feedback'].astype(str).apply(lambda x: len(x.split()))\ntest['Feedback_WordCount'] = test['Customer Feedback'].astype(str).apply(lambda x: len(x.split()))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:07.184568Z","iopub.execute_input":"2024-12-14T16:40:07.185038Z","iopub.status.idle":"2024-12-14T16:40:09.363549Z","shell.execute_reply.started":"2024-12-14T16:40:07.184992Z","shell.execute_reply":"2024-12-14T16:40:09.362413Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# TF-IDF on Customer Feedback (#4)\n#train['Customer Feedback'] = train['Customer Feedback'].fillna('')\n#test['Customer Feedback'] = test['Customer Feedback'].fillna('')\n\n#tfidf = TfidfVectorizer(max_features=10, ngram_range=(1,2), stop_words='english')\n#train_tfidf = tfidf.fit_transform(train['Customer Feedback'])\n#test_tfidf = tfidf.transform(test['Customer Feedback'])\n\n#tfidf_cols = [f\"tfidf_{i}\" for i in range(train_tfidf.shape[1])]\n#train_tfidf_df = pd.DataFrame(train_tfidf.toarray(), columns=tfidf_cols, index=train.index)\n#test_tfidf_df = pd.DataFrame(test_tfidf.toarray(), columns=tfidf_cols, index=test.index)\n\n#train = pd.concat([train, train_tfidf_df], axis=1)\n#test = pd.concat([test, test_tfidf_df], axis=1)\n\n# Drop the raw text if we prefer not to use it directly\ntrain.drop('Customer Feedback', axis=1, inplace=True)\ntest.drop('Customer Feedback', axis=1, inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:09.364998Z","iopub.execute_input":"2024-12-14T16:40:09.365357Z","iopub.status.idle":"2024-12-14T16:40:18.148462Z","shell.execute_reply.started":"2024-12-14T16:40:09.365325Z","shell.execute_reply":"2024-12-14T16:40:18.147369Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# For categorical variables, we will use one-hot or label encoding\n# Let's use one-hot encoding for simplicity\nall_data = pd.concat([train, test], sort=False)\nall_data = pd.get_dummies(all_data, columns=cat_features, drop_first=True)\n\n# Now split back into train and test\ntrain = all_data.iloc[:train.shape[0], :]\ntest = all_data.iloc[train.shape[0]:, :]","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create interaction features (#3)\n# For example: 'Annual Income' and 'Health Score'\n#train['Income_Health_Interaction'] = train['Annual Income'] * train['Health Score']\n#test['Income_Health_Interaction'] = test['Annual Income'] * test['Health Score']\n\n# Another interaction: 'Age' and 'Vehicle Age'\n#train['Age_VehicleAge_Interaction'] = train['Age'] * train['Vehicle Age']\n#test['Age_VehicleAge_Interaction'] = test['Age'] * test['Vehicle Age']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:18.149703Z","iopub.execute_input":"2024-12-14T16:40:18.150066Z","iopub.status.idle":"2024-12-14T16:40:18.173698Z","shell.execute_reply.started":"2024-12-14T16:40:18.150033Z","shell.execute_reply":"2024-12-14T16:40:18.172594Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Another example: Binning 'Age' into categories (Young, Middle-aged, Senior)\n#train['Age_Bin'] = pd.cut(\n#    train['Age'],\n#    bins=[0, 30, 60, train['Age'].max()],\n#    labels=['Young', 'Middle_aged', 'Senior']\n#)\n\n#test['Age_Bin'] = pd.cut(\n#    test['Age'],\n#    bins=[0, 30, 60, test['Age'].max()],\n#    labels=['Young', 'Middle_aged', 'Senior']\n#)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:18.175414Z","iopub.execute_input":"2024-12-14T16:40:18.1759Z","iopub.status.idle":"2024-12-14T16:40:18.241291Z","shell.execute_reply.started":"2024-12-14T16:40:18.175828Z","shell.execute_reply":"2024-12-14T16:40:18.240147Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Handle categorical features\n#te_cat_features = ['Marital Status','Education Level','Occupation','Location',\n#                   'Policy Type','Exercise Frequency','Property Type'\n#                  ]\n\n#non_te_cat_features = ['Gender', 'Smoking Status']\n\n# Example of target encoding for one categorical variable, say 'Occupation'\n# We will do target encoding on the training set and map to test\n# (after applying transformation on the target already)\n#for col in te_cat_features:\n#    _means = train.groupby(col)[TARGET].mean().to_dict()\n#    train[col + '_TE'] = train[col].map(_means)\n#    test[col + '_TE'] = test[col].map(_means)\n\n#    train.drop(col, axis=1, inplace=True)\n#    test.drop(col, axis=1, inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:18.242575Z","iopub.execute_input":"2024-12-14T16:40:18.242945Z","iopub.status.idle":"2024-12-14T16:40:21.708371Z","shell.execute_reply.started":"2024-12-14T16:40:18.242903Z","shell.execute_reply":"2024-12-14T16:40:21.707237Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# For categorical variables, we will use one-hot or label encoding\n# Let's use one-hot encoding for simplicity\n#all_data = pd.concat([train, test], sort=False)\n#all_data = pd.get_dummies(all_data, columns=non_te_cat_features, drop_first=True)\n\n# Now split back into train and test\n#train = all_data.iloc[:train.shape[0], :]\n#test = all_data.iloc[train.shape[0]:, :]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:21.715293Z","iopub.execute_input":"2024-12-14T16:40:21.715643Z","iopub.status.idle":"2024-12-14T16:40:23.562846Z","shell.execute_reply.started":"2024-12-14T16:40:21.71561Z","shell.execute_reply":"2024-12-14T16:40:23.561839Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:23.564026Z","iopub.execute_input":"2024-12-14T16:40:23.56435Z","iopub.status.idle":"2024-12-14T16:40:23.571097Z","shell.execute_reply.started":"2024-12-14T16:40:23.564319Z","shell.execute_reply":"2024-12-14T16:40:23.569952Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:23.572208Z","iopub.execute_input":"2024-12-14T16:40:23.572501Z","iopub.status.idle":"2024-12-14T16:40:23.587158Z","shell.execute_reply.started":"2024-12-14T16:40:23.572473Z","shell.execute_reply":"2024-12-14T16:40:23.585991Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:23.588592Z","iopub.execute_input":"2024-12-14T16:40:23.589103Z","iopub.status.idle":"2024-12-14T16:40:23.673334Z","shell.execute_reply.started":"2024-12-14T16:40:23.589044Z","shell.execute_reply":"2024-12-14T16:40:23.672304Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Make sure target is still in train\n# We moved our target earlier, so it's safe. Just reconfirm structure\ny = train[TARGET]\nX = train.drop([TARGET, ID_COL], axis=1)\nX_test = test.drop([TARGET, ID_COL], axis=1, errors='ignore') # test doesn't have target\n\nprint(\"Final training shape:\", X.shape)\nprint(\"Final test shape:\", X_test.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:23.674728Z","iopub.execute_input":"2024-12-14T16:40:23.675071Z","iopub.status.idle":"2024-12-14T16:40:23.814116Z","shell.execute_reply.started":"2024-12-14T16:40:23.675039Z","shell.execute_reply":"2024-12-14T16:40:23.812954Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ========================================\n# 6. Train-Validation Split (Local)\n# ========================================\nX_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:23.815376Z","iopub.execute_input":"2024-12-14T16:40:23.815666Z","iopub.status.idle":"2024-12-14T16:40:24.342669Z","shell.execute_reply.started":"2024-12-14T16:40:23.815639Z","shell.execute_reply":"2024-12-14T16:40:24.34168Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ========================================\n# 7. Baseline Modeling with XGBoost\n# ========================================\n\ndef rmsle(y_true, y_pred):\n    return np.sqrt(mean_squared_log_error(y_true, y_pred))\n\n# Preliminary model\nxgb_reg = xgb.XGBRegressor(random_state=42,enable_categorical=True)\nxgb_reg.fit(X_train, y_train)\n\ny_pred_val = xgb_reg.predict(X_val)\nval_score = rmsle(np.expm1(y_val), np.expm1(y_pred_val))  # transform back the exponent\nprint(\"Baseline RMSLE on validation:\", val_score)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:24.344426Z","iopub.execute_input":"2024-12-14T16:40:24.344836Z","iopub.status.idle":"2024-12-14T16:40:38.453818Z","shell.execute_reply.started":"2024-12-14T16:40:24.344789Z","shell.execute_reply":"2024-12-14T16:40:38.452481Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ========================================\n# 8. Bayesian Optimization with Optuna\n# ========================================\ndef objective(trial):\n    param = {\n        'verbosity': 0,\n        'objective': 'reg:squarederror',\n        'random_state': 42,\n        'tree_method': 'hist',\n        'enable_categorical': True,\n        'n_estimators': trial.suggest_int('n_estimators', 1000, 3000),\n        'learning_rate': trial.suggest_float('learning_rate', 0.01, 0.1),\n        'max_depth': trial.suggest_int('max_depth', 3, 12),\n        'subsample': trial.suggest_float('subsample', 0.5, 1.0),\n        'colsample_bytree': trial.suggest_float('colsample_bytree', 0.5, 1.0),\n        'reg_alpha': trial.suggest_float('reg_alpha', 0, 5),\n        'reg_lambda': trial.suggest_float('reg_lambda', 0, 5)\n    }\n    \n    # Using K-Fold for evaluation\n    kf = KFold(n_splits=5, shuffle=True, random_state=42)\n    rmsle_scores = []\n    for train_idx, val_idx in kf.split(X, y):\n        X_tr, X_vl = X.iloc[train_idx], X.iloc[val_idx]\n        y_tr, y_vl = y.iloc[train_idx], y.iloc[val_idx]\n        \n        model = xgb.XGBRegressor(**param)\n        model.fit(X_tr, y_tr, eval_set=[(X_vl, y_vl)], early_stopping_rounds=50, verbose=False)\n        preds = model.predict(X_vl)\n        fold_score = rmsle(np.expm1(y_vl), np.expm1(preds))\n        rmsle_scores.append(fold_score)\n    \n    return np.mean(rmsle_scores)\n\nstudy = optuna.create_study(direction='minimize')\nstudy.optimize(objective, n_trials=50, show_progress_bar=True)\n\nprint(\"Best RMSLE:\", study.best_value)\nprint(\"Best parameters:\", study.best_params)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:38.455433Z","iopub.execute_input":"2024-12-14T16:40:38.455772Z","iopub.status.idle":"2024-12-14T16:40:38.465401Z","shell.execute_reply.started":"2024-12-14T16:40:38.455739Z","shell.execute_reply":"2024-12-14T16:40:38.464206Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Train final model with best params\nbest_params = study.best_params\n\n#best_params = {\n#    'n_estimators': 2000, \n#    'learning_rate': 0.04244891943130667, \n#    'max_depth': 8, \n#    'subsample': 0.9193559662883617, \n#    'colsample_bytree': 0.9701872660117554, \n#    'reg_alpha': 1.2835369666881167, \n#    'reg_lambda': 6.170697645520114,\n#    'tree_method': 'hist',\n#    'enable_categorical': True\n#}\n\nfinal_model = xgb.XGBRegressor(**best_params, \n                               random_state=42)\nfinal_model.fit(X, y)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:40:38.466793Z","iopub.execute_input":"2024-12-14T16:40:38.467265Z","iopub.status.idle":"2024-12-14T16:43:40.551306Z","shell.execute_reply.started":"2024-12-14T16:40:38.467229Z","shell.execute_reply":"2024-12-14T16:43:40.5501Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ========================================\n# 9. Model Evaluation using RMSLE on Validation\n# ========================================\n# Already done cross-validation; \n# If we have a separate test set (we do - but no target), we will just generate predictions:\n# We'll trust cross-validation results. For final submission on Kaggle:\ny_pred_test = final_model.predict(X_test)\n# Remember to invert the log transform for predictions\ny_pred_test = np.expm1(y_pred_test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:43:40.553026Z","iopub.execute_input":"2024-12-14T16:43:40.553467Z","iopub.status.idle":"2024-12-14T16:44:09.867554Z","shell.execute_reply.started":"2024-12-14T16:43:40.553421Z","shell.execute_reply":"2024-12-14T16:44:09.866245Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Save predictions\nsubmission = pd.DataFrame({ID_COL: test[ID_COL], TARGET: y_pred_test})\nsubmission.to_csv(\"submission.csv\", index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:44:09.869069Z","iopub.execute_input":"2024-12-14T16:44:09.869497Z","iopub.status.idle":"2024-12-14T16:44:11.087023Z","shell.execute_reply.started":"2024-12-14T16:44:09.869451Z","shell.execute_reply":"2024-12-14T16:44:11.085944Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ========================================\n# 10. Explainability with SHAP\n# ========================================\n# Use a subset of data to speed up SHAP\nexplainer = shap.TreeExplainer(final_model)\nshap_values = explainer.shap_values(X.sample(1000, random_state=42))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:44:11.088424Z","iopub.execute_input":"2024-12-14T16:44:11.08878Z","iopub.status.idle":"2024-12-14T16:45:15.597713Z","shell.execute_reply.started":"2024-12-14T16:44:11.088747Z","shell.execute_reply":"2024-12-14T16:45:15.596716Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Global Feature Importance\nshap.summary_plot(shap_values, X.sample(1000, random_state=42), plot_type='bar')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:45:15.598738Z","iopub.execute_input":"2024-12-14T16:45:15.599076Z","iopub.status.idle":"2024-12-14T16:45:15.941298Z","shell.execute_reply.started":"2024-12-14T16:45:15.599037Z","shell.execute_reply":"2024-12-14T16:45:15.939936Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Detailed SHAP summary plot\nshap.summary_plot(shap_values, X.sample(1000, random_state=42))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:45:15.943184Z","iopub.execute_input":"2024-12-14T16:45:15.943652Z","iopub.status.idle":"2024-12-14T16:45:16.755337Z","shell.execute_reply.started":"2024-12-14T16:45:15.943584Z","shell.execute_reply":"2024-12-14T16:45:16.754251Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Local explanation for a single instance\ninstance_idx = 10\n\nshap.force_plot(explainer.expected_value, shap_values[instance_idx,:], \n                X.iloc[instance_idx,:])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:45:16.756765Z","iopub.execute_input":"2024-12-14T16:45:16.757217Z","iopub.status.idle":"2024-12-14T16:45:16.767754Z","shell.execute_reply.started":"2024-12-14T16:45:16.75717Z","shell.execute_reply":"2024-12-14T16:45:16.766615Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ========================================\n# 11. Conclusion\n# ========================================\n\n# - We performed extensive EDA, cleaning, and feature engineering.\n# - We log-transformed skewed features and target to handle skewness.\n# - We used Bayesian optimization to fine-tune XGBoost parameters.\n# - We evaluated the model using RMSLE.\n# - We used SHAP to gain insights into feature importance both globally and locally.\n\n# The final submission.csv file contains our predictions on the test dataset.","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-14T16:45:16.769103Z","iopub.execute_input":"2024-12-14T16:45:16.76942Z","iopub.status.idle":"2024-12-14T16:45:16.780137Z","shell.execute_reply.started":"2024-12-14T16:45:16.769389Z","shell.execute_reply":"2024-12-14T16:45:16.77884Z"}},"outputs":[],"execution_count":null}]}