{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Quick Intro\n\nHi, this is my first public notebook, so all constructive criticism is highly appreciated.\n\nCurrently it is very bare-bones, meant to be used as a starting benchmark point and will be progressively enriched in the coming days. Wanted to make it with simplicity and efficiency in mind, a set of helper functions, to minimize the fluff and repetitive boilerplate code during the actual data exploration and training.\n\nIt is CPU friendly, i.e not using grid searches, or other computationally intensive automated approaches.","metadata":{}},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"# Core imports\nimport pandas as pd\nimport numpy as np\nimport re\n\nfrom fastai.imports import *\n\n# Data Visualization\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\nsns.set(style='darkgrid', font_scale=1.2)\npd.set_option('display.max_columns', None)\n# pd.set_option('display.max_rows', None)\n\n# Models\nfrom xgboost import XGBRegressor\nfrom lightgbm import LGBMRegressor\nfrom catboost import CatBoostRegressor\n\n# Preprocessing\nfrom sklearn.impute import SimpleImputer\n\n# Metrics\nfrom sklearn.metrics import mean_squared_log_error\n\n# Model Selection\nfrom sklearn.model_selection import KFold\n\n# Seed\nRANDOM_STATE = 42\nnp.random.seed(RANDOM_STATE)\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T14:34:59.486254Z","iopub.execute_input":"2024-12-02T14:34:59.486643Z","iopub.status.idle":"2024-12-02T14:35:02.785898Z","shell.execute_reply.started":"2024-12-02T14:34:59.486605Z","shell.execute_reply":"2024-12-02T14:35:02.784761Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Utility Functions\n\n* **merge_train_and_test_to_X** - combines train and test data into a single DataFrame for less boilerplate during data preprocessing\n\n * **split_X_back_to_train_and_test** - does the reverse of 'merge_train_and_test_to_X'. Split the DataFrame back to train and test before training\n\n * **map_column** - maps a categorical column to an int, given a predefined map. It's assumed all NaN values are imputed\n\n * **impute_missing_values** - imputes missing values in a column by a selected strategy. Generally use 'mean' strategy for high-cardinality features and 'most-frequent' for most_frequent features\n\n * **extract_columns_from_date** - split timestamp column into partial columns, such as 'year', 'month', 'day-of-month', 'day-of-week', etc.\n\n * **plot_kfold_learning_curves** - plot each training fold's (train, eval) metrics, to monitor for overfitting. If you are going to use it for LightGBM & CatBoost, perhaps you'have to play around with the metrics.\n\n * **select_most_important_features** - passes all engineered features through a very basic model ('xgb', 'lgbm', 'catboost') and prints ranking of feature importances for the particular model. Also allows to cap N top performing features. I have taken the idea from this notebook, so props to the author: https://www.kaggle.com/code/arunklenin/ps4e11-mental-health-prediction-classification\n\n * **kfold_mean_prediction** - run k fold predictions for the final submission. The reason is, you can use early-stopping parameter when you have the train/val splits. Then select the mean of all folds as a final answer. My observations, lead me to believe it improves performance slightly, yet to confirm if it's true.\n\n * **kfold_cross_val_regressor** - Kfold cross validation for training","metadata":{}},{"cell_type":"code","source":"# Data Preprocessing \ndef merge_train_and_test_to_X(train, test):\n    df = pd.concat([train, test])\n    return df\n\ndef split_X_back_to_train_and_test(df, train, test):\n    n_train, n_test = train.shape[0], test.shape[0]\n    test_chunk = df.iloc[n_train:, :]\n    train_chunk = df.iloc[:n_train, :]\n    return train_chunk, test_chunk\n\ndef map_column(df, col, map):\n    df[col] = df[col].map(map)\n    return df\n\ndef impute_missing_values(df, columns, strategy):\n    imputer = SimpleImputer(strategy=strategy)\n    df[columns] = imputer.fit_transform(df[columns])\n    return df\n\ndef get_unique_values_for_object_columns(df):\n    for col in df.select_dtypes('object').columns:\n        print(f'{col} - {df[col].unique()}')\n\ndef extract_columns_from_date(df, column):\n\n    def parse_date_to_columns(x):\n        date_format_pattern = r\"(\\d{4})-(\\d{2})-(\\d{2}) (\\d{2}:\\d{2})(:\\d{2})(\\.\\d+)\" \n        match_groups = re.match(date_format_pattern, x).groups()\n        return match_groups[0], match_groups[1], match_groups[2], match_groups[3].replace(':', '')\n\n    parsed_dates = df[column].apply(parse_date_to_columns)\n    df[f'{column}_year'] = parsed_dates.apply(lambda x: x[0]).astype(int)\n    df[f'{column}_month'] = parsed_dates.apply(lambda x: x[1]).astype(int)\n    df[f'{column}_day_of_month'] = parsed_dates.apply(lambda x: x[2]).astype(int)\n    df[f'{column}_time'] = parsed_dates.apply(lambda x: x[3]).astype(int)\n    df[f'{column}_day_of_week'] = pd.to_datetime(parsed_dates.apply(lambda x: (x[0] + '-' + x[1] + '-' + x[2]))).dt.dayofweek.astype(int)\n\n    return df\n\n# Training Visualization\n\ndef plot_kfold_learning_curves(eval_results, model_type):\n    folds = len(eval_results)\n\n    key1, key2 = None, None\n    metric = None\n    match model_type:\n        case 'xgb':\n            key1, key2 = 'validation_0', 'validation_1'\n            metric = 'rmsle'\n        case 'lgbm':\n            key1, key2 = 'training', 'valid_1'\n            metric = 'gamma'\n        case 'cat':\n            key1, key2 = 'validation_0', 'validation_1'\n            metric = 'RMSE'\n\n    fig = plt.figure(figsize=(30, 5*((folds+10)//2)))\n\n    for i, results in enumerate(eval_results):\n        plt.subplot((folds+1)//2, 2, i+1)\n        # plot learning curves\n        plt.title(f'Fold {i}')\n        plt.plot(results[key1][metric], label='train')\n        plt.plot(results[key2][metric], label='val')\n        # show the legend\n        plt.legend()\n\n# Feature Selection\n\ndef select_most_important_features(features, target, n, model_type, seed, n_splits=5):\n    \n    lgbm_params = {\n        'random_state': RANDOM_STATE,\n        'verbose': -1,\n    }\n\n    xgb_params = {\n        'random_state': RANDOM_STATE,\n        'verbosity': 0,\n        'objective': 'reg:squaredlogerror',\n        'n_estimators': 100\n    }\n    \n    cat_params = {\n        'random_state': RANDOM_STATE,\n        'verbose': 0,\n        'iterations': 100,\n    }\n\n    match model_type:\n        case 'lgbm':\n            model = LGBMRegressor(**lgbm_params)\n        case 'xgb':\n            model = XGBRegressor(**xgb_params)\n        case 'cat':\n            model = CatBoostRegressor(**cat_params)\n\n    kf = KFold(n_splits=n_splits, shuffle=True, random_state=RANDOM_STATE)\n    kfold_errors = []\n    kfold_feature_importance_list = []\n\n    for train_index, val_index in kf.split(features, target):\n        train_x, val_x = features.iloc[train_index], features.iloc[val_index]\n        train_y, val_y = target.iloc[train_index], target.iloc[val_index]\n\n        model.fit(train_x, train_y)\n\n        pred_y = model.predict(val_x)\n        \n        kfold_errors.append(np.sqrt(mean_squared_log_error(val_y, pred_y)))\n        kfold_feature_importance_list.append(model.feature_importances_)\n\n    avg_feature_importances = np.mean(kfold_feature_importance_list, axis=0)\n\n    feature_importance_list = list(zip(features, avg_feature_importances))\n    sorted_features = sorted(feature_importance_list, key=lambda x: x[1], reverse=True)\n    top_n_features = [feat[0] for feat in sorted_features[:n]]\n\n    # display_features = top_n_features[:n_display_features]\n    for i, feat in enumerate(top_n_features):\n        print(f'{i}: {feat} with feature importance score of {sorted_features[i]}')\n\n    return top_n_features, avg_feature_importances\n\n# Validation and Prediction\n\ndef kfold_mean_prediction(model, X_train, Y_train, X_test, seed, n_splits=5): \n    skf = KFold(n_splits=n_splits, shuffle=True, random_state=seed)\n    preds = np.zeros((X_test.shape[0], n_splits))\n\n    for fold, (train_index, val_index) in enumerate(skf.split(X_train, Y_train)):\n        train_x, val_x = X_train.iloc[train_index], X_train.iloc[val_index]\n        train_y, val_y = Y_train.iloc[train_index], Y_train.iloc[val_index]\n        \n        model.fit(train_x, train_y, eval_set=[(train_x, train_y), (val_x, val_y)])\n        \n        preds[:, fold] = model.predict(X_test)\n        print(f\"fold {fold+1}/{n_splits} completed.\")\n\n    return np.mean(preds, axis=1)\n\ndef kfold_cross_val_regressor(model, features, target, seed, n_splits=5): \n    skf = KFold(n_splits=n_splits, shuffle=True, random_state=seed)\n\n    evals_results = []\n    errors = []\n    i = 0\n\n    for train_index, val_index in skf.split(features, target):\n        train_x, val_x = features.iloc[train_index], features.iloc[val_index]\n        train_y, val_y = target.iloc[train_index], target.iloc[val_index]\n        try:\n            model.fit(train_x, train_y, eval_set=[(train_x, train_y), (val_x, val_y)])\n            if hasattr(model, 'evals_result'):\n                evals_results.append(model.evals_result())\n            elif hasattr(model, '_evals_result'):\n                evals_results.append(model._evals_result)\n            elif hasattr(model, 'evals_result_'):\n                evals_results.append(model.evals_result_)\n        except: \n            model.fit(train_x, train_y)\n        preds = model.predict(val_x)\n\n        err = np.sqrt(mean_squared_log_error(val_y, preds))\n        errors.append(err)\n        i+=1\n        print(f'fold {i}: error: {err:.8f}')\n\n    average_error = np.mean(errors)\n    std_error = np.std(errors, dtype=np.float64)\n    print(f'Average RMSLE: {average_error:.8f} with STD: {std_error:.4f}') \n    return average_error, std_error, evals_results\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T14:35:27.875273Z","iopub.execute_input":"2024-12-02T14:35:27.876022Z","iopub.status.idle":"2024-12-02T14:35:27.907091Z","shell.execute_reply.started":"2024-12-02T14:35:27.875967Z","shell.execute_reply":"2024-12-02T14:35:27.905639Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Dictionaries & Mappings","metadata":{}},{"cell_type":"code","source":"gender_map = {\n    'Female': 1,\n    'Male': 0\n}\n\nmarital_status_map = {\n    'Married': 2,\n    'Divorced': 1,\n    'Single': 0,\n    'Other': -1\n}\n\neducation_level_map = {\n    'PhD': 3,\n    'Master\\'s': 2,\n    'Bachelor\\'s': 1,\n    'High School': 0\n}\n\noccupation_map = {\n    'Self-Employed': 2,\n    'Employed': 1,\n    'Unemployed': 0\n}\n\nlocation_map = {\n    'Urban': 2,\n    'Suburban': 1,\n    'Rural': 0\n}\n\npolicy_type_map = {\n    'Premium': 2,\n    'Comprehensive': 1,\n    'Basic': 0\n}\n\ncustomer_feedback_map = {\n    'Good': 2,\n    'Average': 1,\n    'Poor': 0,\n    'Other': -1\n}\n\nyes_no_map = {\n    'Yes': 1,\n    'No': 0\n}\n\nexercise_frequency_map = {\n    'Daily': 3,\n    'Weekly': 2,\n    'Monthly': 1,\n    'Rarely': 0\n}\n\nproperty_type = {\n    'House': 2,\n    'Condo': 1,\n    'Apartment': 0\n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T14:35:51.371333Z","iopub.execute_input":"2024-12-02T14:35:51.371856Z","iopub.status.idle":"2024-12-02T14:35:51.380184Z","shell.execute_reply.started":"2024-12-02T14:35:51.371807Z","shell.execute_reply":"2024-12-02T14:35:51.378917Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Data Load, Visualization & Preparation","metadata":{}},{"cell_type":"code","source":"target_name = 'Premium Amount'\nX_train = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\nX_test = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')\nY_train = X_train[target_name]\nX_train = X_train.iloc[:, :-1]\n\nX = merge_train_and_test_to_X(X_train, X_test)\n\n# Impute missing values\ncolumns_w_missing_values = X.columns[X.isnull().any()]\nhigh_cardinality_cols = ['Age', 'Annual Income', 'Number of Dependents', 'Health Score', 'Credit Score']\nlow_cardinality_cols = ['Marital Status', 'Occupation', 'Previous Claims', 'Vehicle Age', 'Insurance Duration']\n\nX = impute_missing_values(X, high_cardinality_cols, 'mean')\nX = impute_missing_values(X, low_cardinality_cols, 'most_frequent')\n\n# Columns that are arguably not applicable to each individual\nX['Customer Feedback'] = X['Customer Feedback'].fillna('Other')\n\n# Encode Categorical columns with low cardinality\nX = map_column(X, 'Gender', gender_map)\nX = map_column(X, 'Marital Status', marital_status_map)\nX = map_column(X, 'Education Level', education_level_map)\nX = map_column(X, 'Occupation', occupation_map)\nX = map_column(X, 'Location', location_map)\nX = map_column(X, 'Policy Type', policy_type_map)\nX = map_column(X, 'Customer Feedback', customer_feedback_map)\nX = map_column(X, 'Smoking Status', yes_no_map)\nX = map_column(X, 'Exercise Frequency', exercise_frequency_map)\nX = map_column(X, 'Property Type', property_type)\n\n# Split timestamp to multiple columns to extract maximum signal ( hopefully ;) )\nX = extract_columns_from_date(X, 'Policy Start Date')\n\n# Convert any remaining cols of type 'object' to int or float\nobject_to_int_columns = ['Vehicle Age', 'Previous Claims', 'Insurance Duration']\nX[object_to_int_columns] = X[object_to_int_columns].astype(int)\n\ncolumns_to_drop = ['id', 'Policy Start Date']\nX = X.drop(columns_to_drop, axis=1).reset_index(drop=True)\n\n\nX_train, X_test = split_X_back_to_train_and_test(X, X_train, X_test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T14:38:40.229908Z","iopub.execute_input":"2024-12-02T14:38:40.230333Z","iopub.status.idle":"2024-12-02T14:39:03.635258Z","shell.execute_reply.started":"2024-12-02T14:38:40.230298Z","shell.execute_reply":"2024-12-02T14:39:03.634154Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Feature Selection","metadata":{}},{"cell_type":"code","source":"n_top_features = 30\n\nprint('=== LGBM Regressor feature importance')\ntop_feats_lgbm, avg_feat_importances_lgbm = select_most_important_features(X_train, Y_train, n_top_features, 'lgbm', RANDOM_STATE, n_splits=5)\nprint('=== XGB Regressor feature importance')\ntop_feats_xgb, avg_feat_importances_xgb = select_most_important_features(X_train, Y_train, n_top_features, 'xgb', RANDOM_STATE, n_splits=5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T14:39:23.245061Z","iopub.execute_input":"2024-12-02T14:39:23.245481Z","iopub.status.idle":"2024-12-02T14:41:34.687134Z","shell.execute_reply.started":"2024-12-02T14:39:23.245449Z","shell.execute_reply":"2024-12-02T14:41:34.686023Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Models","metadata":{}},{"cell_type":"markdown","source":"## XGB Regressor","metadata":{}},{"cell_type":"code","source":"xgb_params = {\n    'verbosity': 0,\n    'random_state': RANDOM_STATE,\n\n    'n_estimators': 30,\n    'subsample': 0.3,\n    'learning_rate': 0.5,\n\n    'tree_method': 'hist',\n\n    'objective': 'reg:gamma',\n    'metric': 'rmsle',\n    'eval_metric': 'rmsle'\n}\n\nxgb = XGBRegressor(**xgb_params, early_stopping_rounds=50)\n\n_, _, eval_results = kfold_cross_val_regressor(xgb, X_train, Y_train, 42)\n\nplot_kfold_learning_curves(eval_results, 'xgb')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Final Submission","metadata":{}},{"cell_type":"code","source":"preds = kfold_mean_prediction(xgb, X_train, Y_train, X_test, RANDOM_STATE, 7)\nsub = pd.read_csv('/kaggle/input/playground-series-s4e12/sample_submission.csv')\n\nsub[\"Premium Amount\"] = preds\nsub.to_csv(\"submission.csv\", index=False)","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}