{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"},{"sourceId":211205059,"sourceType":"kernelVersion"}],"dockerImageVersionId":30823,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This Python script defines a class `Andro` to automate the process of data preparation, feature engineering, training, and prediction for a machine learning pipeline. Here’s a detailed step-by-step explanation of the code:\n\n---\n\n# 1. **Setup and Initialization**\n- **Imports**: Libraries like `pandas`, `numpy`, `CatBoostRegressor` (a gradient boosting model), and other utilities for data processing, evaluation, and model persistence are imported.\n- **Warnings**: Suppresses warnings for cleaner output.\n- **Class `Andro`**: Encapsulates the entire process within a class to make it modular and reusable.\n\n# 2. **Constructor (`__init__` method)**\n- The constructor takes four inputs:\n  - `train_path`: Path to the training dataset.\n  - `test_path`: Path to the test dataset.\n  - `sample_path`: Path to the sample submission file.\n  - `model_path`: Path to a saved model or feature file (`.pkl` file).\n- These paths are saved as instance variables.\n\n---\n\n# 3. **Loading Data (`load_data` method)**\n- Reads the training (`train.csv`), test (`test.csv`), and sample submission (`sample_submission.csv`) files into pandas DataFrames.\n- Drops the `id` column since it’s not useful for modeling.\n- Loads pre-computed non-logarithmically transformed features (`nonlog_fe`) and values (`nonlog`) from the `.pkl` file and adds them as columns in `train` and `test`.\n\n---\n\n# 4. **Date Features (`date_features` method)**\n- Extracts temporal features from the `Policy Start Date` column, such as:\n  - **Year**, **Month**, **Day**, **Week**, **Day of the week**, etc.\n  - **Sinusoidal transformations** for cyclic features (e.g., `Month_sin`, `Month_cos`) to handle periodicity.\n  - **Grouping** to identify policies in the same month/year.\n- Removes the original `Policy Start Date` column after extracting features.\n\n---\n\n# 5. **Feature Engineering (`feature_engineering` method)**\n- Creates a new feature `contract length`:\n  - Categorizes the `Insurance Duration` into bins and assigns labels (e.g., 0 for short contracts, 1 for medium, 2 for long).\n- Handles missing values in `Insurance Duration` by replacing them with 99.\n\n---\n\n# 6. **Frequency Encoding (`frequency_encode` method)**\n- Encodes categorical variables using their frequency in the dataset:\n  - Combines `train` and `test` for consistency in encoding.\n  - For each categorical column, maps the frequency of each category into a new feature (e.g., `col_freq`).\n  - Adds new encoded columns and optionally removes the original categorical columns.\n\n---\n\n# 7. **Preprocessing (`preprocess` method)**\n- Calls:\n  - `date_features` to extract temporal features.\n  - `feature_engineering` to create the `contract length` feature.\n  - `frequency_encode` to encode categorical variables.\n- Updates `train` and `test` datasets to only include engineered features for consistency.\n\n---\n\n# 8. **Root Mean Squared Logarithmic Error (`rmsle` method)**\n- Calculates the RMSLE between actual and predicted values.\n  - Converts predicted values back from log-transformed space before evaluation.\n\n---\n\n# 9. **Model Training (`train_model` method)**\n- Prepares features (`X`) and target (`y`) from the training dataset.\n- Applies a 5-fold cross-validation approach:\n  - Splits the data into training and validation subsets.\n  - Trains a `CatBoostRegressor` model on each fold with specified hyperparameters:\n    - **GPU support** for faster computation.\n    - **Early stopping** to avoid overfitting.\n  - Collects predictions and calculates RMSLE for each fold.\n- Outputs:\n  - List of trained models.\n  - Out-of-fold predictions for overall RMSLE computation.\n\n---\n\n# 10. **Prediction and Submission (`predict_and_submit` method)**\n- Uses trained models to predict test data:\n  - Averages predictions across models for robustness.\n  - Ensures predictions are non-negative using `np.maximum(0, ...)`.\n- Updates the `sample_submission.csv` file with predictions and saves it as `submission.csv`.\n\n---\n\n# 11. **Workflow**\nThe overall workflow looks like this:\n1. **Initialize** an `Andro` object with the required paths.\n2. **Load Data** using `load_data`.\n3. **Preprocess** data with `preprocess`.\n4. **Train the Model** using `train_model`.\n5. **Predict and Submit** results with `predict_and_submit`.\n\n---","metadata":{"execution":{"iopub.status.busy":"2024-12-20T03:12:23.47378Z","iopub.execute_input":"2024-12-20T03:12:23.474129Z","iopub.status.idle":"2024-12-20T03:12:23.491293Z","shell.execute_reply.started":"2024-12-20T03:12:23.474097Z","shell.execute_reply":"2024-12-20T03:12:23.489427Z"}}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom catboost import CatBoostRegressor\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import mean_squared_log_error\nimport joblib\nimport warnings\n\nwarnings.filterwarnings(\"ignore\")\n\nclass Andro:\n    def __init__(self, train_path, test_path, sample_path, model_path):\n        self.train_path = train_path\n        self.test_path = test_path\n        self.sample_path = sample_path\n        self.model_path = model_path\n\n    def load_data(self):\n        self.train = pd.read_csv(self.train_path)\n        self.test = pd.read_csv(self.test_path)\n        self.sample = pd.read_csv(self.sample_path)\n\n        self.train.drop('id', axis=1, inplace=True)\n        self.test.drop('id', axis=1, inplace=True)\n\n        nonlog_fe, nonlog = joblib.load(self.model_path)\n        self.train['nonlog'] = nonlog_fe\n        self.test['nonlog'] = nonlog\n\n    def date_features(self, df):\n        df['Policy Start Date'] = pd.to_datetime(df['Policy Start Date'])\n        df['Year'] = df['Policy Start Date'].dt.year\n        df['Day'] = df['Policy Start Date'].dt.day\n        df['Month'] = df['Policy Start Date'].dt.month\n        df['Month_name'] = df['Policy Start Date'].dt.month_name()\n        df['Day_of_week'] = df['Policy Start Date'].dt.day_name()\n        df['Week'] = df['Policy Start Date'].dt.isocalendar().week\n        df['Year_sin'] = np.sin(2 * np.pi * df['Year'])\n        df['Year_cos'] = np.cos(2 * np.pi * df['Year'])\n        df['Month_sin'] = np.sin(2 * np.pi * df['Month'] / 12)\n        df['Month_cos'] = np.cos(2 * np.pi * df['Month'] / 12)\n        df['Day_sin'] = np.sin(2 * np.pi * df['Day'] / 31)\n        df['Day_cos'] = np.cos(2 * np.pi * df['Day'] / 31)\n        df['Group'] = (df['Year'] - 2020) * 48 + df['Month'] * 4 + df['Day'] // 7\n        df.drop('Policy Start Date', axis=1, inplace=True)\n        return df\n\n    def feature_engineering(self, df):\n        df['contract length'] = pd.cut(\n            df[\"Insurance Duration\"].fillna(99),\n            bins=[-float('inf'), 1, 3, float('inf')],\n            labels=[0, 1, 2]\n        ).astype(int)\n        return df\n\n    def frequency_encode(self, train, test, cat_cols, feature_cols, drop_org=False):\n        combined = pd.concat([train, test], axis=0, ignore_index=True)\n        new_cat_cols = []\n        for col in cat_cols:\n            freq_encoding = combined[col].value_counts().to_dict()\n            train[f\"{col}_freq\"] = train[col].map(freq_encoding).astype('float')\n            test[f\"{col}_freq\"] = test[col].map(freq_encoding).astype('float')\n            new_col_name = f\"{col}_freq\"\n            new_cat_cols.append(new_col_name)\n            feature_cols.append(new_col_name)\n            if drop_org:\n                feature_cols.remove(col)\n        return train, test, new_cat_cols, feature_cols\n\n    def preprocess(self):\n        self.train = self.date_features(self.train)\n        self.test = self.date_features(self.test)\n        self.train = self.feature_engineering(self.train)\n        self.test = self.feature_engineering(self.test)\n\n        cat_cols = [col for col in self.train.columns if self.train[col].dtype == 'object']\n        feature_cols = list(self.test.columns)\n\n        self.train, self.test, cat_cols, feature_cols = self.frequency_encode(\n            self.train, self.test, cat_cols, feature_cols, drop_org=True\n        )\n\n        self.train = self.train[feature_cols + ['Premium Amount']]\n        self.test = self.test[feature_cols]\n\n    def rmsle(self, y_true, y_pred):\n        return np.sqrt(mean_squared_log_error(y_true, y_pred))\n\n    def train_model(self):\n        X = self.train.drop('Premium Amount', axis=1)\n        y = self.train['Premium Amount']\n        y_log = np.log1p(y)\n\n        kf = KFold(n_splits=5, shuffle=True, random_state=42)\n        oof = np.zeros(len(X))\n        models = []\n\n        for fold, (train_idx, valid_idx) in enumerate(kf.split(X)):\n            print(f\"Fold {fold + 1}\")\n            X_train, X_valid = X.iloc[train_idx], X.iloc[valid_idx]\n            y_train, y_valid = y_log.iloc[train_idx], y_log.iloc[valid_idx]\n\n            model = CatBoostRegressor(\n                iterations=3000,\n                learning_rate=0.05,\n                depth=6,\n                eval_metric=\"RMSE\",\n                random_seed=42,\n                verbose=200,\n                task_type='GPU',\n                l2_leaf_reg=0.7,\n            )\n\n            model.fit(\n                X_train,\n                y_train,\n                eval_set=(X_valid, y_valid),\n                early_stopping_rounds=300,\n            )\n\n            models.append(model)\n            oof[valid_idx] = np.maximum(0, model.predict(X_valid))\n            fold_rmsle = self.rmsle(np.expm1(y_valid), np.expm1(oof[valid_idx]))\n            print(f\"Fold {fold + 1} RMSLE: {fold_rmsle}\")\n\n        overall_rmsle = self.rmsle(y, np.expm1(oof))\n        print(f\"Overall RMSLE: {overall_rmsle}\")\n\n        # Save the trained model\n        joblib.dump(models, \"andro_insurance_premium_prediction_catboost_v1.pkl\")\n\n        return models, oof\n\n    def load_trained_model(self, model_path):\n        return joblib.load(model_path)\n\n    def predict_and_submit(self, models):\n        test_predictions = np.zeros(len(self.test))\n        for model in models:\n            test_predictions += np.maximum(0, np.expm1(model.predict(self.test))) / len(models)\n\n        self.sample['Premium Amount'] = test_predictions\n        self.sample.to_csv('submission.csv', index=False)\n        return self.sample.head()\n\n    def explore_pretrained_model(self, models):\n        for i, model in enumerate(models):\n            print(f\"Exploring Model {i+1}\")\n            print(\"Feature Importances:\")\n            feature_importances = model.get_feature_importance()\n            for feature, importance in zip(self.test.columns, feature_importances):\n                print(f\"{feature}: {importance:.4f}\")\n            print(\"\\nModel Parameters:\")\n            print(model.get_all_params())\n            print(\"-\" * 50)\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-20T03:43:31.814714Z","iopub.execute_input":"2024-12-20T03:43:31.815007Z","iopub.status.idle":"2024-12-20T03:43:31.83588Z","shell.execute_reply.started":"2024-12-20T03:43:31.814985Z","shell.execute_reply":"2024-12-20T03:43:31.834868Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Example Usage\n```python\nandro = Andro(\n    train_path=\"/kaggle/input/playground-series-s4e12/train.csv\",\n    test_path=\"/kaggle/input/playground-series-s4e12/test.csv\",\n    sample_path=\"/kaggle/input/playground-series-s4e12/sample_submission.csv\",\n    model_path=\"/kaggle/input/rid-catboost-nonlog/cat_non_loged.pkl\"\n)\n\nandro.load_data()\nandro.preprocess()\nmodels, oof = andro.train_model()\nsubmission = andro.predict_and_submit(models)\nprint(submission)\n```\n\n### Key Benefits\n- **Modularity**: Each functionality is encapsulated in a method.\n- **Reusability**: Easily adapts to different datasets or pipelines.\n- **Efficiency**: Uses GPU acceleration and early stopping for faster training.","metadata":{}},{"cell_type":"code","source":"\n\nandro = Andro(\n    train_path=\"/kaggle/input/playground-series-s4e12/train.csv\",\n    test_path=\"/kaggle/input/playground-series-s4e12/test.csv\",\n    sample_path=\"/kaggle/input/playground-series-s4e12/sample_submission.csv\",\n    model_path=\"/kaggle/input/rid-catboost-nonlog/cat_non_loged.pkl\"\n)\n\nandro.load_data()\nandro.preprocess()\nmodels, oof = andro.train_model()\nsubmission = andro.predict_and_submit(models)\nprint(submission)\n\n#loading and exploring a pretrained model\n\ntrained_models = andro.load_trained_model(\"andro_insurance_premium_prediction_catboost_v1.pkl\")\nandro.explore_pretrained_model(trained_models)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-20T03:33:07.449728Z","iopub.execute_input":"2024-12-20T03:33:07.449973Z","iopub.status.idle":"2024-12-20T03:34:51.201258Z","shell.execute_reply.started":"2024-12-20T03:33:07.449951Z","shell.execute_reply":"2024-12-20T03:34:51.200487Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# ML Model Analytics and Insights\n\nHere’s an analysis based on output:\n\n---\n\n## 1. **Point-by-Point Analysis**\n\n1. **Performance Metrics**:\n   - The overall RMSLE for the model across 5 folds is **1.0341**, indicating the model's logarithmic prediction accuracy is consistent.\n   - Each fold's RMSLE shows minimal variation, highlighting stable model training.\n\n2. **Feature Importances**:\n   - **Top Features**:\n     - `Annual Income`, `nonlog`, and `Credit Score` consistently rank as the most important features.\n     - `Health Score`, `Previous Claims`, and `Age` contribute significantly.\n   - Temporal features such as `Year`, `Day`, `Month_sin`, and `Month_cos` show lower importance, indicating limited temporal influence.\n\n3. **Model Parameters**:\n   - Parameters like `depth: 6`, `learning_rate: 0.05`, and `iterations: 3000` suggest a balance between model complexity and generalization.\n   - The `bootstrap_type` used is `Bayesian`, which ensures robust sample distribution during training.\n\n4. **Iterations and Early Stopping**:\n   - The best iterations for folds vary (e.g., 1451 for Fold 1, 1442 for Fold 2).\n   - The model consistently converges around 1400–1900 iterations, indicating sufficient training cycles.\n\n---\n\n## 2. **Table Format Comparison**\n\n| **Metric/Feature**         | **Fold 1** | **Fold 2** | **Fold 3** | **Fold 4** | **Fold 5** | **Overall** |\n|-----------------------------|------------|------------|------------|------------|------------|-------------|\n| RMSLE                      | 1.0343     | 1.0337     | 1.0355     | 1.0325     | 1.0347     | 1.0341      |\n| Top Feature                | Annual Income | Annual Income | Annual Income | Annual Income | Annual Income | - |\n| Second Feature             | nonlog      | nonlog      | nonlog      | nonlog      | nonlog      | -           |\n| Third Feature              | Credit Score | Credit Score | Credit Score | Credit Score | Credit Score | -           |\n| Number of Iterations       | 1451        | 1442        | 1136        | 1428        | 1873        | -           |\n\n---\n\n## 3. **Final Summary**\n\n- **Model Strengths**:\n  - High importance for financial (`Annual Income`, `Credit Score`) and historical (`Previous Claims`) features suggests robust predictability based on user profiles.\n  - Consistent RMSLE across folds demonstrates reliable generalization.\n\n- **Improvement Areas**:\n  - Temporal features like `Day`, `Month_sin`, and `Month_cos` show limited importance, suggesting potential data enhancement or feature engineering opportunities.\n  - The model's depth (`6`) and iterations (`3000`) are standard; exploring deeper architectures or advanced regularization (e.g., dropout) might yield gains.\n\n- **Actionable Insights**:\n  - Focus data collection efforts on financial and historical metrics as they significantly impact predictions.\n  - Consider improving the temporal data representation for possible performance gains.\n","metadata":{}},{"cell_type":"markdown","source":"## ML Model Folding Visualization","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Fold numbers and their corresponding RMSLE scores\nfold_numbers = ['Fold 1', 'Fold 2', 'Fold 3', 'Fold 4', 'Fold 5']\nrmsle_scores = [1.0343, 1.0337, 1.0355, 1.0325, 1.0347]\n\n# Plotting\nplt.figure(figsize=(10, 6))\nsns.barplot(x=fold_numbers, y=rmsle_scores, palette='viridis')\n\n# Adding titles and labels\nplt.title('RMSLE Scores for Each Fold', fontsize=16)\nplt.xlabel('Fold Number', fontsize=14)\nplt.ylabel('RMSLE Score', fontsize=14)\nplt.ylim(1.030, 1.036)  # Adjust the range for better visualization\n\n# Display the plot\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-20T16:07:27.374995Z","iopub.execute_input":"2024-12-20T16:07:27.375312Z","iopub.status.idle":"2024-12-20T16:07:27.620898Z","shell.execute_reply.started":"2024-12-20T16:07:27.375281Z","shell.execute_reply":"2024-12-20T16:07:27.62011Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}