{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30805,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"border-radius: 15px; border: 2px solid #4CAF50; padding: 20px; background: linear-gradient(135deg, #81C784, #66BB6A, #228b22); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\r\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 10px; font-size: 36px; font-family: 'Roboto', sans-serif;\">\r\n        💡 LightGBM Duo: GBDT + GOSS for Premiums 🚀\r\n    </h1>\r\n</div>\r\n\r\n<!-- Include Google Fonts for a modern font -->\r\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\r\n","metadata":{}},{"cell_type":"markdown","source":"### 📝 **Dataset Overview**\r\n\r\nThe dataset used in this project is designed for **insurance premium prediction** and originates from the **Kaggle Playground Series - Season 4, Episode 12** competition. It includes a variety of numerical, categorical, and date-based features to simulate the complexities faced by insurance companies in real-world premium calculations.\r\n\r\n**Key characteristics of the dataset:**\r\n\r\n- **Total Features**: 20 (excluding the target variable).\r\n- **Target Variable**: `Premium Amount` – a continuous numerical value representing the insurance premium.\r\n- **Feature Types**:\r\n  - **Numerical Features**: Examples include `Age`, `Annual Income`, `Health Score`, `Credit Score`, and `Vehicle Age`.\r\n  - **Categorical Features**: Examples include `Gender`, `Marital Status`, `Education Level`, `Occupation`, and `Policy Type`.\r\n  - **Date Feature**: `Policy Start Date` provides temporal information related to policy initiation.\r\n  - **Derived Features**: Additional features are created through feature engineering, such as cyclical representations of dates and contract length categories.\r\n- **Data Challenges**:\r\n  - **Missing Values**: Present in several features, requiring careful imputation.\r\n  - **Skewed Data**: Features like `Annual Income` and the target `Premium Amount` exhibit skewness, necessitating log transformations.\r\n\r\n### 🎯 **Objective**\r\n\r\nThe goal of this project is to develop a robust regression model to predict the **`Premium Amount`** using customer characteristics and policy details. This objective is achieved through a systematic approach:\r\n\r\n1. **Data Preprocessing**:\r\n   - Imputing missing values in numerical and categorical features.\r\n   - Encoding categorical variables using frequency encoding.\r\n   - Engineering features from the date column (`Policy Start Date`).\r\n   - Applying log transformations to handle skewness in numerical features.\r\n\r\n2. **Model Training**:\r\n   - Training two **LightGBM** models with different boosting types:\r\n     - **GBDT (Gradient Boosting Decision Tree)**\r\n     - **GOSS (Gradient-based One-Side Sampling)**\r\n   - Using **early stopping** to prevent overfitting during training.\r\n\r\n3. **Model Evaluation**:\r\n   - Evaluating models using **Root Mean Squared Logarithmic Error (RMSLE)**.\r\n   - Performing **K-Fold Cross-Validation** to ensure model robustness and generalization.\r\n\r\n4. **Prediction and Submission**:\r\n   - Generating predictions on the test dataset.\r\n   - Combining predictions from both models using an **ensemble averaging** technique.\r\n   - Preparing a submission file in the required format for the competition.\r\n\r\nThe project aims to achieve a **competitive RMSLE score** by leveraging effective preprocessing, feature engineering, and advanced modeling techniques, providing a reliable approach for predicting insurance premiums. 🚀","metadata":{}},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Import Libraries</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #4CAF50; padding: 20px; background: linear-gradient(135deg, #81C784, #66BB6A, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 10px; font-size: 36px; font-family: 'Roboto', sans-serif;\">\n        Import Libraries\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\nfrom sklearn.model_selection import train_test_split, KFold, cross_val_score\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import LabelEncoder\nfrom sklearn.metrics import mean_squared_log_error\n\nimport lightgbm as lgb\nfrom lightgbm import early_stopping\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:32:54.766047Z","iopub.execute_input":"2024-12-13T18:32:54.766423Z","iopub.status.idle":"2024-12-13T18:32:58.755363Z","shell.execute_reply.started":"2024-12-13T18:32:54.766383Z","shell.execute_reply":"2024-12-13T18:32:58.754574Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Load the Datasets</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #4CAF50; padding: 20px; background: linear-gradient(135deg, #81C784, #66BB6A, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 10px; font-size: 36px; font-family: 'Roboto', sans-serif;\">\n        Load the Datasets\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{}},{"cell_type":"code","source":"# Load the datasets\ndf_train = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\ndf_test = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')\nsample_sub = pd.read_csv('/kaggle/input/playground-series-s4e12/sample_submission.csv')\n\n# Display first few rows\ndisplay(df_train.head())\ndisplay(df_test.head())\n\n# Verify shapes\nprint(\"Train Data Shape:\", df_train.shape)\nprint(\"Test Data Shape:\", df_test.shape)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:32:58.757712Z","iopub.execute_input":"2024-12-13T18:32:58.758759Z","iopub.status.idle":"2024-12-13T18:33:07.60227Z","shell.execute_reply.started":"2024-12-13T18:32:58.758726Z","shell.execute_reply":"2024-12-13T18:33:07.601447Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Display information for the training dataset\nprint(\"Training Dataset Information: \\n\")\ntrain_info = df_train.info()\ndisplay(train_info)\nprint('\\n')\n# Display information for the test dataset\nprint(\"Test Dataset Information: \\n\")\ntest_info = df_test.info()\ndisplay(test_info)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:33:07.603219Z","iopub.execute_input":"2024-12-13T18:33:07.603477Z","iopub.status.idle":"2024-12-13T18:33:08.51559Z","shell.execute_reply.started":"2024-12-13T18:33:07.603452Z","shell.execute_reply":"2024-12-13T18:33:08.514762Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Preprocessing</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #4CAF50; padding: 20px; background: linear-gradient(135deg, #81C784, #66BB6A, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 10px; font-size: 36px; font-family: 'Roboto', sans-serif;\">\n        Preprocessing\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{}},{"cell_type":"markdown","source":"## Drop Unnecessary Columns","metadata":{}},{"cell_type":"code","source":"# Retain 'id' for submission purposes\ntest_ids = df_test['id']\ndf_train.drop(columns=['id'], inplace=True)\ndf_test.drop(columns=['id'], inplace=True)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:33:08.516866Z","iopub.execute_input":"2024-12-13T18:33:08.517638Z","iopub.status.idle":"2024-12-13T18:33:08.824352Z","shell.execute_reply.started":"2024-12-13T18:33:08.517577Z","shell.execute_reply":"2024-12-13T18:33:08.823664Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Date Feature Engineering","metadata":{}},{"cell_type":"code","source":"def date_features(df):\n    df['Policy Start Date'] = pd.to_datetime(df['Policy Start Date'])\n    df['Year'] = df['Policy Start Date'].dt.year\n    df['Month'] = df['Policy Start Date'].dt.month\n    df['Day'] = df['Policy Start Date'].dt.day\n    df['Day_of_Week'] = df['Policy Start Date'].dt.dayofweek\n\n    # Cyclical features\n    df['Month_sin'] = np.sin(2 * np.pi * df['Month'] / 12)\n    df['Month_cos'] = np.cos(2 * np.pi * df['Month'] / 12)\n    df['Day_sin'] = np.sin(2 * np.pi * df['Day'] / 31)\n    df['Day_cos'] = np.cos(2 * np.pi * df['Day'] / 31)\n    \n    # Drop original 'Policy Start Date'\n    df.drop(columns=['Policy Start Date'], inplace=True)\n    return df\n\ndf_train = date_features(df_train)\ndf_test = date_features(df_test)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:33:08.8253Z","iopub.execute_input":"2024-12-13T18:33:08.825536Z","iopub.status.idle":"2024-12-13T18:33:10.290398Z","shell.execute_reply.started":"2024-12-13T18:33:08.825513Z","shell.execute_reply":"2024-12-13T18:33:10.289663Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Feature Engineering","metadata":{}},{"cell_type":"code","source":"def add_features(df):\n    # Contract length feature based on Insurance Duration\n    df['contract_length'] = pd.cut(\n        df['Insurance Duration'].fillna(99),\n        bins=[-float('inf'), 1, 3, float('inf')],\n        labels=[0, 1, 2]\n    ).astype(int)\n    return df\n\ndf_train = add_features(df_train)\ndf_test = add_features(df_test)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:33:10.291414Z","iopub.execute_input":"2024-12-13T18:33:10.291722Z","iopub.status.idle":"2024-12-13T18:33:10.349393Z","shell.execute_reply.started":"2024-12-13T18:33:10.291696Z","shell.execute_reply":"2024-12-13T18:33:10.34852Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Handle Missing Values","metadata":{}},{"cell_type":"code","source":"# Function to calculate missing values and percentages\ndef missing_values_table(df):\n    missing_count = df.isnull().sum()\n    missing_percentage = 100 * missing_count / len(df)\n    return pd.DataFrame({'Missing Values': missing_count, 'Percentage (%)': missing_percentage})\n\n# Create tables for train and test datasets\ntrain_missing_table = missing_values_table(df_train)\ntest_missing_table = missing_values_table(df_test)\n\n# Display the tables\nprint(\"Missing Values Table - Training Dataset:\\n\")\ndisplay(train_missing_table[train_missing_table['Missing Values'] > 0])  # Display only features with missing values\nprint(\"\\n\")\n\nprint(\"Missing Values Table - Test Dataset:\\n\")\ndisplay(test_missing_table[test_missing_table['Missing Values'] > 0])  # Display only features with missing values","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:33:10.352437Z","iopub.execute_input":"2024-12-13T18:33:10.352768Z","iopub.status.idle":"2024-12-13T18:33:11.207126Z","shell.execute_reply.started":"2024-12-13T18:33:10.352739Z","shell.execute_reply":"2024-12-13T18:33:11.206192Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Separate numerical and categorical columns\nnumerical_cols = df_train.select_dtypes(include=['float64']).columns.drop('Premium Amount')\ncategorical_cols = df_train.select_dtypes(include=['object']).columns\n\n# Impute missing numerical values with median\nimputer_num = SimpleImputer(strategy='median')\ndf_train[numerical_cols] = imputer_num.fit_transform(df_train[numerical_cols])\ndf_test[numerical_cols] = imputer_num.transform(df_test[numerical_cols])\n\n# Impute missing categorical values with mode\nimputer_cat = SimpleImputer(strategy='most_frequent')\ndf_train[categorical_cols] = imputer_cat.fit_transform(df_train[categorical_cols])\ndf_test[categorical_cols] = imputer_cat.transform(df_test[categorical_cols])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:33:11.208069Z","iopub.execute_input":"2024-12-13T18:33:11.208402Z","iopub.status.idle":"2024-12-13T18:33:16.707173Z","shell.execute_reply.started":"2024-12-13T18:33:11.208371Z","shell.execute_reply":"2024-12-13T18:33:16.706434Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Encode Categorical Variables","metadata":{}},{"cell_type":"code","source":"def frequency_encode(train, test, cat_cols):\n    for col in cat_cols:\n        freq_encoding = train[col].value_counts().to_dict()\n        train[col] = train[col].map(freq_encoding)\n        test[col] = test[col].map(freq_encoding)\n    return train, test\n\ndf_train, df_test = frequency_encode(df_train, df_test, categorical_cols)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:33:16.708434Z","iopub.execute_input":"2024-12-13T18:33:16.708747Z","iopub.status.idle":"2024-12-13T18:33:18.563893Z","shell.execute_reply.started":"2024-12-13T18:33:16.708718Z","shell.execute_reply":"2024-12-13T18:33:18.563124Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Log Transformation","metadata":{}},{"cell_type":"code","source":"# Log-transform skewed features\ndf_train['Annual Income'] = np.log1p(df_train['Annual Income'])\ndf_test['Annual Income'] = np.log1p(df_test['Annual Income'])\n\n# Log-transform target variable\ny = np.log1p(df_train['Premium Amount'])\nX = df_train.drop(columns=['Premium Amount'])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:33:18.564933Z","iopub.execute_input":"2024-12-13T18:33:18.565293Z","iopub.status.idle":"2024-12-13T18:33:18.732939Z","shell.execute_reply.started":"2024-12-13T18:33:18.565255Z","shell.execute_reply":"2024-12-13T18:33:18.732048Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Train-Test Split","metadata":{}},{"cell_type":"code","source":"X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:33:18.733965Z","iopub.execute_input":"2024-12-13T18:33:18.734213Z","iopub.status.idle":"2024-12-13T18:33:19.240431Z","shell.execute_reply.started":"2024-12-13T18:33:18.734189Z","shell.execute_reply":"2024-12-13T18:33:19.239673Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Model Training</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #4CAF50; padding: 20px; background: linear-gradient(135deg, #81C784, #66BB6A, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 10px; font-size: 36px; font-family: 'Roboto', sans-serif;\">\n        Model Training\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{}},{"cell_type":"markdown","source":"## LightGBM GBDT Model","metadata":{}},{"cell_type":"code","source":"lgbm_gbdt_model = lgb.LGBMRegressor(\n    boosting_type='gbdt',\n    n_estimators=1000,\n    learning_rate=0.05,\n    max_depth=10,\n    num_leaves=50,\n    device='gpu',\n    random_state=42\n)\n\n# Fit the model with early stopping\nlgbm_gbdt_model.fit(\n    X_train, y_train,\n    eval_set=[(X_val, y_val)],\n    callbacks=[early_stopping(stopping_rounds=50, verbose=True)]\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:33:19.241391Z","iopub.execute_input":"2024-12-13T18:33:19.241666Z","iopub.status.idle":"2024-12-13T18:33:30.682372Z","shell.execute_reply.started":"2024-12-13T18:33:19.241625Z","shell.execute_reply":"2024-12-13T18:33:30.681505Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## LightGBM GOSS Model","metadata":{}},{"cell_type":"code","source":"lgbm_goss_model = lgb.LGBMRegressor(\n    boosting_type='goss',\n    n_estimators=1000,\n    learning_rate=0.05,\n    max_depth=10,\n    num_leaves=50,\n    device='gpu',\n    random_state=42\n)\n\n# Fit the model with early stopping\nlgbm_goss_model.fit(\n    X_train, y_train,\n    eval_set=[(X_val, y_val)],\n    callbacks=[early_stopping(stopping_rounds=50, verbose=True)]\n)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:33:30.683435Z","iopub.execute_input":"2024-12-13T18:33:30.683775Z","iopub.status.idle":"2024-12-13T18:33:49.549286Z","shell.execute_reply.started":"2024-12-13T18:33:30.683746Z","shell.execute_reply":"2024-12-13T18:33:49.548377Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Cross-Validation</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #4CAF50; padding: 20px; background: linear-gradient(135deg, #81C784, #66BB6A, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 10px; font-size: 36px; font-family: 'Roboto', sans-serif;\">\n        Cross-Validation\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{}},{"cell_type":"code","source":"kf = KFold(n_splits=5, shuffle=True, random_state=42)\n\n# LightGBM GBDT Cross-Validation\ncv_scores_lgbm_gbdt = cross_val_score(lgbm_gbdt_model, X, y, cv=kf, scoring='neg_mean_squared_log_error')\nrmsle_lgbm_gbdt = np.mean(np.sqrt(-cv_scores_lgbm_gbdt))\nprint(f'LightGBM GBDT RMSLE: {rmsle_lgbm_gbdt:.5f}')\n\n# LightGBM GOSS Cross-Validation\ncv_scores_lgbm_goss = cross_val_score(lgbm_goss_model, X, y, cv=kf, scoring='neg_mean_squared_log_error')\nrmsle_lgbm_goss = np.mean(np.sqrt(-cv_scores_lgbm_goss))\nprint(f'LightGBM GOSS RMSLE: {rmsle_lgbm_goss:.5f}')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:33:49.550437Z","iopub.execute_input":"2024-12-13T18:33:49.550825Z","iopub.status.idle":"2024-12-13T18:41:21.817318Z","shell.execute_reply.started":"2024-12-13T18:33:49.550784Z","shell.execute_reply":"2024-12-13T18:41:21.816344Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Ensemble Predictions</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #4CAF50; padding: 20px; background: linear-gradient(135deg, #81C784, #66BB6A, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 10px; font-size: 36px; font-family: 'Roboto', sans-serif;\">\n        Ensemble Predictions\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{}},{"cell_type":"code","source":"# Fit the models on the full training data\nlgbm_gbdt_model.fit(X, y)\nlgbm_goss_model.fit(X, y)\n\n# Predict on the test data\ngbdt_preds = lgbm_gbdt_model.predict(df_test)\ngoss_preds = lgbm_goss_model.predict(df_test)\n\n# Average the predictions (ensemble)\nfinal_preds_log = (gbdt_preds + goss_preds) / 2\nfinal_preds = np.expm1(final_preds_log)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:41:21.818427Z","iopub.execute_input":"2024-12-13T18:41:21.818782Z","iopub.status.idle":"2024-12-13T18:43:59.261472Z","shell.execute_reply.started":"2024-12-13T18:41:21.818745Z","shell.execute_reply":"2024-12-13T18:43:59.260661Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"color:transparent;\">Create Submission File</span>\n\n<div style=\"border-radius: 15px; border: 2px solid #4CAF50; padding: 20px; background: linear-gradient(135deg, #81C784, #66BB6A, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 10px; font-size: 36px; font-family: 'Roboto', sans-serif;\">\n        Create Submission File\n    </h1>\n</div>\n\n<!-- Include Google Fonts for a modern font -->\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\n","metadata":{}},{"cell_type":"code","source":"# Prepare submission file\nsubmission = pd.DataFrame({\n    'id': test_ids,\n    'Premium Amount': final_preds\n})\n\n# Display the first few rows of the submission file\nprint(submission.head(10))\n\nsubmission.to_csv('submission.csv', index=False)\nprint(\"Submission file created successfully!\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-13T18:43:59.262717Z","iopub.execute_input":"2024-12-13T18:43:59.263153Z","iopub.status.idle":"2024-12-13T18:44:00.593248Z","shell.execute_reply.started":"2024-12-13T18:43:59.263111Z","shell.execute_reply":"2024-12-13T18:44:00.592362Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<div style=\"border-radius: 15px; border: 2px solid #4CAF50; padding: 20px; background: linear-gradient(135deg, #81C784, #66BB6A, #4CAF50); text-align: center; box-shadow: 0px 4px 8px rgba(0, 0, 0, 0.5);\">\r\n    <h1 style=\"color: #ffffff; text-shadow: 2px 2px 4px rgba(0, 0, 0, 0.7); font-weight: bold; margin-bottom: 10px; font-size: 28px; font-family: 'Roboto', sans-serif;\">\r\n        🙏 Thank You for Reading the Notebook! 🚀\r\n    </h1>\r\n    <p style=\"color: #ffffff; font-size: 18px; text-align: center;\">\r\n        Your support is truly appreciated. If you found this helpful, feel free to upvote and leave feedback.\r\n    </p>\r\n    <p style=\"color: #ffffff; font-size: 18px; text-align: center;\">\r\n        Happy Coding! 🙌😊\r\n    </p>\r\n</div>\r\n\r\n<!-- Include Google Fonts for a modern font -->\r\n<link href=\"https://fonts.googleapis.com/css2?family=Roboto:wght@700&display=swap\" rel=\"stylesheet\">\r\n","metadata":{}}]}