{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30823,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# ❓ Dividing the Dataset into NaNs & Non NaNs\n\nHi! Welcome to my notebook. For those who want to know the context of this approach, please refer to [this discussion comment here](https://www.kaggle.com/competitions/playground-series-s4e12/discussion/552568#3077137) \n\n## About the approach \n\nAs an abstract of the linked discussion, here's the explanation of the approach:\n\nHow about we split the dataset between two groups like so: abs\n\n- **NaN Group:** The split that includes all NaN values from the entire train & test set\n- **Non NaN Group:** The split that includes zero NaN values.\n\nIf we group the data like so, here will be the proportions of the splits: \n\n| Dataset | Total Number of Rows | NaNs Group's Rows | No NaN's Group's Rows\n| --- | --- | --- | --- |\n| Train | 1,200,000 | 815,996 ~(68%) | 384,004 ~(32%) | \n| Test | 800,000 | 544,642 ~(68%) | 255,358 ~(32%) |\n\n\n### What's the purpose of such splitting?\n\nwe can test out how the models perform in this manner. For baseline I will be ensembling XGBoost, CatBoost & LightGBM.  \n\n","metadata":{}},{"cell_type":"code","source":"print('50 TRIALS PER MODEL | STARTED')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd \nimport numpy as np\n\nfrom matplotlib import pyplot as plt \nimport seaborn as sns","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:07:21.520801Z","iopub.execute_input":"2024-12-21T19:07:21.521116Z","iopub.status.idle":"2024-12-21T19:07:22.253739Z","shell.execute_reply.started":"2024-12-21T19:07:21.52109Z","shell.execute_reply":"2024-12-21T19:07:22.253053Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\ntest = pd.read_csv('/kaggle/input/playground-series-s4e12/test.csv')\n\ntrain.drop(columns=['id'], inplace=True) ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:07:22.254887Z","iopub.execute_input":"2024-12-21T19:07:22.255337Z","iopub.status.idle":"2024-12-21T19:07:28.370659Z","shell.execute_reply.started":"2024-12-21T19:07:22.255301Z","shell.execute_reply":"2024-12-21T19:07:28.369962Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Null Values Visualization 📊\n\nLet's see the sequence of null values in the dataset first","metadata":{}},{"cell_type":"code","source":"def visualize_nulls(df, title='Nulls in Train'):\n    plt.figure(figsize=(10, 12))\n    sns.heatmap(df.isnull(), cbar=False, cmap='viridis')\n\n    plt.title(title)\n    plt.show() ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:07:28.372243Z","iopub.execute_input":"2024-12-21T19:07:28.372563Z","iopub.status.idle":"2024-12-21T19:07:28.376735Z","shell.execute_reply.started":"2024-12-21T19:07:28.372538Z","shell.execute_reply":"2024-12-21T19:07:28.37603Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualize_nulls(train) \nvisualize_nulls(test, title='Nulls in Test')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:07:28.378025Z","iopub.execute_input":"2024-12-21T19:07:28.378303Z","iopub.status.idle":"2024-12-21T19:08:01.986236Z","shell.execute_reply.started":"2024-12-21T19:07:28.378281Z","shell.execute_reply":"2024-12-21T19:08:01.985372Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Interesting, now let's settle down the NaN group in the bottom in both the dfs ","metadata":{}},{"cell_type":"code","source":"def consolidate_nulls(df):\n    nulls = df.isnull().sum(axis=1)\n    return df.iloc[nulls.argsort()]\n\ntrain = consolidate_nulls(train)\ntest = consolidate_nulls(test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:08:01.987309Z","iopub.execute_input":"2024-12-21T19:08:01.98767Z","iopub.status.idle":"2024-12-21T19:08:03.57538Z","shell.execute_reply.started":"2024-12-21T19:08:01.987636Z","shell.execute_reply":"2024-12-21T19:08:03.574673Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualize_nulls(train) \nvisualize_nulls(test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:08:03.576139Z","iopub.execute_input":"2024-12-21T19:08:03.576384Z","iopub.status.idle":"2024-12-21T19:08:36.90953Z","shell.execute_reply.started":"2024-12-21T19:08:03.576362Z","shell.execute_reply":"2024-12-21T19:08:36.908587Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Okay cool, so let's finally create the splits\n\n## Splitting the datasets ✂\n\n- `{df}_nn` = the {df} is from `not null` group\n- `{df}_n` = the {df} is from `null` group","metadata":{}},{"cell_type":"code","source":"def split_df(df):\n    non_nulls = df.dropna()\n    nulls = df[df.isnull().any(axis=1)]\n\n    return non_nulls, nulls\n\n\ntrain_nn, train_n = split_df(train)\ntest_nn, test_n = split_df(test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:08:36.910531Z","iopub.execute_input":"2024-12-21T19:08:36.910859Z","iopub.status.idle":"2024-12-21T19:08:39.112294Z","shell.execute_reply.started":"2024-12-21T19:08:36.910825Z","shell.execute_reply":"2024-12-21T19:08:39.111594Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_nn.shape, test_nn.shape ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T20:39:25.991119Z","iopub.execute_input":"2024-12-21T20:39:25.991427Z","iopub.status.idle":"2024-12-21T20:39:25.996418Z","shell.execute_reply.started":"2024-12-21T20:39:25.991402Z","shell.execute_reply":"2024-12-21T20:39:25.995729Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_n.shape, test_n.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T20:39:44.405862Z","iopub.execute_input":"2024-12-21T20:39:44.406214Z","iopub.status.idle":"2024-12-21T20:39:44.41149Z","shell.execute_reply.started":"2024-12-21T20:39:44.406185Z","shell.execute_reply":"2024-12-21T20:39:44.410585Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🔎 Verifying the splits \n\nLet's verify if the splits are rightly done","metadata":{}},{"cell_type":"code","source":"visualize_nulls(train_nn, title='Train Non Nulls Split') \nvisualize_nulls(train_n, title='Train Nulls Split')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:08:39.11491Z","iopub.execute_input":"2024-12-21T19:08:39.115147Z","iopub.status.idle":"2024-12-21T19:09:00.134513Z","shell.execute_reply.started":"2024-12-21T19:08:39.115127Z","shell.execute_reply":"2024-12-21T19:09:00.133635Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualize_nulls(test_nn, title='Test Non Nulls Split') \nvisualize_nulls(test_n, title='Test Nulls Split')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:09:00.135762Z","iopub.execute_input":"2024-12-21T19:09:00.136082Z","iopub.status.idle":"2024-12-21T19:09:14.396383Z","shell.execute_reply.started":"2024-12-21T19:09:00.136057Z","shell.execute_reply":"2024-12-21T19:09:14.395383Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ⁉ Some Unexpected Results\n\nWhen calculating %age of missing values in the splits we created, well.. only ~7% !?","metadata":{}},{"cell_type":"code","source":"def percent_nulls_in_df(df):\n    return df.isnull().sum().sum() / (df.shape[0] * df.shape[1]) * 100\n\nprint(f'Nulls in test group: {percent_nulls_in_df(test_n):.2f}%')\nprint(f'Nulls in train group: {percent_nulls_in_df(train_n):.2f}%')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:09:14.397277Z","iopub.execute_input":"2024-12-21T19:09:14.397573Z","iopub.status.idle":"2024-12-21T19:09:15.060867Z","shell.execute_reply.started":"2024-12-21T19:09:14.397543Z","shell.execute_reply":"2024-12-21T19:09:15.060125Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(f'Nulls in test (no-split): {percent_nulls_in_df(test):.2f}%')\nprint(f'Nulls in train (no-split): {percent_nulls_in_df(train):.2f}%')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:09:15.061629Z","iopub.execute_input":"2024-12-21T19:09:15.061847Z","iopub.status.idle":"2024-12-21T19:09:16.063564Z","shell.execute_reply.started":"2024-12-21T19:09:15.061828Z","shell.execute_reply":"2024-12-21T19:09:16.062843Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Okay... Makes sense now. Even though the NaN %age is so low, the groups occupy 68%-32% of the entire dfs. The reason is simple - not all NaN values are consecutively present and not all columns have same distribution of nan values.","metadata":{}},{"cell_type":"markdown","source":"## Preprocessing The Splits \n\nFor each split, the preprocess steps are as follows:\n\n- Convert the `Policy Start Date` to numerical by converting the timestamps to milliseconds since epoch. For reference, [see this discussion comment here](https://www.kaggle.com/competitions/playground-series-s4e12/discussion/549691#3062788)\n\n- Label NaN values completely uniquely. For reference [see this discussion](https://www.kaggle.com/competitions/playground-series-s4e12/discussion/552165) and [this one](https://www.kaggle.com/competitions/playground-series-s4e12/discussion/552568)\n\n- Perform One-Hot Encoding on object (categorical) columns. (personal experiment), as I've tried categorical encoding already [here in my earlier notebook](https://www.kaggle.com/code/eraakash/insurance-regress-detailed-notebook-x6-lgbm)\n\n- If the split originates from train, tranform the target variable to `log1p` (or exponent logarithm). For reference, [here's the discussion](https://www.kaggle.com/competitions/playground-series-s4e12/discussion/549336)","metadata":{}},{"cell_type":"code","source":"import warnings  \nwarnings.filterwarnings('ignore')\n\ndef preprocess_split(df, is_train=False): \n    df['Policy Start Date'] = pd.to_datetime(df['Policy Start Date'])\n    df['Policy Start Date'] = df['Policy Start Date'].astype('int64') / 10**9\n\n    obj_cols = df.select_dtypes('object').columns\n    non_obj_cols = [col for col in df.columns if col not in obj_cols] \n\n    for col in df.columns: \n        if df[col].isnull().sum() > 0:\n            \n            substitute = 'Other' if col in obj_cols else -1\n            df[col].fillna(substitute, inplace=True) \n\n    if is_train: \n        df['Premium Amount'] = np.log1p(df['Premium Amount'])\n\n    df = pd.get_dummies(df)\n    return df \n\n\ntrain_nn = preprocess_split(train_nn, True) \ntrain_n = preprocess_split(train_n, True) \n\ntest_nn = preprocess_split(test_nn) \ntest_n = preprocess_split(test_n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:09:16.064417Z","iopub.execute_input":"2024-12-21T19:09:16.064742Z","iopub.status.idle":"2024-12-21T19:09:20.270578Z","shell.execute_reply.started":"2024-12-21T19:09:16.064708Z","shell.execute_reply":"2024-12-21T19:09:20.269824Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ⚙ Custom Model Training Setup \n\nLet's setup some optimized approach. \n\n- `models_db` will act as a database storing all models created during hyperparameter tuning. I will pick best models from there.\n\n### Evaluation \n\nI will be evaluating using RMSE because the target is already converted into `log1p` ","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import mean_squared_error as mse \nimport optuna \n\nfrom lightgbm import LGBMRegressor\nfrom catboost import CatBoostRegressor\n\nfrom sklearn.model_selection import KFold\nfrom xgboost import XGBRegressor\n\nmodels_db = {\n    'XGB_TrainNulls': [],\n    'XGB_TrainNoNulls': [],\n    'LGB_TrainNulls': [],\n    'LGB_TrainNoNulls': [],\n    'CAT_TrainNulls': [],\n    'CAT_TrainNoNulls': [],\n}\n\ncurr_idx = 0\n\ndef save_model(mdl, score, group):\n    models_db[f'{group}'].append((score, mdl))\n\n\ndef rmse(y_true, y_pred):    \n    return mse(y_true, y_pred) ** 0.5 \n\n\ndef train_model(df, target, params={}, model=None, rs=84, use_gpu=False):\n    \"\"\"\n    Trains a classification model with optional GPU support.\n\n    Args:\n        df (DataFrame): The dataset containing features and target.\n        target (str): The target column name.\n        params (dict): Model-specific hyperparameters.\n        model (str): The type of model to train ('lgb', 'cat', or 'xgb').\n        rs (int): Random state for reproducibility.\n        use_gpu (bool): Whether to use GPU for training. Default is False.\n\n    Returns:\n        tuple: The trained model and mean accuracy across folds.\n    \"\"\"\n\n    device = 'GPU' if use_gpu else 'CPU'\n    print(f'[!]; {device} | Training started.')\n\n    X = df.drop(target, axis=1)\n    y = df[target]\n\n    skf = KFold(n_splits=10, shuffle=True, random_state=rs)\n    acc = []\n\n    def progress_bar(style='-', current=1, total=5, length=40, label='folds'):\n        progress = current / total * length\n        bar = style * int(progress)\n\n        print(f'[{bar.ljust(length)}] {current}/{total} {label}', end='\\r')\n\n    fold = 0\n\n    for train_index, test_index in skf.split(X, y):\n        progress_bar(current=fold+1, total=10, label='folds')\n\n        X_train, X_test = X.iloc[train_index], X.iloc[test_index]\n        y_train, y_test = y.iloc[train_index], y.iloc[test_index]\n\n        if model == 'lgb':\n            if use_gpu:\n                params.update({'device': 'gpu'})\n            clf = LGBMRegressor(**params)\n            clf.fit(\n                X_train, y_train,\n                eval_set=[(X_test, y_test)]\n            )\n\n        elif model == 'cat':\n            if use_gpu:\n                params.update({'task_type': 'GPU'})\n            clf = CatBoostRegressor(**params)\n            clf.fit(\n                X_train, y_train,\n                eval_set=[(X_test, y_test)],\n                early_stopping_rounds=50,\n                verbose=False\n            )\n\n        else:  # XGBoost\n            if use_gpu:\n                params.update({'tree_method': 'gpu_hist'})\n            clf = XGBRegressor(**params, early_stopping_rounds=50)\n\n            clf.fit(\n                X_train, y_train,\n                eval_set=[(X_test, y_test)],\n                verbose=False\n            )\n\n        y_pred = np.abs(clf.predict(X_test))\n        acc.append(rmse(np.abs(y_test), y_pred))\n\n        fold += 1\n\n    print(f'[!]; Training completed.')\n    return clf, np.mean(acc)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:13:03.847382Z","iopub.execute_input":"2024-12-21T19:13:03.847695Z","iopub.status.idle":"2024-12-21T19:13:03.858983Z","shell.execute_reply.started":"2024-12-21T19:13:03.847671Z","shell.execute_reply":"2024-12-21T19:13:03.858155Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def optimize(model='XGB', y='Premium Amount', df=train_nn, use_gpu=True, group=''):\n\n    if model not in ['XGB', 'LGB', 'CAT']:\n        print('Invalid Model Selected | Choose from XGB/LGB/CAT')\n        return \n\n    def objective(trial):\n\n        obj_params = {\n            'XGB': {\n                # 'n_estimators': trial.suggest_int('n_estimators', 100, 1000),\n                'learning_rate': trial.suggest_float('learning_rate', 1e-3, 0.25),\n                'max_depth': trial.suggest_int('max_depth', 3, 10),\n                'subsample': trial.suggest_float('subsample', 0.1, 1.0),\n                'colsample_bytree': trial.suggest_float('colsample_bytree', 0.1, 1.0),\n                'gamma': trial.suggest_float('gamma', 0.1, 1.0),\n                'random_state': 84,\n                'eval_metric': 'rmse'\n            },\n    \n            'LGB': {\n                # 'n_estimators': trial.suggest_int('n_estimators', 100, 1000),\n                'learning_rate': trial.suggest_float('learning_rate', 1e-3, 0.25),\n                'max_depth': trial.suggest_int('max_depth', 3, 10),\n                'subsample': trial.suggest_float('subsample', 0.1, 1.0),\n                'colsample_bytree': trial.suggest_float('colsample_bytree', 0.1, 1.0),\n                'random_state': 84,\n                'n_jobs': -1,\n                'verbosity': -1,\n                'early_stopping_round': 50,\n                'eval_metric': 'rmse'\n            },\n    \n            'CAT': {\n                # 'n_estimators': trial.suggest_int('n_estimators', 100, 1000),\n                'learning_rate': trial.suggest_float('learning_rate', 1e-3, 0.25),\n                'max_depth': trial.suggest_int('max_depth', 3, 10),\n                'random_state': 84,\n                'l2_leaf_reg': trial.suggest_float('l2_leaf_reg', 1.0, 10.0),\n                'grow_policy': trial.suggest_categorical('grow_policy', ['SymmetricTree', 'Depthwise', 'Lossguide']),\n                'min_data_in_leaf': trial.suggest_int('min_data_in_leaf', 1, 100),\n                'eval_metric': 'RMSE'\n            }\n        }\n\n        params = obj_params[model]\n        \n        mdl, acc = train_model(df, y, params, model.lower(), use_gpu=use_gpu)\n        save_model(mdl,acc, group)\n\n        return acc \n\n    return objective ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:13:07.578749Z","iopub.execute_input":"2024-12-21T19:13:07.579109Z","iopub.status.idle":"2024-12-21T19:13:07.586395Z","shell.execute_reply.started":"2024-12-21T19:13:07.579077Z","shell.execute_reply":"2024-12-21T19:13:07.585461Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"group_names = ['TrainNulls', 'TrainNoNulls']\n\nstudies = [\n    {\n        'name': 'XGBoost | Train - No Nulls',\n        'study': optimize(model='XGB', df=train_nn, group=f'XGB_{group_names[1]}')\n    },\n    {\n        'name': 'XGBoost | Train - Nulls',\n        'study': optimize(model='XGB', df=train_n, group=f'XGB_{group_names[0]}')\n    },\n    {\n        'name': 'LightGBM | Train - No Nulls',\n        'study': optimize(model='LGB', df=train_nn, use_gpu=False, group=f'LGB_{group_names[1]}')\n    },\n    {\n        'name': 'LightGBM | Train - Nulls',\n        'study': optimize(model='LGB', df=train_n, use_gpu=False, group=f'LGB_{group_names[0]}')\n    },\n    {\n        'name': 'CatBoost | Train - No Nulls',\n        'study': optimize(model='CAT', df=train_nn, group=f'CAT_{group_names[1]}')\n    },\n    {\n        'name': 'CatBoost | Train - Nulls',\n        'study': optimize(model='CAT', df=train_n, group=f'CAT_{group_names[0]}')\n    },\n]\n\n\nfor study in studies:\n    print(f\"\\n\\n\\n{study['name']} \\n\\n\\n\")\n    \n    objective = study['study']\n    obj_study = optuna.create_study(direction='minimize')\n\n    obj_study.optimize(objective, n_trials=50)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T19:13:10.290803Z","iopub.execute_input":"2024-12-21T19:13:10.291178Z","iopub.status.idle":"2024-12-21T19:49:52.094575Z","shell.execute_reply.started":"2024-12-21T19:13:10.291145Z","shell.execute_reply":"2024-12-21T19:49:52.093548Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Collecting Best Models & Predicting\n\nWell, we got plenty of models now. Let's pick the best models from each group, and finally generate our predictions","metadata":{}},{"cell_type":"code","source":"def find_best_models(db):\n    groups = db.keys()\n    best_models = {}\n\n    for group in groups:\n        models = db[group]\n        best = sorted(models, key=lambda x: x[0])[0]\n\n        best_models[group] = best\n\n    return best_models\n\n\ndef determine_test_dataset(group):\n    if 'NoNulls' in group:\n        return test_nn \n    else:\n        return test_n \n\n\ndef predict_test_data(group, model):\n    test_data = determine_test_dataset(group)\n    model = model[1]\n\n    ids = test_data['id']\n    preds = model.predict(test_data.drop(columns=['id']))\n\n    res = (ids, preds)\n    return res ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T20:10:05.834541Z","iopub.execute_input":"2024-12-21T20:10:05.834834Z","iopub.status.idle":"2024-12-21T20:10:05.839983Z","shell.execute_reply.started":"2024-12-21T20:10:05.834811Z","shell.execute_reply":"2024-12-21T20:10:05.839122Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def gather_all_predictions():\n    all_preds = []\n    best_models = find_best_models(models_db)\n    \n    for group, model in best_models.items():\n        print(f'[#] | {group}')\n        preds = predict_test_data(group, model)\n        all_preds.append(preds)\n\n    return all_preds\n\n\nall_predictions = gather_all_predictions()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T20:10:08.925461Z","iopub.execute_input":"2024-12-21T20:10:08.925785Z","iopub.status.idle":"2024-12-21T20:10:13.68861Z","shell.execute_reply.started":"2024-12-21T20:10:08.925758Z","shell.execute_reply":"2024-12-21T20:10:13.68791Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Consolidate Predictions \n\nNow let's combine all the floating chunks of predictions into one submission file.","metadata":{}},{"cell_type":"code","source":"def save_predictions(all_predictions):\n    all_preds = {\n        'id': [],\n        'Premium Amount': []\n    }\n\n    for preds in all_predictions:\n        ids = preds[0].tolist()\n        estimates = np.expm1(preds[1]) \n        \n        all_preds['id'].extend(ids)\n        all_preds['Premium Amount'].extend(estimates)\n\n    all_preds_df = pd.DataFrame(all_preds)\n    \n    all_preds_df.to_csv('/kaggle/working/raw_sub.csv', index=False)\n    all_preds_df = all_preds_df.groupby('id').median().reset_index()\n\n    \n    all_preds_df = all_preds_df.sort_values(by='id', ascending=True)\n    all_preds_df.to_csv('/kaggle/working/submission.csv', index=False)\n    \n    return all_preds_df\n\n\nsubmission = save_predictions(all_predictions)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T21:19:41.589754Z","iopub.execute_input":"2024-12-21T21:19:41.590124Z","iopub.status.idle":"2024-12-21T21:19:44.278749Z","shell.execute_reply.started":"2024-12-21T21:19:41.590091Z","shell.execute_reply":"2024-12-21T21:19:44.278024Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission.head(10)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-21T21:19:47.120865Z","iopub.execute_input":"2024-12-21T21:19:47.12121Z","iopub.status.idle":"2024-12-21T21:19:47.12919Z","shell.execute_reply.started":"2024-12-21T21:19:47.121182Z","shell.execute_reply":"2024-12-21T21:19:47.128471Z"}},"outputs":[],"execution_count":null}]}