{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30823,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🌟**S4E12 | Neural Network Regression** 🌟\n\nWelcome to the **Neural Network Regression** phase of our journey! 🚀 \n\nWe’ve been progressing step-by-step, building a deeper understanding of our dataset and leveraging the power of deep learning. Here’s a quick recap of the milestones we’ve achieved so far:  \n\n---\n\n### 📚 **Our Journey So Far**\n1. **[Starter Pack](https://www.kaggle.com/code/isaranja/s4e12-starter-pack-by-ama)**  \n   *Launched with an overview of the dataset and initial setup.*  \n\n2. **[EDA and Transformation](https://www.kaggle.com/code/isaranja/s4e12-eda-transformation-by-ama)**  \n   *Explored the data, uncovered patterns, and applied necessary transformations.*  \n\n3. **[Feature Engineering](https://www.kaggle.com/code/isaranja/s4e12-feature-engineering-by-ama)**  \n   *Crafted and refined features to highlight hidden insights.*  \n\n4. **[Algorithm Spot Check](https://www.kaggle.com/code/isaranja/s4e12-algorithm-spot-check-by-ama)**  \n   *Assessed multiple algorithms, including traditional ML models, to establish a baseline.*  \n\n---\n\n### 🧠 **Neural Network Approach**\nNow, it’s time to harness the power of **Deep Learning** with Neural Networks! 🔍  \nWe’ll train a fully connected neural network model with all the transformations, engineered features, and preprocessing steps that we have meticulously crafted. Let’s dive in! 🚀  \n\n---\n\n#### 💡 **What’s Next?**\n- Preprocess the data (scaling, encoding, and handling missing values).  \n- Design a neural network architecture tailored to regression.  \n- Train the model using GPU for faster computation.  \n- Evaluate the model using **Root Mean Squared Error (RMSE)** as the evaluation metric.  \n- Add enhancements like **batch normalization**, **dropout**, or **early stopping** for optimal performance.  \n\n---\n\nStay tuned as we unlock the potential of Neural Networks for our regression problem! 🌟","metadata":{}},{"cell_type":"markdown","source":"### Library loading and configurations","metadata":{}},{"cell_type":"code","source":"\nfrom IPython.core.interactiveshell import InteractiveShell #\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom scipy.stats import f_oneway, pointbiserialr, pearsonr, spearmanr # statistical analysis\n\nfrom sklearn.preprocessing import LabelEncoder,OneHotEncoder, MinMaxScaler, StandardScaler # transformation\n\nimport lightgbm as lgb # lightgbm \nfrom lightgbm import LGBMRegressor # lightgbm sklern api\n\nfrom sklearn.model_selection import train_test_split # data split\nfrom sklearn.metrics import mean_squared_log_error, mean_squared_error # evaluation metrices\nimport optuna # Hyperparameter Tuning\n\nfrom tabulate import tabulate # tabulate printing\n\nimport seaborn as sns # plots\nimport matplotlib.pyplot as plt # plots\n\nimport warnings #\n\n#********** Settings **************************************************************************************\n# settings for jupyter envioronment\n\n# This ensures that plots are rendered inline\n%matplotlib inline\n\n# This ensures that all output, including text and plots, is shown automatically\nInteractiveShell.ast_node_interactivity = \"all\"\n\n# Switching off the future warrnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)\n\ndef add_binned_feature_with_target_mean_and_frequency(column_name, target_column, bins=100):\n    \"\"\"\n    Adds two new columns to the global DataFrame:\n    - The mean of the target variable for each bin of the specified column.\n    - The frequency of each bin.\n\n    Args:\n    - column_name (str): The name of the column to be binned.\n    - target_column (str): The name of the target column whose mean will be calculated within bins.\n    - bins (int): The number of bins to use for discretization (default is 100).\n    \"\"\"\n    \n    # Check if the column exists in the DataFrame\n    if column_name not in df.columns or target_column not in df.columns:\n        print(f\"Error: {column_name} or {target_column} not found in the DataFrame.\")\n        return\n    \n    # Bin the specified column using pd.cut\n    df[f'{column_name}_binned'] = pd.cut(df[column_name], bins=bins, labels=False, include_lowest=True)\n    \n    # Calculate the mean of the target variable within each bin\n    mean_target_per_bin = df.groupby(f'{column_name}_binned')[target_column].mean()\n    \n    # Map the mean target values back to the original dataframe\n    df[f'{column_name}_binned_mean'] = df[f'{column_name}_binned'].map(mean_target_per_bin)\n    \n    # Calculate the frequency of each bin\n    bin_frequency = df[f'{column_name}_binned'].value_counts(normalize=False)\n    \n    # Map the frequency values back to the original dataframe\n    df[f'{column_name}_binned_freq'] = df[f'{column_name}_binned'].map(bin_frequency)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-31T17:15:05.487883Z","iopub.execute_input":"2024-12-31T17:15:05.488205Z","iopub.status.idle":"2024-12-31T17:15:09.551053Z","shell.execute_reply.started":"2024-12-31T17:15:05.488164Z","shell.execute_reply":"2024-12-31T17:15:09.550166Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Loading data","metadata":{}},{"cell_type":"code","source":"\ntrain_df = pd.read_csv(\"/kaggle/input/playground-series-s4e12/train.csv\",parse_dates=['Policy Start Date'])\ntest_df = pd.read_csv(\"/kaggle/input/playground-series-s4e12/test.csv\",parse_dates=['Policy Start Date'])\n\n# Merging two dataframes after adding trn and tst tag\ntrain_df['src']='trn'\ntest_df['src']='tst'\n\ndf = pd.concat([train_df, test_df], ignore_index=True)\n\ndf.columns = df.columns.str.replace(' ', '_', regex=False) # Removing whitespaces","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T17:15:09.552096Z","iopub.execute_input":"2024-12-31T17:15:09.552601Z","iopub.status.idle":"2024-12-31T17:15:18.700151Z","shell.execute_reply.started":"2024-12-31T17:15:09.552568Z","shell.execute_reply":"2024-12-31T17:15:18.699219Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Feature Engineering","metadata":{}},{"cell_type":"code","source":"\n#++++++++++++++++++++++++\ndf['annual_income_null'] = df['Annual_Income'].isnull().map({True: 1, False: 0})\ndf['credit_score_null'] = df['Credit_Score'].isnull().map({True: 1, False: 0})\ndf['previous_claims_null'] = df['Previous_Claims'].isnull().map({True: 1, False: 0})\n#++++++++++++++++++++++++\n\n#********** Transformation *********************************************************************************\n# Replace 'inf' and '-inf' with NaN\ndf.replace([np.inf, -np.inf], np.nan, inplace=True)\n\n# Missing value imputation\nfill_values = {'Age': 0, \n               'Annual_Income': df['Annual_Income'].max(), # since very low count there\n               'Marital_Status': 'unknown',\n               'Number_of_Dependents':5.0,\n               'Occupation':'unknown',\n               'Health_Score':0.0,\n               'Previous_Claims':10.0,\n               'Vehicle_Age':0.0,\n               'Credit_Score':900,\n               'Insurance_Duration':0.0,\n               'Customer_Feedback':'unknown',\n               'Premium_Amount':1\n              }\n\ndf.fillna(value=fill_values, inplace=True)\n\n# Log transformation\ndf['annual_income_log'] = np.log1p(df['Annual_Income'])\ndf['premium_amount_log'] = np.log1p(df['Premium_Amount'])\n\n#********** Feature Engineering *****************************************************************************\n\n# deriving year from the policy start date\ndf['policy_start_year'] = df['Policy_Start_Date'].dt.year\n\n# deriving the month from the policy start date \ndf['policy_start_month'] = df['Policy_Start_Date'].dt.month\n\n# deriving the day of policy start date\ndf['policy_start_day'] = df['Policy_Start_Date'].dt.day\n\nmax_date = df['Policy_Start_Date'].max()\ndf['policy_age'] = (max_date - df['Policy_Start_Date']).dt.days\ndf['policy_age_bins'] = pd.cut(df['policy_age'], bins=[df['policy_age'].min()-1,1700,df['policy_age'].max()], labels=['new','old'])\nadd_binned_feature_with_target_mean_and_frequency('policy_age', 'premium_amount_log', bins=100)\n\n# Annual Income\ndf['annual_income_bins'] = pd.cut(df['annual_income_log'], bins=[df['annual_income_log'].min()-1,1,10.9,df['annual_income_log'].max()], labels=['no-data','mid','high'])\nadd_binned_feature_with_target_mean_and_frequency('Annual_Income', 'premium_amount_log', bins=100)\n\n# Credit Score\ndf['credit_score_bins'] = pd.cut(df['Credit_Score'], bins=[df['Credit_Score'].min()-1,380,550,805,840,df['Credit_Score'].max()], labels=['S','M','L','XL','no-data'])\nadd_binned_feature_with_target_mean_and_frequency('Credit_Score', 'premium_amount_log', bins=100)\n\n# Health Score\ndf['health_score_bins'] = pd.cut(df['Health_Score'], bins=[df['Health_Score'].min()-1,1,50,df['Health_Score'].max()], labels=['S','M','L'])\nadd_binned_feature_with_target_mean_and_frequency('Health_Score', 'premium_amount_log', bins=50)\n\n# Annual Income per dependents\ndf['annual_income_to_dependent'] = np.log1p(df['Annual_Income']/(df['Number_of_Dependents']+1))\n\n# Health score and Credit Score combined\ndf['health_and_credit_score'] = df['Health_Score']+(df['Credit_Score'])\ndf['health_and_credit_score_bins'] = pd.cut(df['health_and_credit_score'], bins=[df['health_and_credit_score'].min()-1,440,575,df['health_and_credit_score'].max()], labels=['S','M','L'])\nadd_binned_feature_with_target_mean_and_frequency('health_and_credit_score', 'premium_amount_log', bins=100)\n\n#Annual Income per Insurance Duration\ndf['annual_income_to_insurance_duration'] = np.log1p(df['Annual_Income']/(df['Insurance_Duration']+1))\ndf['annual_income_to_insurance_duration_bins'] = pd.cut(df['annual_income_to_insurance_duration'], bins=[df['annual_income_to_insurance_duration'].min()-1,8.5,9.75,10.9,df['annual_income_to_insurance_duration'].max()], labels=['S','M','L','XL'])\n\n# Annual Income per Age\ndf['annual_income_to_age'] = np.log1p(df['Annual_Income']/(df['Age']+1))\ndf['annual_income_to_age_bins'] = pd.cut(df['annual_income_to_age'], bins=[df['annual_income_to_age'].min()-1,6.6,8.2,df['annual_income_to_age'].max()], labels=['S','M','L'])\nadd_binned_feature_with_target_mean_and_frequency('annual_income_to_age', 'premium_amount_log', bins=100)\n\n# Annual Income to Health Score ratio\ndf['annual_income_to_health_score'] = np.log1p(df['Annual_Income']/(df['Health_Score']+1))\ndf['annual_income_to_health_score_bins'] = pd.cut(df['annual_income_to_health_score'], bins=[df['annual_income_to_health_score'].min()-1,6.8,9,df['annual_income_to_health_score'].max()], labels=['S','M','L'])\nadd_binned_feature_with_target_mean_and_frequency('annual_income_to_health_score', 'premium_amount_log', bins=100)\n\n# Credit Score to Dependents ratio\ndf['credit_score_to_dependent'] = np.log1p(df['Credit_Score']/(df['Number_of_Dependents']+1))\ndf['credit_score_to_dependent_bins'] = pd.cut(df['credit_score_to_dependent'], bins=[df['credit_score_to_dependent'].min()-1,5,6.25,df['credit_score_to_dependent'].max()], labels=['S','M','L'])\nadd_binned_feature_with_target_mean_and_frequency('credit_score_to_dependent', 'premium_amount_log', bins=100)\n\n# Annual Income to Policy Age\ndf['annual_income_to_policy_age'] = np.log1p(df['Annual_Income']/(df['policy_age']+1))\nadd_binned_feature_with_target_mean_and_frequency('annual_income_to_policy_age', 'premium_amount_log', bins=100)\n\n# Credit Score to Previous Claims\ndf['credit_score_to_previous_claims'] = df['Credit_Score']/(df['Previous_Claims']+1)\ndf['credit_score_to_previous_claims_bins'] = pd.cut(df['credit_score_to_previous_claims'], bins=[df['credit_score_to_previous_claims'].min()-1,300,440,565,df['credit_score_to_previous_claims'].max()], labels=['S','M','L','XL'])\nadd_binned_feature_with_target_mean_and_frequency('credit_score_to_previous_claims', 'premium_amount_log', bins=100)\n\n# Annual Income to Previous Claims\ndf['annual_income_to_previous_claims'] = np.log1p(df['Annual_Income']/(df['Previous_Claims']+1))\nadd_binned_feature_with_target_mean_and_frequency('annual_income_to_previous_claims', 'premium_amount_log', bins=100)\n\n# Dependents and previous claims\ndf['dependents_and_previous_claims'] = (df['Number_of_Dependents'])*(df['Previous_Claims'])\n\n#********** Final Dataset *****************************************************************************\n# Dataset looks like                                             \nwith pd.option_context('display.max_columns', None): # setting the max rows\n    display(df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T17:15:18.701834Z","iopub.execute_input":"2024-12-31T17:15:18.702061Z","iopub.status.idle":"2024-12-31T17:15:25.739385Z","shell.execute_reply.started":"2024-12-31T17:15:18.702042Z","shell.execute_reply":"2024-12-31T17:15:25.73844Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Features","metadata":{}},{"cell_type":"code","source":"\nfeature_list= [\n 'Age',\n 'Gender',\n 'Marital_Status',\n 'Number_of_Dependents',\n 'Education_Level',\n 'Occupation',\n 'Health_Score',\n 'Location',\n 'Policy_Type',\n 'Previous_Claims',\n 'Vehicle_Age',\n 'Credit_Score',\n 'Insurance_Duration',\n 'Customer_Feedback',\n 'Smoking_Status',\n 'Exercise_Frequency',\n 'Property_Type',\n 'annual_income_log',\n 'policy_start_year',\n 'policy_start_month',\n 'policy_start_day',\n 'policy_age',\n 'policy_age_bins','annual_income_bins','credit_score_bins','health_score_bins',\n 'annual_income_to_dependent','health_and_credit_score','annual_income_to_insurance_duration','annual_income_to_age','annual_income_to_health_score',\n 'credit_score_to_dependent','annual_income_to_policy_age','credit_score_to_previous_claims','annual_income_to_previous_claims','dependents_and_previous_claims'\n]\n\ncat_cols = ['Gender','Marital_Status','Occupation','Location','Smoking_Status','Property_Type']\n\nnum_cols = ['Age','Number_of_Dependents','Health_Score','Previous_Claims','Vehicle_Age','Credit_Score','Insurance_Duration','annual_income_log','policy_start_year','policy_start_month','policy_start_day','policy_age',\n            'annual_income_to_dependent','health_and_credit_score','annual_income_to_insurance_duration','annual_income_to_age','annual_income_to_health_score','credit_score_to_dependent',\n            'annual_income_to_policy_age','credit_score_to_previous_claims','annual_income_to_previous_claims','dependents_and_previous_claims'\n           ]\n\nord_cols = {\n    'Education_Level': ['High School', 'Bachelor\\'s', 'Master\\'s', 'PhD'],\n    'Policy_Type': ['Basic','Comprehensive','Premium'],\n    'Customer_Feedback': ['unknown','Poor','Average','Good'],\n    'Exercise_Frequency':['Rarely','Monthly','Weekly','Daily'],\n    'policy_age_bins':['new','old'],\n    'annual_income_bins':['no-data','mid','high'],\n    'credit_score_bins':['S','M','L','XL','no-data'],\n    'health_score_bins':['S','M','L'],\n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T17:15:25.740787Z","iopub.execute_input":"2024-12-31T17:15:25.741102Z","iopub.status.idle":"2024-12-31T17:15:25.745085Z","shell.execute_reply.started":"2024-12-31T17:15:25.74107Z","shell.execute_reply":"2024-12-31T17:15:25.744441Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### OneHot Encoding","metadata":{}},{"cell_type":"code","source":"# 1. One-hot encode categorical features\n\n# Initialize OneHotEncoder\nohe = OneHotEncoder(sparse=False, drop=None)  # Set `sparse=False` to return a dense array\n\n# Create a new DataFrame with one-hot encoded columns\ndf_cat = pd.DataFrame(ohe.fit_transform(df[cat_cols]), columns=ohe.get_feature_names_out(cat_cols))\n\ndf_cat.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T17:15:25.780831Z","iopub.execute_input":"2024-12-31T17:15:25.781003Z","iopub.status.idle":"2024-12-31T17:15:28.766971Z","shell.execute_reply.started":"2024-12-31T17:15:25.780988Z","shell.execute_reply":"2024-12-31T17:15:28.766257Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Ordinal Encoding","metadata":{}},{"cell_type":"code","source":"# 2. Ordinal encode ordinal features (if necessary)\nencoder = OrdinalEncoder(categories=[ord_cols[col] for col in ord_cols])\n# Select columns to encode and apply the encoder\ndf_ord = pd.DataFrame(encoder.fit_transform(df[list(ord_cols.keys())]),columns=[list(ord_cols.keys())])\n\nsc_ord = StandardScaler()  # You can also use StandardScaler() depending on the need\ndf_ord = pd.DataFrame(sc_ord.fit_transform(df_ord), columns=df_ord.columns)\n\ndf.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T17:15:28.768737Z","iopub.execute_input":"2024-12-31T17:15:28.769033Z","iopub.status.idle":"2024-12-31T17:15:32.184077Z","shell.execute_reply.started":"2024-12-31T17:15:28.769003Z","shell.execute_reply":"2024-12-31T17:15:32.183227Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Standard Scalling","metadata":{}},{"cell_type":"code","source":"\nnum_sc = StandardScaler()\ndf_num = pd.DataFrame(num_sc.fit_transform(df[num_cols]),columns=num_cols)\ndf_num.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T17:15:32.185085Z","iopub.execute_input":"2024-12-31T17:15:32.185342Z","iopub.status.idle":"2024-12-31T17:15:33.141609Z","shell.execute_reply.started":"2024-12-31T17:15:32.185321Z","shell.execute_reply":"2024-12-31T17:15:33.140524Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\ntgt_sc = StandardScaler()\ndf_tgt = pd.DataFrame(tgt_sc.fit_transform(df['premium_amount_log'].values.reshape(-1, 1)),columns=['premium_amount_log_scalled'])\ndf_tgt.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T17:15:33.142461Z","iopub.execute_input":"2024-12-31T17:15:33.142687Z","iopub.status.idle":"2024-12-31T17:15:33.170287Z","shell.execute_reply.started":"2024-12-31T17:15:33.142667Z","shell.execute_reply":"2024-12-31T17:15:33.169341Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Pre-processed dataset","metadata":{}},{"cell_type":"code","source":"\ndf_trn = pd.concat([df_cat, df_ord, df_num, df_tgt, df[['src','id','premium_amount_log']]], axis=1)\ndf_trn.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T17:15:33.171094Z","iopub.execute_input":"2024-12-31T17:15:33.171331Z","iopub.status.idle":"2024-12-31T17:15:33.873317Z","shell.execute_reply.started":"2024-12-31T17:15:33.171311Z","shell.execute_reply":"2024-12-31T17:15:33.872537Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Data splitting","metadata":{}},{"cell_type":"code","source":"\n# Splitting features and target\nX = df_trn.loc[df_trn['src']=='trn'].drop(['src','id','premium_amount_log','premium_amount_log_scalled'], axis=1)\ny = df_trn.loc[df_trn['src']=='trn',[\"premium_amount_log_scalled\"]]\n\n# Split data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T17:15:33.874005Z","iopub.execute_input":"2024-12-31T17:15:33.874226Z","iopub.status.idle":"2024-12-31T17:15:35.012418Z","shell.execute_reply.started":"2024-12-31T17:15:33.874208Z","shell.execute_reply":"2024-12-31T17:15:35.011503Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Neural Network training","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras import layers\n\n# Check if GPU is available\nprint(\"Num GPUs Available: \", len(tf.config.experimental.list_physical_devices('GPU')))\n\nmodel = tf.keras.Sequential([\n    layers.InputLayer(shape=(X_train.shape[1],)),  # Number of input features\n    layers.Dense(64, activation='relu'),  # Hidden layer with 64 neurons\n    layers.Dense(64, activation='relu'),  # Hidden layer with 64 neurons\n    layers.Dense(64, activation='relu'),  # Hidden layer with 64 neurons\n    layers.Dense(64, activation='relu'),  # Hidden layer with 64 neurons\n    layers.Dense(64, activation='relu'),  # Hidden layer with 64 neurons\n    layers.Dense(1)  # Output layer for regression (1 neuron)\n])\n\n# Compile the model with mean squared error loss and an optimizer\nmodel.compile(optimizer='adam', loss='mse')\n\nearly_stopping = tf.keras.callbacks.EarlyStopping(\n    monitor='val_loss',  # Metric to monitor (validation loss)\n    patience=5,  # Number of epochs to wait for improvement\n    restore_best_weights=True,  # Restore weights of the best epoch\n    verbose=1  # Print messages when training is stopped\n)\n\n# Train the model\nwith tf.device('/GPU:0'):\n    _ = model.fit(X_train.values, y_train.values, epochs=5, batch_size=64,  validation_split=0.2, callbacks=[early_stopping], verbose=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T17:15:35.013273Z","iopub.execute_input":"2024-12-31T17:15:35.013582Z","iopub.status.idle":"2024-12-31T17:17:20.181157Z","shell.execute_reply.started":"2024-12-31T17:15:35.01355Z","shell.execute_reply":"2024-12-31T17:17:20.180493Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Evaluation","metadata":{}},{"cell_type":"code","source":"\n# Predict on the test set\ny_pred = model.predict(X_test.values)\nrmsle = mean_squared_error(tgt_sc.inverse_transform(y_test), tgt_sc.inverse_transform(y_pred),squared=False)\nprint(f'\\n RMSLE : {rmsle}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T16:48:37.388897Z","iopub.execute_input":"2024-12-31T16:48:37.389824Z","iopub.status.idle":"2024-12-31T16:48:50.606598Z","shell.execute_reply.started":"2024-12-31T16:48:37.389787Z","shell.execute_reply":"2024-12-31T16:48:50.605785Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Submission","metadata":{}},{"cell_type":"code","source":"\ndf['premium_amount_log_nn'] = tgt_sc.inverse_transform(model.predict(df_trn.drop(['src','id','premium_amount_log','premium_amount_log_scalled'], axis=1).values))\ndf[['id','src','Premium_Amount','premium_amount_log','premium_amount_log_nn']].to_csv('/kaggle/working/nn.csv', index=False)\n\nsubmission_df = df.loc[df['src']=='tst',['id','Premium_Amount']]\nsubmission_df['Premium Amount'] = np.expm1(df.loc[df['src']=='tst','premium_amount_log_nn']).round(decimals=3)\n\nsubmission_df.to_csv('submission.csv', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-31T17:02:24.060248Z","iopub.execute_input":"2024-12-31T17:02:24.060573Z","iopub.status.idle":"2024-12-31T17:04:09.086291Z","shell.execute_reply.started":"2024-12-31T17:02:24.06055Z","shell.execute_reply":"2024-12-31T17:04:09.085373Z"}},"outputs":[],"execution_count":null}]}