{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30804,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# How To Average Numbers\nIn this notebook we compare 3 different ways to \"average\" numbers. From school, we remember that the \"middle\" of a set of numbers can be either the `mean` or `median`. However, there are many more \"middles\" for example we can take `arthmetic mean` `(a+b)/2` or `geometric mean` `sqrt(a*b)`. There is also the `exponented mean log` which is `exp( (log(a)+log(b))/2 )`. All of these find \"middles\" between numbers.\n\nIn this competition, the best \"middle\" is to use the `exponented mean log 1p` because this competition metric is `RMSLE`. Therefore when creating ensembles where we average model targets or other times when we average target numbers we should consider the `exponented mean log 1p`.\n\nIn this notebook, we demonstrate the CV score using 3 different middles. Namely we compute CV for `median`, `mean`, and `exponented mean log 1p`. We see that `exponented mean log 1p` produces the best CV score and LB score! There is a discussion about this [here][1]\n\n[1]: https://www.kaggle.com/competitions/playground-series-s4e12/discussion/549909","metadata":{}},{"cell_type":"code","source":"import pandas as pd, numpy as np\n\ntrain = pd.read_csv(\"/kaggle/input/playground-series-s4e12/train.csv\")\nprint(\"Train shape:\",train.shape)\ntrain.head()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-04T13:46:57.754878Z","iopub.execute_input":"2024-12-04T13:46:57.755315Z","iopub.status.idle":"2024-12-04T13:47:02.443343Z","shell.execute_reply.started":"2024-12-04T13:46:57.755263Z","shell.execute_reply":"2024-12-04T13:47:02.442229Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Define Metric RMSLE\nThe competition metric is `Root Mean Squared Log Error`.","metadata":{}},{"cell_type":"code","source":"def RMSLE(true,pred):\n    true_log = np.log1p(true)\n    pred_log = np.log1p(pred)\n    m = np.sqrt(np.mean( (true_log-pred_log)**2.0 ))\n    return m","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T13:47:02.445518Z","iopub.execute_input":"2024-12-04T13:47:02.445956Z","iopub.status.idle":"2024-12-04T13:47:02.452069Z","shell.execute_reply.started":"2024-12-04T13:47:02.445905Z","shell.execute_reply":"2024-12-04T13:47:02.450933Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Try MAE's Median\nWhen metric is `MAE` using `medians` is preferred.","metadata":{}},{"cell_type":"code","source":"pred = np.median( train[\"Premium Amount\"] )\nm = RMSLE(train[\"Premium Amount\"].values, pred)\nprint(f\"Median produces CV RMSLE = {m}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T13:47:02.453442Z","iopub.execute_input":"2024-12-04T13:47:02.453802Z","iopub.status.idle":"2024-12-04T13:47:02.514193Z","shell.execute_reply.started":"2024-12-04T13:47:02.453746Z","shell.execute_reply":"2024-12-04T13:47:02.513074Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Try MSE/RMSE's Mean\nWhen metric is `MSE or RMSE` using `means` is preferred.","metadata":{}},{"cell_type":"code","source":"pred = np.mean( train[\"Premium Amount\"] )\nm = RMSLE(train[\"Premium Amount\"].values, pred)\nprint(f\"Mean produces CV RMSLE = {m}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T13:47:02.516568Z","iopub.execute_input":"2024-12-04T13:47:02.517036Z","iopub.status.idle":"2024-12-04T13:47:02.55438Z","shell.execute_reply.started":"2024-12-04T13:47:02.516986Z","shell.execute_reply":"2024-12-04T13:47:02.553085Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Try RMSLE's Mean Log!\nWhen metric is `RMSLE` using `exponented mean log 1p` is preferred.","metadata":{}},{"cell_type":"code","source":"pred = np.exp( np.mean( np.log1p(train[\"Premium Amount\"]) ) )-1\nm = RMSLE(train[\"Premium Amount\"].values, pred)\nprint(f\"Exponented Mean Log 1p produces CV RMSLE = {m}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T13:47:02.555685Z","iopub.execute_input":"2024-12-04T13:47:02.556048Z","iopub.status.idle":"2024-12-04T13:47:02.615171Z","shell.execute_reply.started":"2024-12-04T13:47:02.556013Z","shell.execute_reply":"2024-12-04T13:47:02.614056Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Create Submission with Mean Log!\nLet's make a single number submission using `exponented mean log 1p`. This will be a CV and LB baseline. Any model that isn't beating this CV and LB is doing poorly in this competition.","metadata":{}},{"cell_type":"code","source":"pred = np.exp( np.mean( np.log1p(train[\"Premium Amount\"]) ) )-1\nsub = pd.read_csv(\"/kaggle/input/playground-series-s4e12/sample_submission.csv\")\nsub[\"Premium Amount\"] = pred\nsub.to_csv(\"submission.csv\",index=False)\nprint(\"Sub shape:\",sub.shape)\nsub.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-04T13:47:02.616608Z","iopub.execute_input":"2024-12-04T13:47:02.617039Z","iopub.status.idle":"2024-12-04T13:47:04.462861Z","shell.execute_reply.started":"2024-12-04T13:47:02.616992Z","shell.execute_reply":"2024-12-04T13:47:04.461514Z"}},"outputs":[],"execution_count":null}]}