{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Published on December 01, 2024. By Prata, Marília (mpwolke)","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.graph_objs as go\nimport plotly.offline as py\nimport plotly.express as px\n\n#Ignore warnings\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-12-01T00:54:33.213417Z","iopub.execute_input":"2024-12-01T00:54:33.213827Z","iopub.status.idle":"2024-12-01T00:54:37.043633Z","shell.execute_reply.started":"2024-12-01T00:54:33.21377Z","shell.execute_reply":"2024-12-01T00:54:37.042381Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Competition Citation:\n\n@misc{playground-series-s4e12,\n\r\n    author = {Walter Reade and Elizabeth Park}\n    ,\r\n    title = {Regression with an Insurance Dataset\n    },\r\n    year = {2024},\r\n    howpublished = {\\url{https://kaggle.com/competitions/playground-series-s4e12}},\r\n    note = {Kaggle}\r\n}","metadata":{}},{"cell_type":"markdown","source":"## Optimizing insurance risk assessment\n\nOptimizing insurance risk assessment: a regression model based on a risk-loaded approach\n\nAuthors: Zinoviy Landsman and Tomer Shushi\n\nPublished online by Cambridge University Press:  31 May 2024\n\n\"In this paper, the authors proposed a hybrid risk-loaded regression model that simultaneously minimizes the classical least squared loss function and penalty loss function of suggested loss intensity of the problem, being the sum of expected squared historical data of responses.\"\n\n\"Imposing rather general conditions on the model and the loss function, they found that the explicit solution of the minimization problem takes a proportional form of the least squared estimator, where the coefficient of proportionality depends on response vector  and design matrix. Special attention was given to the powered penalty loss function, which essentially generalized the classical case of identical loss function.\"\n\n\"In addition to the unconditional problem, the authors also considered a situation in which the solution satisfies the linear equality constraints. In that case, they also found the explicit closed-form solution. In both the unconditional and conditional cases, they demonstrated that the ratio of empirical intensities for classical least squared estimator and for hybrid risk loaded estimator proposed in the paper is always less than 1, which speaks in favor of the proposed method.\"\n\nhttps://www.cambridge.org/core/journals/annals-of-actuarial-science/article/optimizing-insurance-risk-assessment-a-regression-model-based-on-a-riskloaded-approach/3196BD763DDF261223A0FB1A9898F4A8","metadata":{}},{"cell_type":"code","source":"#We have 1200000 rows\n\ndf = pd.read_csv('/kaggle/input/playground-series-s4e12/train.csv', nrows=1000)\ndf.tail()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T01:59:05.411058Z","iopub.execute_input":"2024-12-01T01:59:05.411543Z","iopub.status.idle":"2024-12-01T01:59:05.446773Z","shell.execute_reply.started":"2024-12-01T01:59:05.411505Z","shell.execute_reply":"2024-12-01T01:59:05.445574Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"The (Row,Column) is:\\n\", df.shape)\n    \nprint(\"Data type of each column:\\n\", df.dtypes)\n    \nprint(\"The number of null values in each column are:\\n\", df.isnull().sum())\n    \nprint(\"Numeric summary:\\n\", df.describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T01:15:31.222783Z","iopub.execute_input":"2024-12-01T01:15:31.223242Z","iopub.status.idle":"2024-12-01T01:15:32.632398Z","shell.execute_reply.started":"2024-12-01T01:15:31.223209Z","shell.execute_reply":"2024-12-01T01:15:32.630906Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"g = sns.catplot(x=\"Smoking Status\", y=\"Premium Amount\",col_wrap=3, col=\"Gender\",data= df, kind=\"box\",height=5, aspect=0.8);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T01:35:12.49035Z","iopub.execute_input":"2024-12-01T01:35:12.49079Z","iopub.status.idle":"2024-12-01T01:35:15.994178Z","shell.execute_reply.started":"2024-12-01T01:35:12.490751Z","shell.execute_reply":"2024-12-01T01:35:15.992898Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### It seems that smoking status by gender is quite the same.\n\nThe whiskers:\n\r\nThe lines coming out from each box extend from the maximum to the minimum values of each set. Together with the box, the whiskers show how big a range there is between those two extremes. Larger ranges indicate wider distribution, that is, more scattered data\n\nWider ranges (whisker length, box size) indicate more variable data.\n\nhttps://bioturing.medium.com/how-to-compare-box-plots-3da8e2adfa8f.","metadata":{}},{"cell_type":"code","source":"g = sns.catplot(x=\"Smoking Status\", y=\"Age\",col_wrap=3, col=\"Gender\",data= df, kind=\"box\",height=5, aspect=0.8);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T01:30:45.098185Z","iopub.execute_input":"2024-12-01T01:30:45.098638Z","iopub.status.idle":"2024-12-01T01:30:48.357911Z","shell.execute_reply.started":"2024-12-01T01:30:45.098601Z","shell.execute_reply":"2024-12-01T01:30:48.356549Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### It's a sign of gender similarity on smoking status (behaviour)","metadata":{}},{"cell_type":"code","source":"X = [\"Age\", \"Gender\", \"Annual Income\",\"Education Level\",\"Occupation\",\"Health Score\",\"Premium Amount\"]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T02:01:18.728695Z","iopub.execute_input":"2024-12-01T02:01:18.729773Z","iopub.status.idle":"2024-12-01T02:01:18.735198Z","shell.execute_reply.started":"2024-12-01T02:01:18.729734Z","shell.execute_reply":"2024-12-01T02:01:18.733643Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#That was to work with Elbow Method though it failed. However, I need it to the pairplot.\n\ndata = df[X]\ny = df[\"Premium Amount\"]\ndata.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T02:06:00.935965Z","iopub.execute_input":"2024-12-01T02:06:00.936345Z","iopub.status.idle":"2024-12-01T02:06:00.951503Z","shell.execute_reply.started":"2024-12-01T02:06:00.936312Z","shell.execute_reply":"2024-12-01T02:06:00.950308Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sns.pairplot(data, hue = \"Premium Amount\");","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T02:07:24.893767Z","iopub.execute_input":"2024-12-01T02:07:24.89422Z","iopub.status.idle":"2024-12-01T02:07:51.082126Z","shell.execute_reply.started":"2024-12-01T02:07:24.894182Z","shell.execute_reply":"2024-12-01T02:07:51.081004Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Input X contains NaN.  I removed them. Unfortunately, the issue persisted.\n\ndf.isnull().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T02:22:48.02401Z","iopub.execute_input":"2024-12-01T02:22:48.024914Z","iopub.status.idle":"2024-12-01T02:22:48.037469Z","shell.execute_reply.started":"2024-12-01T02:22:48.024871Z","shell.execute_reply":"2024-12-01T02:22:48.036279Z"},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"cat=df.select_dtypes(\"object\")\nfor column in cat:\n    df[column].fillna(df[column].mode()[0], inplace=True)\n    #dataset[column].fillna(\"NA\", inplace=True)\n\n\nfl=df.select_dtypes([\"float64\",\"int64\"]).drop(\"Premium Amount\",axis=1)\nfor column in fl:\n    df[column].fillna(df[column].median(), inplace=True)\n    #dataset[column].fillna(0, inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T02:32:23.371925Z","iopub.execute_input":"2024-12-01T02:32:23.372308Z","iopub.status.idle":"2024-12-01T02:32:23.393054Z","shell.execute_reply.started":"2024-12-01T02:32:23.372276Z","shell.execute_reply":"2024-12-01T02:32:23.391703Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.preprocessing import PolynomialFeatures\nfrom scipy import stats\nimport statsmodels.api as sm\nfrom statsmodels.formula.api import ols","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T02:38:53.345337Z","iopub.execute_input":"2024-12-01T02:38:53.345721Z","iopub.status.idle":"2024-12-01T02:38:54.932863Z","shell.execute_reply.started":"2024-12-01T02:38:53.345687Z","shell.execute_reply":"2024-12-01T02:38:54.931546Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#Label Encoder","metadata":{}},{"cell_type":"code","source":"#https://www.kaggle.com/code/mpwolke/engagement-with-boruta-feature-selection\n\nfrom sklearn.preprocessing import LabelEncoder\n\n#fill in mean for floats\nfor c in df.columns:\n    if df[c].dtype=='float16' or  df[c].dtype=='float32' or  df[c].dtype=='float64':\n        df[c].fillna(df[c].mean())\n\n#fill in -999 for categoricals\ndf = df.fillna(-999)\n# Label Encoding\nfor f in df.columns:\n    if df[f].dtype=='object': \n        lbl = LabelEncoder()\n        lbl.fit(list(df[f].values))\n        df[f] = lbl.transform(list(df[f].values))\n        \nprint('Labelling done.')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T02:53:20.166694Z","iopub.execute_input":"2024-12-01T02:53:20.167754Z","iopub.status.idle":"2024-12-01T02:53:20.200375Z","shell.execute_reply.started":"2024-12-01T02:53:20.167712Z","shell.execute_reply":"2024-12-01T02:53:20.199167Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#Performing a linear regression","metadata":{}},{"cell_type":"code","source":"#https://www.kaggle.com/code/kianwee/linear-regression-insurance-dataset\n\n# Creating training and testing dataset\ny = df['Premium Amount']\nX = df.drop(['Premium Amount'], axis = 1)\n\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.20, random_state = 42)\n\nlr = LinearRegression().fit(X_train,y_train)\ny_train_pred = lr.predict(X_train)\ny_test_pred = lr.predict(X_test)\n\nprint(lr.score(X_test,y_test))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T02:53:30.124782Z","iopub.execute_input":"2024-12-01T02:53:30.125472Z","iopub.status.idle":"2024-12-01T02:53:30.191781Z","shell.execute_reply.started":"2024-12-01T02:53:30.125434Z","shell.execute_reply":"2024-12-01T02:53:30.18919Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Shame on me. That score is ridiculous.","metadata":{}},{"cell_type":"code","source":"dfcorr=df.corr()\ndfcorr\nplt.figure(figsize=(10,4))\nsns.heatmap(df.corr(),annot=False,cmap='summer')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T02:57:14.402966Z","iopub.execute_input":"2024-12-01T02:57:14.403355Z","iopub.status.idle":"2024-12-01T02:57:14.886952Z","shell.execute_reply.started":"2024-12-01T02:57:14.403322Z","shell.execute_reply":"2024-12-01T02:57:14.885813Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Try removing non-correlated features","metadata":{}},{"cell_type":"code","source":"# Creating training and testing dataset\ny = df['Premium Amount']\nX = df.drop(['Premium Amount', 'Location'], axis = 1)\n\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.20, random_state = 42)\n\nlr = LinearRegression().fit(X_train,y_train)\ny_train_pred = lr.predict(X_train)\ny_test_pred = lr.predict(X_test)\n\nprint(lr.score(X_test,y_test))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T03:00:01.273293Z","iopub.execute_input":"2024-12-01T03:00:01.2737Z","iopub.status.idle":"2024-12-01T03:00:01.319213Z","shell.execute_reply.started":"2024-12-01T03:00:01.273664Z","shell.execute_reply":"2024-12-01T03:00:01.315698Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## No comments. Even worst.","metadata":{}},{"cell_type":"markdown","source":"![](https://pbs.twimg.com/media/DcHQHRvX0AAu9sk.jpg)","metadata":{}},{"cell_type":"code","source":"# Creating training and testing dataset\ny = df['Premium Amount']\nX = df.drop(['Premium Amount', 'Location'], axis = 1)\n\npoly_reg  = PolynomialFeatures(degree=2)\nX_poly = poly_reg.fit_transform(X)\n\nX_train,X_test,y_train,y_test = train_test_split(X_poly,y,test_size=0.20, random_state = 42)\n\nlin_reg = LinearRegression()\nlin_reg  = lin_reg.fit(X_train,y_train)\n\nprint(lin_reg.score(X_test,y_test))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T03:02:07.223436Z","iopub.execute_input":"2024-12-01T03:02:07.223879Z","iopub.status.idle":"2024-12-01T03:02:07.276412Z","shell.execute_reply.started":"2024-12-01T03:02:07.223842Z","shell.execute_reply":"2024-12-01T03:02:07.274892Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### That sucks. No insurance Premium.\n\nI was about to work with clusters and Elbow Method, however the imputation didn't work no matter it shows zeros on Nan values. Then I jumped to Plan B and I jinxed the code.","metadata":{}},{"cell_type":"markdown","source":"## In fact, this code is dead. After 2h:24m, time of death: 00:17 December 01, 2024. \n\n![](https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcRAyiJ_NZKqMN4b_krSrzxnNJD4HwNFBhUm2-i5e7jtuZgDcNVVJFl0_tPgjItF03kGAbg&usqp=CAU)","metadata":{}},{"cell_type":"markdown","source":"#Acknowledgements:\n\nMichael Scott https://www.kaggle.com/code/kianwee/linear-regression-insurance-dataset","metadata":{}}]}