{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84896,"databundleVersionId":10305135,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Kaggle Tabular Playground Series Season 4 Eps 12: The data Exploration","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport polars as pl\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-02T01:50:07.534253Z","iopub.execute_input":"2024-12-02T01:50:07.534826Z","iopub.status.idle":"2024-12-02T01:50:09.196664Z","shell.execute_reply.started":"2024-12-02T01:50:07.534737Z","shell.execute_reply":"2024-12-02T01:50:09.195376Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df = pl.read_csv('/kaggle/input/playground-series-s4e12/train.csv')\ndf.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T01:50:15.064437Z","iopub.execute_input":"2024-12-02T01:50:15.065009Z","iopub.status.idle":"2024-12-02T01:50:16.206072Z","shell.execute_reply.started":"2024-12-02T01:50:15.064969Z","shell.execute_reply":"2024-12-02T01:50:16.204865Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<font size=\"3\">Here is a brief explanation of what each variable represents:\n\n1. Age: Numeric; the age of the insurance holder \n2. Gender: Categorical; the gender of the insurance holder\n3. Annual Income: The income of the insurance holder per year. There is no information on the currency unit\n4. Marital Status: Categorical; the marital status of the insurance holder\n5. Number of Dependents: Numeric; the total people the insurance holder needs to support financially.\n6. Education level: Categorical; the highest degree the insurance holder obtain.\n7. Occupation: Categorical; the employment status of the insurance holder.\n8. Health Score: Numeric; the score represents how healthy the insurance holder is.\n9. Location: Categorical; the environmental type where the insurance holder lives.\n10. Policy Type: Categorical; the insurance type each person has.\n11. Previous Claims: Categorical (Numeric); the total claims previously done.\n12. Vehicle Age: Numeric; the age of the insurance holder's vehicle.\n13. Credit Score: Numeric; the insurance holder's credit score.\n14. Insurance Duration: Numeric; the total years an insurance holder has had the insurance.\n15. Policy Start Date: Datetime; the starting date (day 1) when an insurance holder first buy the insurance.\n16. Customer Feedback: Categorical; the insurance holder's satisfaction.\n17. Smoking Status: Categorical; whether the insurance holder is smoking.\n18. Exercise Frequency: Categorical; how often the insurance holder does exercise.\n19. Property Type: Categorical; the type of the property the insurance holder has.\n20. Premium Amount: Numeric; the amount of money the insurance holder needs to pay (usually annually, even though there is no explicit information on this dataset). This is the target feature.</font>\n\n<font size=\"3\"> The table below shows the total missing values for each column in this dataset.</font>","metadata":{}},{"cell_type":"code","source":"print(f\"The shape of the dataframe: {df.shape}\")\ndf.null_count()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T01:53:00.776578Z","iopub.execute_input":"2024-12-02T01:53:00.777391Z","iopub.status.idle":"2024-12-02T01:53:00.786196Z","shell.execute_reply.started":"2024-12-02T01:53:00.777345Z","shell.execute_reply":"2024-12-02T01:53:00.784999Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<font size=\"3\"> If we consider using only the data without missing values, than we would have only a third of the original data's size.</font>.","metadata":{}},{"cell_type":"code","source":"df.drop_nulls().shape #drop any row missing any data, even that subject has only 1 missing value","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T01:54:53.527468Z","iopub.execute_input":"2024-12-02T01:54:53.527934Z","iopub.status.idle":"2024-12-02T01:54:53.658459Z","shell.execute_reply.started":"2024-12-02T01:54:53.527896Z","shell.execute_reply":"2024-12-02T01:54:53.657352Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# we will fill the NAs later\ndf_filled = pl.read_csv('/kaggle/input/playground-series-s4e12/train.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T17:00:48.65252Z","iopub.execute_input":"2024-12-01T17:00:48.653375Z","iopub.status.idle":"2024-12-01T17:00:49.355553Z","shell.execute_reply.started":"2024-12-01T17:00:48.653335Z","shell.execute_reply":"2024-12-01T17:00:49.354361Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Numerical Values\n<font size=\"3\">Let's begin exploring the dataset. I'll start with the numerical data, and then the categorical ones.</font>","metadata":{},"attachments":{}},{"cell_type":"markdown","source":"## First, the target 🎯\n<font size=\"3\">The following is the histogram of `Premium Amount`. From this histogram, we can see it has multimodal, indicating that **there might be a mixture of distributions.** The first mode is around 100 and it has the highest frequency. The second mode is around 700, and the third mode is a bit higher than 2000. As the mode goes higher, the frequency goes lower.</font>","metadata":{}},{"cell_type":"code","source":"sns.histplot(df['Premium Amount'], bins=100, color='#377ec8');","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:06:08.736403Z","iopub.execute_input":"2024-12-02T02:06:08.736846Z","iopub.status.idle":"2024-12-02T02:06:10.118823Z","shell.execute_reply.started":"2024-12-02T02:06:08.736805Z","shell.execute_reply":"2024-12-02T02:06:10.117406Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Age 🎂\n<font size=\"3\"> The age ranges from 18 to 64, and it has roughly uniform distribution.</font>","metadata":{},"attachments":{}},{"cell_type":"code","source":"df['Age'].describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:09:26.243219Z","iopub.execute_input":"2024-12-02T02:09:26.243856Z","iopub.status.idle":"2024-12-02T02:09:26.325816Z","shell.execute_reply.started":"2024-12-02T02:09:26.243757Z","shell.execute_reply":"2024-12-02T02:09:26.323869Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"_, axs = plt.subplots(1, 2, figsize=(12,6))\nsns.histplot(df['Age'], discrete=True, color='#377ec8', ax=axs[0])\nsns.boxplot(data=df['Age'], color='#7EC837', ax=axs[1]);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:10:36.204786Z","iopub.execute_input":"2024-12-02T02:10:36.205191Z","iopub.status.idle":"2024-12-02T02:10:37.642364Z","shell.execute_reply.started":"2024-12-02T02:10:36.205156Z","shell.execute_reply":"2024-12-02T02:10:37.641252Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Annual Income 💰\n<font size=\"3\"> As we can see from the histogram, this feature has a heavily right-skewed distribution. The histogram excludes the 3% missing data.</font>","metadata":{}},{"cell_type":"code","source":"df['Annual Income'].describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:14:19.306568Z","iopub.execute_input":"2024-12-02T02:14:19.307028Z","iopub.status.idle":"2024-12-02T02:14:19.369845Z","shell.execute_reply.started":"2024-12-02T02:14:19.306987Z","shell.execute_reply":"2024-12-02T02:14:19.368313Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"_, axs = plt.subplots(1, 2, figsize=(12,6))\nsns.histplot(df['Annual Income'], bins=100, color='#377ec8', ax=axs[0])\nsns.boxplot(data=df['Annual Income'], color='#7ec837', ax=axs[1]);","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-12-01T17:12:46.038077Z","iopub.execute_input":"2024-12-01T17:12:46.038494Z","iopub.status.idle":"2024-12-01T17:12:47.734641Z","shell.execute_reply.started":"2024-12-01T17:12:46.038457Z","shell.execute_reply":"2024-12-01T17:12:47.733633Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Number of Dependents 👶👴👵\n<font size=\"3\">","metadata":{},"attachments":{}},{"cell_type":"code","source":"df['Number of Dependents'].describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:29:09.224588Z","iopub.execute_input":"2024-12-02T02:29:09.225906Z","iopub.status.idle":"2024-12-02T02:29:09.279084Z","shell.execute_reply.started":"2024-12-02T02:29:09.225841Z","shell.execute_reply":"2024-12-02T02:29:09.277715Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"_, axs = plt.subplots(1, 2, figsize=(12,6))\nsns.histplot(df['Number of Dependents'], discrete=True, color='#377ec8', ax=axs[0])\nsns.boxenplot(data=df['Number of Dependents'], color='#7ec837', ax=axs[1]);","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-12-01T17:14:01.988008Z","iopub.execute_input":"2024-12-01T17:14:01.988446Z","iopub.status.idle":"2024-12-01T17:14:03.532603Z","shell.execute_reply.started":"2024-12-01T17:14:01.988376Z","shell.execute_reply":"2024-12-01T17:14:03.531648Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Health score 💪🏻","metadata":{}},{"cell_type":"code","source":"df['Health Score'].describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:32:06.42397Z","iopub.execute_input":"2024-12-02T02:32:06.424392Z","iopub.status.idle":"2024-12-02T02:32:06.495412Z","shell.execute_reply.started":"2024-12-02T02:32:06.424354Z","shell.execute_reply":"2024-12-02T02:32:06.49414Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"_, axs = plt.subplots(1, 2, figsize=(12,6))\nsns.histplot(df['Health Score'], color='#377ec8', bins=75, ax=axs[0])\naxs[0].set_title(f\"mean: {np.round(df['Health Score'].mean(), 2)}, variance: {np.round(df['Health Score'].var(), 2)}\")\nsns.boxenplot(data=df['Health Score'], color='#7ec837', ax=axs[1]);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:32:13.481294Z","iopub.execute_input":"2024-12-02T02:32:13.481699Z","iopub.status.idle":"2024-12-02T02:32:15.36505Z","shell.execute_reply.started":"2024-12-02T02:32:13.48166Z","shell.execute_reply":"2024-12-02T02:32:15.363848Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Previous Claims 📑\n<font size=\"3\">The majority of the data have either zero or one previous claims, and even more don't provide this information. There are some outliers, too, with very few have 6 or more claims.</font>","metadata":{}},{"cell_type":"code","source":"df['Previous Claims'].describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:39:07.392751Z","iopub.execute_input":"2024-12-02T02:39:07.393292Z","iopub.status.idle":"2024-12-02T02:39:07.440988Z","shell.execute_reply.started":"2024-12-02T02:39:07.393241Z","shell.execute_reply":"2024-12-02T02:39:07.439845Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"_, axs = plt.subplots(1, 2, figsize=(12,6))\nsns.histplot(df['Previous Claims'], discrete=True, color='#377ec8', ax=axs[0])\nsns.boxplot(data=df['Previous Claims'], color='#7ec837', ax=axs[1]);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:35:16.534187Z","iopub.execute_input":"2024-12-02T02:35:16.534611Z","iopub.status.idle":"2024-12-02T02:35:17.915483Z","shell.execute_reply.started":"2024-12-02T02:35:16.534571Z","shell.execute_reply":"2024-12-02T02:35:17.914054Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<font size=\"3\"> This is the histogram and boxplot of the variable, including NAs.</font>","metadata":{}},{"cell_type":"code","source":"_, axs = plt.subplots(1, 2, figsize=(12,6))\nsns.histplot(df['Previous Claims'].fill_null(-1), discrete=True, color='#377ec8', ax=axs[0])\nsns.boxplot(data=df['Previous Claims'].fill_null(-1), color='#7ec837', ax=axs[1]);","metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-12-02T02:36:24.35703Z","iopub.execute_input":"2024-12-02T02:36:24.357417Z","iopub.status.idle":"2024-12-02T02:36:25.588309Z","shell.execute_reply.started":"2024-12-02T02:36:24.35738Z","shell.execute_reply":"2024-12-02T02:36:25.586887Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Vehicle Age 🚗🛵\n","metadata":{}},{"cell_type":"code","source":"df['Vehicle Age'].describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:39:55.641789Z","iopub.execute_input":"2024-12-02T02:39:55.642198Z","iopub.status.idle":"2024-12-02T02:39:55.688659Z","shell.execute_reply.started":"2024-12-02T02:39:55.642159Z","shell.execute_reply":"2024-12-02T02:39:55.687315Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"_, axs = plt.subplots(1, 2, figsize=(12,6))\nsns.histplot(df['Vehicle Age'], discrete=True, color='#377ec8', ax=axs[0])\nsns.boxenplot(data=df['Vehicle Age'], color='#7cc837', ax=axs[1]);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-01T17:24:53.431082Z","iopub.execute_input":"2024-12-01T17:24:53.432041Z","iopub.status.idle":"2024-12-01T17:24:55.013081Z","shell.execute_reply.started":"2024-12-01T17:24:53.431997Z","shell.execute_reply":"2024-12-01T17:24:55.011924Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Credit Score 💯\n<font size=\"3\"> The credit score is slightly left-skewed.</font>","metadata":{}},{"cell_type":"code","source":"df['Credit Score'].describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:41:35.481557Z","iopub.execute_input":"2024-12-02T02:41:35.482Z","iopub.status.idle":"2024-12-02T02:41:35.544132Z","shell.execute_reply.started":"2024-12-02T02:41:35.481958Z","shell.execute_reply":"2024-12-02T02:41:35.54216Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"_, axs = plt.subplots(1,2, figsize=(12,6))\nsns.histplot(df['Credit Score'], bins=50, color='#377ec8', ax=axs[0])\nsns.boxplot(data=df['Credit Score'], color='#7cc837', ax=axs[1]);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:41:49.155275Z","iopub.execute_input":"2024-12-02T02:41:49.155667Z","iopub.status.idle":"2024-12-02T02:41:50.59398Z","shell.execute_reply.started":"2024-12-02T02:41:49.155631Z","shell.execute_reply":"2024-12-02T02:41:50.592691Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Insurance Duration ⏳️\n<font size=\"3\"> The distribution of this data is uniform. The duration is in years.</font>","metadata":{}},{"cell_type":"code","source":"df['Insurance Duration'].describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:44:02.354738Z","iopub.execute_input":"2024-12-02T02:44:02.357104Z","iopub.status.idle":"2024-12-02T02:44:02.424499Z","shell.execute_reply.started":"2024-12-02T02:44:02.357024Z","shell.execute_reply":"2024-12-02T02:44:02.422873Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"_, axs = plt.subplots(1,2, figsize=(12,6))\nsns.histplot(df['Insurance Duration'], discrete=True, color='#377ec8', ax=axs[0])\nsns.boxplot(data=df['Insurance Duration'], color='#7ec837', ax=axs[1]);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:44:05.508978Z","iopub.execute_input":"2024-12-02T02:44:05.509383Z","iopub.status.idle":"2024-12-02T02:44:06.870703Z","shell.execute_reply.started":"2024-12-02T02:44:05.509346Z","shell.execute_reply":"2024-12-02T02:44:06.869372Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Correlation of the Numeric Features\n<font size=\"3\"> Looking at the pairplots of the numeric features, it seems like there's no multicorrelation between the regressors. Furthermore, there is no clear correlation between `Premium Amount` and each the numeric regressors alone.</font>","metadata":{}},{"cell_type":"code","source":"sns.pairplot(df[['Age', 'Annual Income', 'Health Score', 'Number of Dependents', 'Vehicle Age', 'Premium Amount']].to_pandas(), \n             corner=True);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T02:48:25.103271Z","iopub.execute_input":"2024-12-02T02:48:25.103662Z","iopub.status.idle":"2024-12-02T02:49:40.011941Z","shell.execute_reply.started":"2024-12-02T02:48:25.103625Z","shell.execute_reply":"2024-12-02T02:49:40.010474Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Categorical Features\n<font size=\"3\"> We have several categorical variables in this dataset, such as `Age`, `Education Level`, just to name a few. **All of them have a uniform marginal distribution** (excluding the NAs). </font>","metadata":{}},{"cell_type":"code","source":"freq = df['Gender'].value_counts()\nplt.bar(freq['Gender'], freq['count'], color='#377ec8')\nplt.title('Gender');","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-12-02T03:15:26.346151Z","iopub.execute_input":"2024-12-02T03:15:26.346517Z","iopub.status.idle":"2024-12-02T03:15:26.607284Z","shell.execute_reply.started":"2024-12-02T03:15:26.346485Z","shell.execute_reply":"2024-12-02T03:15:26.606058Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"freq = df['Marital Status'].fill_null(\"Missing\").value_counts()\nplt.bar(freq['Marital Status'], freq['count'], color='#377ec8')\nplt.title('Marital Status');","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-12-02T03:15:00.121869Z","iopub.execute_input":"2024-12-02T03:15:00.122241Z","iopub.status.idle":"2024-12-02T03:15:00.471014Z","shell.execute_reply.started":"2024-12-02T03:15:00.12221Z","shell.execute_reply":"2024-12-02T03:15:00.469848Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"freq = df['Education Level'].value_counts()\nplt.bar(freq['Education Level'], freq['count'], color='#377ec8')\nplt.title('Education Level');","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-12-02T03:14:27.270865Z","iopub.execute_input":"2024-12-02T03:14:27.271375Z","iopub.status.idle":"2024-12-02T03:14:27.559515Z","shell.execute_reply.started":"2024-12-02T03:14:27.271327Z","shell.execute_reply":"2024-12-02T03:14:27.558359Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"freq = df['Occupation'].fill_null(\"Missing\").value_counts()\nplt.bar(freq['Occupation'], freq['count'], color='#377ec8')\nplt.title('Occupation');","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-12-02T03:13:39.97865Z","iopub.execute_input":"2024-12-02T03:13:39.979053Z","iopub.status.idle":"2024-12-02T03:13:40.350074Z","shell.execute_reply.started":"2024-12-02T03:13:39.979018Z","shell.execute_reply":"2024-12-02T03:13:40.348886Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"freq = df['Location'].fill_null(\"Missing\").value_counts()\nplt.bar(freq['Location'], freq['count'], color='#377ec8')\nplt.title('Location');","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-12-02T03:13:02.78273Z","iopub.execute_input":"2024-12-02T03:13:02.783149Z","iopub.status.idle":"2024-12-02T03:13:03.069158Z","shell.execute_reply.started":"2024-12-02T03:13:02.783113Z","shell.execute_reply":"2024-12-02T03:13:03.068004Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"freq = df['Policy Type'].fill_null(\"Missing\").value_counts()\nplt.bar(freq['Policy Type'], freq['count'], color='#377ec8')\nplt.title('Policy Type');","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-12-02T03:12:32.717822Z","iopub.execute_input":"2024-12-02T03:12:32.718248Z","iopub.status.idle":"2024-12-02T03:12:33.084076Z","shell.execute_reply.started":"2024-12-02T03:12:32.718213Z","shell.execute_reply":"2024-12-02T03:12:33.082846Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"policy_start_date = df['Policy Start Date'].str.to_datetime()\nmonth_start = policy_start_date.dt.month()\nyear_start = policy_start_date.dt.year()\n\n\n_, axs = plt.subplots(1, 2, figsize=(12,6))\nsns.histplot(month_start, discrete=True, color='#377ec8', ax=axs[0])\naxs[0].set_title('Month start')\nsns.histplot(year_start, discrete=True, color='#7ec837', ax=axs[1])\naxs[1].set_title('Year start');","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-02T03:08:15.381744Z","iopub.execute_input":"2024-12-02T03:08:15.382659Z","iopub.status.idle":"2024-12-02T03:08:17.882408Z","shell.execute_reply.started":"2024-12-02T03:08:15.382609Z","shell.execute_reply":"2024-12-02T03:08:17.881217Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"freq = df['Customer Feedback'].fill_null(\"Missing\").value_counts()\nplt.bar(freq['Customer Feedback'], freq['count'], color='#377ec8')\nplt.title('Customer Feedback');","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-12-02T03:15:55.92653Z","iopub.execute_input":"2024-12-02T03:15:55.92697Z","iopub.status.idle":"2024-12-02T03:15:56.275069Z","shell.execute_reply.started":"2024-12-02T03:15:55.926932Z","shell.execute_reply":"2024-12-02T03:15:56.27396Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"freq = df['Smoking Status'].fill_null(\"Missing\").value_counts()\nplt.bar(freq['Smoking Status'], freq['count'], color='#377ec8')\nplt.title('Smoking Status');","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-12-02T03:16:12.499749Z","iopub.execute_input":"2024-12-02T03:16:12.500177Z","iopub.status.idle":"2024-12-02T03:16:12.75443Z","shell.execute_reply.started":"2024-12-02T03:16:12.500143Z","shell.execute_reply":"2024-12-02T03:16:12.75329Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"freq = df['Exercise Frequency'].fill_null(\"Missing\").value_counts()\nplt.bar(freq['Exercise Frequency'], freq['count'], color='#377ec8')\nplt.title('Exercise Frequency');","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-12-02T03:16:29.503633Z","iopub.execute_input":"2024-12-02T03:16:29.504119Z","iopub.status.idle":"2024-12-02T03:16:29.854617Z","shell.execute_reply.started":"2024-12-02T03:16:29.504082Z","shell.execute_reply":"2024-12-02T03:16:29.853133Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"freq = df['Property Type'].fill_null(\"Missing\").value_counts()\nplt.bar(freq['Property Type'], freq['count'], color='#377ec8')\nplt.title('Property Type');","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-12-02T03:16:41.672088Z","iopub.execute_input":"2024-12-02T03:16:41.672492Z","iopub.status.idle":"2024-12-02T03:16:41.970602Z","shell.execute_reply.started":"2024-12-02T03:16:41.672456Z","shell.execute_reply":"2024-12-02T03:16:41.96932Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<font size=\"3\"> This is a work-in-progress notebook, so I will add more insights.</font>\n\n<font size=\"3\"> Thank you so much for checking out this notebook!</font>","metadata":{}}]}