{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.10","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":56537,"databundleVersionId":8015876,"sourceType":"competition"}],"dockerImageVersionId":30497,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Table of Contents\n* [Data description](#data)\n* [Init](#init)\n* [Import and Overview (subset)](#import)\n* [Targets and Features](#features_target)\n* [Correlations](#corr)\n* [Other Explorations](#other)\n* [Import individual columns (full data)](#import_col)","metadata":{}},{"cell_type":"markdown","source":"<a id='data'></a>\n# Data description\n\n## Inputs:\n### Arrays (dimension 60):\n* state_t: air temperature\n* state_q0001: specific humidity\n* state_q0002: cloud liquid mixing ratio\n* state_q0003: cloud ice mixing ratio\n* state_u: zonal wind speed\n* state_v: meridional wind speed\n* pbuf_ozone: ozone volume mixing ratio\n* pbuf_CH4: methane volume mixing ratio\n* pbuf_N2O: nitrous oxide volume mixing ratio\n\n### Scalars:\n* state_ps: surface pressure\n* pbuf_SOLIN: solar insolation\n* pbuf_LHFLX: surface latent heat flux\n* pbuf_SHFLX: surface sensible heat flux\n* pbuf_TAUX: zonal surface stress\n* pbuf_TAUY: meridional surface stress\n* pbuf_COSZRS: cosine of solar zenith angle\n* cam_in_ALDIF: albedo for diffuse longwave radiation\n* cam_in_ALDIR: albedo for direct longwave radiation\n* cam_in_ASDIF: albedo for diffuse shortwave radiation\n* cam_in_ASDIR: albedo for direct shortwave radiation\n* cam_in_LWUP: upward longwave flux\n* cam_in_ICEFRAC: sea-ice areal fraction\n* cam_in_LANDFRAC: land areal fraction\n* cam_in_OCNFRAC: ocean areal fraction\n* cam_in_SNOWHLAND: snow depth over land\n\n## Targets:\n### Arrays (dimension 60):\n* ptend_t: heating tendency\n* ptend_q0001: moistening tendency\n* ptend_q0002: cloud liquid mixing ratio change over time\t\n* ptend_q0003: cloud ice mixing ratio change over time\n* ptend_u: zonal wind acceleration\n* ptend_v: meridional wind acceleration\n\n### Scalars:\n* cam_out_NETSW: net shortwave flux at surface\n* cam_out_FLWDS: downward longwave flux at surface\n* cam_out_PRECSC: snow rate (liquid water equivalent)\n* cam_out_PRECC: rain rate\n* cam_out_SOLS: downward visible direct solar flux to surface\n* cam_out_SOLL: downward near-infrared direct solar flux to surface\n* cam_out_SOLSD: downward diffuse solar flux to surface\n* cam_out_SOLLD: downward diffuse near-infrared solar flux to surface\n","metadata":{}},{"cell_type":"markdown","source":"<a id='init'></a>\n# Init","metadata":{}},{"cell_type":"code","source":"# packages\n\n# standard\nimport numpy as np\nimport pandas as pd\nimport time\n\n# plots\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# faster alternative to pandas\nimport polars as pl","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-04-22T17:53:51.516613Z","iopub.execute_input":"2024-04-22T17:53:51.517049Z","iopub.status.idle":"2024-04-22T17:53:53.619478Z","shell.execute_reply.started":"2024-04-22T17:53:51.517015Z","shell.execute_reply":"2024-04-22T17:53:53.618035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# configs\npd.set_option('display.max_columns', None) # we want to display all columns in this notebook\n\n# aesthetics\ndefault_color_1 = 'darkblue'\ndefault_color_2 = 'darkgreen'\ndefault_color_3 = 'darkred'\n\n# random seed\nmy_random_seed = 111","metadata":{"execution":{"iopub.status.busy":"2024-04-22T17:53:53.62224Z","iopub.execute_input":"2024-04-22T17:53:53.623729Z","iopub.status.idle":"2024-04-22T17:53:53.630586Z","shell.execute_reply.started":"2024-04-22T17:53:53.623673Z","shell.execute_reply":"2024-04-22T17:53:53.629264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='import'></a>\n# Import and Overview (subset)","metadata":{}},{"cell_type":"code","source":"# file overview\n!ls -l '../input/leap-atmospheric-physics-ai-climsim/'","metadata":{"execution":{"iopub.status.busy":"2024-04-22T17:53:53.63224Z","iopub.execute_input":"2024-04-22T17:53:53.632864Z","iopub.status.idle":"2024-04-22T17:53:54.753383Z","shell.execute_reply.started":"2024-04-22T17:53:53.632817Z","shell.execute_reply":"2024-04-22T17:53:54.751548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Datasets are HUGE, let's start with a small subset:","metadata":{}},{"cell_type":"code","source":"# import SUBSET of data\nn_rows = 100000\nfolder = 'leap-atmospheric-physics-ai-climsim'\nt1 = time.time()\ndf_train = pl.read_csv('../input/'+folder+'/train.csv', n_rows=n_rows).to_pandas()\ndf_test = pl.read_csv('../input/'+folder+'/test.csv', n_rows=n_rows).to_pandas()\ndf_sub = pl.read_csv('../input/'+folder+'/sample_submission.csv', n_rows=n_rows).to_pandas()\nt2 = time.time()\nprint('Elapsed time [s]: ', np.round(t2-t1,2))","metadata":{"execution":{"iopub.status.busy":"2024-04-22T17:53:54.755699Z","iopub.execute_input":"2024-04-22T17:53:54.757082Z","iopub.status.idle":"2024-04-22T17:54:07.200507Z","shell.execute_reply.started":"2024-04-22T17:53:54.757019Z","shell.execute_reply":"2024-04-22T17:54:07.199055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# preview - train\ndf_train.head(10)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T17:55:07.70237Z","iopub.execute_input":"2024-04-22T17:55:07.702899Z","iopub.status.idle":"2024-04-22T17:55:08.802168Z","shell.execute_reply.started":"2024-04-22T17:55:07.702848Z","shell.execute_reply":"2024-04-22T17:55:08.801109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train set overview\ndf_train.info(verbose=True, show_counts=True)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-22T17:55:18.475317Z","iopub.execute_input":"2024-04-22T17:55:18.476029Z","iopub.status.idle":"2024-04-22T17:55:18.757034Z","shell.execute_reply.started":"2024-04-22T17:55:18.475978Z","shell.execute_reply":"2024-04-22T17:55:18.755524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# preview - test\ndf_test.head(10)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T17:55:25.834664Z","iopub.execute_input":"2024-04-22T17:55:25.835121Z","iopub.status.idle":"2024-04-22T17:55:26.445814Z","shell.execute_reply.started":"2024-04-22T17:55:25.835087Z","shell.execute_reply":"2024-04-22T17:55:26.444645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# test set overview\ndf_test.info(verbose=True, show_counts=True)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-22T17:55:28.874094Z","iopub.execute_input":"2024-04-22T17:55:28.874552Z","iopub.status.idle":"2024-04-22T17:55:28.992035Z","shell.execute_reply.started":"2024-04-22T17:55:28.874516Z","shell.execute_reply":"2024-04-22T17:55:28.99055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='features_target'></a>\n# Targets and Features","metadata":{}},{"cell_type":"code","source":"# targets (extract from submission file)\ntargets = [x for x in df_sub.columns.tolist() if x not in ['sample_id']]\n\n# numerical features\nfeatures_num = [x for x in df_train.columns.tolist() if x not in ['sample_id']+targets]\n\n# categorical features\nfeatures_cat = []\n\n# all features combined\nfeatures = features_num + features_cat","metadata":{"execution":{"iopub.status.busy":"2024-04-22T17:55:34.891341Z","iopub.execute_input":"2024-04-22T17:55:34.891905Z","iopub.status.idle":"2024-04-22T17:55:34.91009Z","shell.execute_reply.started":"2024-04-22T17:55:34.891862Z","shell.execute_reply":"2024-04-22T17:55:34.90884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ouput of dimensions\nprint(\"Number of numerical features:\", len(features_num))\nprint(\"Number of categorical features:\", len(features_cat))\nprint(\"Number of targets:\", len(targets))\nprint(\"Size of subset:\", n_rows)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T17:55:49.36263Z","iopub.execute_input":"2024-04-22T17:55:49.363081Z","iopub.status.idle":"2024-04-22T17:55:49.369773Z","shell.execute_reply.started":"2024-04-22T17:55:49.363045Z","shell.execute_reply":"2024-04-22T17:55:49.368837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Targets","metadata":{}},{"cell_type":"code","source":"# basic stats - targets\ndf_train[targets].describe()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T17:55:52.498923Z","iopub.execute_input":"2024-04-22T17:55:52.499304Z","iopub.status.idle":"2024-04-22T17:55:54.866012Z","shell.execute_reply.started":"2024-04-22T17:55:52.499277Z","shell.execute_reply":"2024-04-22T17:55:54.865089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot target distributions in compact matrix form\nfig, axs = plt.subplots(92, 4, figsize=(16,350))\ni = 0\nfor t in targets:\n    current_ax = axs.flat[i]\n    current_ax.hist(df_train[t], bins=100, color=default_color_3)\n    current_ax.set_title('Target ' + str(t))\n    current_ax.grid()\n    i = i + 1","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-22T06:46:27.803098Z","iopub.execute_input":"2024-04-22T06:46:27.803822Z","iopub.status.idle":"2024-04-22T06:49:11.002059Z","shell.execute_reply.started":"2024-04-22T06:46:27.80378Z","shell.execute_reply":"2024-04-22T06:49:11.000106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Features","metadata":{}},{"cell_type":"code","source":"# basic stats - train\ndf_train[features_num].describe()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T06:49:11.003912Z","iopub.execute_input":"2024-04-22T06:49:11.004375Z","iopub.status.idle":"2024-04-22T06:49:14.53358Z","shell.execute_reply.started":"2024-04-22T06:49:11.004337Z","shell.execute_reply":"2024-04-22T06:49:14.532396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# basic stats - test\ndf_test[features_num].describe()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T06:49:14.535299Z","iopub.execute_input":"2024-04-22T06:49:14.535769Z","iopub.status.idle":"2024-04-22T06:49:17.898954Z","shell.execute_reply.started":"2024-04-22T06:49:14.535728Z","shell.execute_reply":"2024-04-22T06:49:17.897832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot histograms for numerical features (train and test)\nfor f in features_num:\n    plt.figure(figsize=(12,2))\n    ax1 = plt.subplot(1,2,1)\n    df_train[f].plot(kind='hist', bins=100, color=default_color_1)\n    plt.title(f + ' - Train')\n    plt.grid()\n    ax2 = plt.subplot(1,2,2, sharex=ax1)\n    df_test[f].plot(kind='hist', bins=100, color=default_color_2)\n    plt.title(f + ' - Test')\n    plt.grid()\n    plt.show()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-22T06:49:17.900229Z","iopub.execute_input":"2024-04-22T06:49:17.900649Z","iopub.status.idle":"2024-04-22T06:57:01.961275Z","shell.execute_reply.started":"2024-04-22T06:49:17.900612Z","shell.execute_reply":"2024-04-22T06:57:01.959892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# compact boxplot of all features - train only\nn_plot_rows = 10\nn_plot_cols = 60\nn = len(features_num)\nfor i in range(n_plot_rows):\n    a = n_plot_cols*i+1\n    b = min(n_plot_cols*i+n_plot_cols, n)\n    print('Columns', a, 'to', b)\n    df_train.iloc[:,a:(b+1)].plot(kind='box', figsize=(15,5))\n    plt.xticks(rotation=90)\n    plt.grid()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T18:39:43.493428Z","iopub.execute_input":"2024-04-22T18:39:43.494002Z","iopub.status.idle":"2024-04-22T18:39:58.188329Z","shell.execute_reply.started":"2024-04-22T18:39:43.493955Z","shell.execute_reply":"2024-04-22T18:39:58.186877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# boxplots (train and test)\nfor f in features_num:\n    plt.figure(figsize=(14,0.5))\n    ax1 = plt.subplot(1,2,1)\n    df_temp = df_train[f].dropna() # boxplot does not like missings...\n    plt.boxplot(df_temp, vert=False)\n    plt.title(f + ' - Train')\n    plt.grid()\n    ax2 = plt.subplot(1,2,2, sharex=ax1)\n    df_temp = df_test[f].dropna()\n    plt.boxplot(df_temp, vert=False)\n    plt.title(f + ' - Test')\n    plt.grid()\n    plt.show()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-22T06:57:18.14195Z","iopub.execute_input":"2024-04-22T06:57:18.142395Z","iopub.status.idle":"2024-04-22T07:00:45.458945Z","shell.execute_reply.started":"2024-04-22T06:57:18.142352Z","shell.execute_reply":"2024-04-22T07:00:45.457498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='corr'></a>\n# Correlations","metadata":{}},{"cell_type":"markdown","source":"### Targets","metadata":{}},{"cell_type":"code","source":"# calc and plot correlation matrix\ncor_p_target = df_train[targets].corr(method='pearson')\nplt.figure(figsize=(14,12))\nsns.heatmap(cor_p_target, annot=False, cmap='RdYlGn',\n            vmin=-1, vmax=+1)\nplt.title('Targets - Pearson Correlation')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T07:00:45.460511Z","iopub.execute_input":"2024-04-22T07:00:45.460884Z","iopub.status.idle":"2024-04-22T07:01:06.247258Z","shell.execute_reply.started":"2024-04-22T07:00:45.460854Z","shell.execute_reply":"2024-04-22T07:01:06.245902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Features (train)","metadata":{}},{"cell_type":"code","source":"# calc and plot correlation matrix\ncor_p_train = df_train[features_num].corr(method='pearson')\nplt.figure(figsize=(14,12))\nsns.heatmap(cor_p_train, annot=False, cmap='RdYlGn',\n            vmin=-1, vmax=+1)\nplt.title('Features - Pearson Correlation (train)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T07:01:06.249197Z","iopub.execute_input":"2024-04-22T07:01:06.249758Z","iopub.status.idle":"2024-04-22T07:01:51.096072Z","shell.execute_reply.started":"2024-04-22T07:01:06.249713Z","shell.execute_reply":"2024-04-22T07:01:51.094861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Features (test)","metadata":{}},{"cell_type":"code","source":"# calc and plot correlation matrix\ncor_p_test = df_test[features_num].corr(method='pearson')\nplt.figure(figsize=(14,12))\nsns.heatmap(cor_p_test, annot=False, cmap='RdYlGn',\n            vmin=-1, vmax=+1)\nplt.title('Features - Pearson Correlation (test)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T07:01:51.09781Z","iopub.execute_input":"2024-04-22T07:01:51.098225Z","iopub.status.idle":"2024-04-22T07:02:33.718371Z","shell.execute_reply.started":"2024-04-22T07:01:51.09819Z","shell.execute_reply":"2024-04-22T07:02:33.716688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# export results\ndf_train.to_csv('df_train_subset.csv')\ncor_p_target.to_csv('cor_p_target.csv')\ncor_p_train.to_csv('cor_p_train.csv')\ncor_p_test.to_csv('cor_p_test.csv')","metadata":{"execution":{"iopub.status.busy":"2024-04-22T07:02:33.719671Z","iopub.status.idle":"2024-04-22T07:02:33.720132Z","shell.execute_reply.started":"2024-04-22T07:02:33.719919Z","shell.execute_reply":"2024-04-22T07:02:33.71994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='other'></a>\n# Other Explorations","metadata":{}},{"cell_type":"markdown","source":"### Target vs row index","metadata":{}},{"cell_type":"code","source":"# plot target values\nfor t in targets:\n    plt.figure(figsize=(14,2))\n    plt.scatter(df_train.index, df_train[t], color=default_color_3,\n                alpha=0.25, s=1)\n    plt.title(t)\n    plt.grid()\n    plt.show()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-22T07:02:33.723379Z","iopub.status.idle":"2024-04-22T07:02:33.724146Z","shell.execute_reply.started":"2024-04-22T07:02:33.723808Z","shell.execute_reply":"2024-04-22T07:02:33.723841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Features vs row index","metadata":{}},{"cell_type":"code","source":"# plot feature values\nfor f in features_num:\n    plt.figure(figsize=(14,2))\n    plt.scatter(df_train.index, df_train[f], color=default_color_1,\n                alpha=0.25, s=1)\n    plt.title(f)\n    plt.grid()\n    plt.show()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-22T07:02:33.726193Z","iopub.status.idle":"2024-04-22T07:02:33.726816Z","shell.execute_reply.started":"2024-04-22T07:02:33.726493Z","shell.execute_reply":"2024-04-22T07:02:33.72652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='import_col'></a>\n# Import individual columns (full data)","metadata":{}},{"cell_type":"markdown","source":"#### In order to approach the full dataset we could try to import just a subset of columns. This is shown in the following section.","metadata":{}},{"cell_type":"code","source":"# define columns (has to be a list)\nn_max = 20 # columns with index 0..n_max\ncols_select = ['state_t_' + str(t) for t in range(0,n_max+1)]\nprint(cols_select)","metadata":{"execution":{"iopub.status.busy":"2024-04-22T17:58:38.19194Z","iopub.execute_input":"2024-04-22T17:58:38.192358Z","iopub.status.idle":"2024-04-22T17:58:38.199607Z","shell.execute_reply.started":"2024-04-22T17:58:38.192328Z","shell.execute_reply":"2024-04-22T17:58:38.198079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# load only selected column\nt1 = time.time()\ndf_col = pl.read_csv('../input/'+folder+'/train.csv', columns=cols_select).to_pandas()\nt2 = time.time()\nprint('Elapsed time [s]: ', np.round(t2-t1,2))\nprint('Number of rows: ', df_col.shape[0])\nprint('Number of cols: ', df_col.shape[1])","metadata":{"execution":{"iopub.status.busy":"2024-04-22T17:58:39.669639Z","iopub.execute_input":"2024-04-22T17:58:39.6704Z","iopub.status.idle":"2024-04-22T18:15:25.80039Z","shell.execute_reply.started":"2024-04-22T17:58:39.670356Z","shell.execute_reply":"2024-04-22T18:15:25.799032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# basic stats\ndf_col.describe(percentiles=[0.01,0.1,0.25,0.5,0.75,0.9,0.99])","metadata":{"execution":{"iopub.status.busy":"2024-04-22T18:15:25.804779Z","iopub.execute_input":"2024-04-22T18:15:25.805188Z","iopub.status.idle":"2024-04-22T18:15:38.204996Z","shell.execute_reply.started":"2024-04-22T18:15:25.805156Z","shell.execute_reply":"2024-04-22T18:15:38.203355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot distributions\nfor f in cols_select:\n    plt.figure(figsize=(10,3))\n    plt.hist(df_col[f], bins=1000, color=default_color_1)\n    plt.title(f + ' - full data')\n    plt.grid()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T18:19:29.927406Z","iopub.execute_input":"2024-04-22T18:19:29.927925Z","iopub.status.idle":"2024-04-22T18:20:11.68027Z","shell.execute_reply.started":"2024-04-22T18:19:29.927887Z","shell.execute_reply":"2024-04-22T18:20:11.679099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 💡 state_t_0 shows some unusually high values, let's have a closer look:","metadata":{}},{"cell_type":"code","source":"# boxplot\nplt.figure(figsize=(10,0.5))\nplt.boxplot(df_col.state_t_0, vert=False)\nplt.title('state_t_0')\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T18:23:12.077374Z","iopub.execute_input":"2024-04-22T18:23:12.077807Z","iopub.status.idle":"2024-04-22T18:23:13.171443Z","shell.execute_reply.started":"2024-04-22T18:23:12.077756Z","shell.execute_reply":"2024-04-22T18:23:13.170151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# let's check the most extreme outliers\ndf_col[df_col.state_t_0>400]","metadata":{"execution":{"iopub.status.busy":"2024-04-22T18:40:15.058301Z","iopub.execute_input":"2024-04-22T18:40:15.058723Z","iopub.status.idle":"2024-04-22T18:40:15.099893Z","shell.execute_reply.started":"2024-04-22T18:40:15.058686Z","shell.execute_reply":"2024-04-22T18:40:15.098526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Correlations:","metadata":{}},{"cell_type":"code","source":"# calc and plot correlation matrix\ncor_p_train_few_cols = df_train[cols_select].corr(method='pearson')\nplt.figure(figsize=(14,10))\nsns.heatmap(cor_p_train_few_cols, annot=True, cmap='RdYlGn',\n            fmt='.2f', linecolor='black', linewidth=.5,\n            vmin=-1, vmax=+1)\nplt.title('Features - Pearson Correlation (train/selected columns)')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T18:26:33.091945Z","iopub.execute_input":"2024-04-22T18:26:33.092391Z","iopub.status.idle":"2024-04-22T18:26:34.83497Z","shell.execute_reply.started":"2024-04-22T18:26:33.092358Z","shell.execute_reply":"2024-04-22T18:26:34.833015Z"},"trusted":true},"execution_count":null,"outputs":[]}]}