{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":56537,"databundleVersionId":8015876,"sourceType":"competition"}],"dockerImageVersionId":30732,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# EDA using Clustering and Random Sampling ","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-06-11T20:16:52.137632Z","iopub.execute_input":"2024-06-11T20:16:52.138101Z","iopub.status.idle":"2024-06-11T20:16:52.143469Z","shell.execute_reply.started":"2024-06-11T20:16:52.138066Z","shell.execute_reply":"2024-06-11T20:16:52.142161Z"}}},{"cell_type":"markdown","source":"#### Challenge: Visualize the dataset without bias and with memory constraints.\n\n#### Solution : Random Sampling + Clustering ","metadata":{}},{"cell_type":"markdown","source":"### Why Random Sampling \n\nThe hypothesis is that with enough random samples we will be able to understand the distribution of the data. This is based on Law of Large Numbers.\nUnderstanding the distribution of of data through random sampling can help make inferences about the entire dataset eliminating the need to parse though the entire dataset due to its size.","metadata":{}},{"cell_type":"markdown","source":"### Why Custering\n\nClustering is a powerful technique in machine learning and data analysis for reducing the dimensionality of data and revealing underlying patterns. By grouping similar data points together, clustering can help simplify and summarize the dataset, making it easier to visualize and analyze without having to process the entire dataset at once. ","metadata":{}},{"cell_type":"markdown","source":"### Why not PCA, s-tne and other DR methods\n\nNo information on whether the data follows linear or non-linear pattern. \nPCA is effetcive with linear data. \nEvaluation is much harder for chossing optimal components.\ns-tne is computationallyexpensive for laegr datasets. ","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Code","metadata":{}},{"cell_type":"code","source":"pip install pyarrow","metadata":{"execution":{"iopub.status.busy":"2024-06-10T17:44:45.040049Z","iopub.execute_input":"2024-06-10T17:44:45.040592Z","iopub.status.idle":"2024-06-10T17:44:56.987715Z","shell.execute_reply.started":"2024-06-10T17:44:45.040549Z","shell.execute_reply":"2024-06-10T17:44:56.9865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Prep","metadata":{}},{"cell_type":"code","source":"#import lib \n\nimport pandas as pd\nimport numpy as np \nimport random\nimport pandas as pd \nimport gc  \nimport os\nimport random\nfrom tqdm import tqdm\nimport dask.dataframe as dd\nfrom sklearn.cluster import KMeans\nimport pandas as pd\nimport pyarrow as pa\nimport pyarrow.parquet as pq\nimport plotly.express as px\nimport polars as pl \nimport pandas \nimport os \nfrom sklearn.cluster import KMeans \nimport gc \ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2024-06-11T21:53:59.210654Z","iopub.execute_input":"2024-06-11T21:53:59.211202Z","iopub.status.idle":"2024-06-11T21:54:03.482031Z","shell.execute_reply.started":"2024-06-11T21:53:59.211166Z","shell.execute_reply":"2024-06-11T21:54:03.480568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Variables \nChange these to fit your needs","metadata":{"execution":{"iopub.status.busy":"2024-06-11T20:44:44.344111Z","iopub.execute_input":"2024-06-11T20:44:44.344696Z","iopub.status.idle":"2024-06-11T20:44:44.354275Z","shell.execute_reply.started":"2024-06-11T20:44:44.344655Z","shell.execute_reply":"2024-06-11T20:44:44.352559Z"}}},{"cell_type":"code","source":"train_csv_path = '/kaggle/input/leap-atmospheric-physics-ai-climsim/train.csv'\nparquet_file = '/kaggle/input/train.parquet'\n\n#Change this to fit your needs (This should work fine with kaggle CPU notebook)\nchunksize = 100000\n\n#Change this to decide how much you want to sample from entire population (10%-30% recommended)\nsample_ratio = 0.16\nbatch_size = 50000\n\n#columns from the dataset, only feature with multiple columns are mentioned here,change based on your needs\ncol_names = [\"state_t\",\"state_q0001\",\"state_q0002\",\"state_q0003\",\"state_u\",\"state_v\",\"pbuf_ozone\",\n            \"pbuf_CH4\",\"pbuf_N2O\",\"ptend_v\",\"ptend_u\",\"ptend_q0003\",\"ptend_q0002\",\"ptend_q0001\",\"ptend_t\"\n            ]","metadata":{"execution":{"iopub.status.busy":"2024-06-11T21:54:03.484811Z","iopub.execute_input":"2024-06-11T21:54:03.485679Z","iopub.status.idle":"2024-06-11T21:54:03.494447Z","shell.execute_reply.started":"2024-06-11T21:54:03.485613Z","shell.execute_reply":"2024-06-11T21:54:03.492253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Prepare dataset using Polars and Random Sampling","metadata":{}},{"cell_type":"code","source":"#Get list of column names\nload_cols = []\n\n#Change this based on what column you need \nbase_col_name = col_names[1]\n\n#Imp Change the range if your using columns not mentioned in the list above\nfor col_num in range(0,60):\n    load_cols.append(base_col_name + \"_\" + str(col_num))","metadata":{"execution":{"iopub.status.busy":"2024-06-11T21:54:24.603636Z","iopub.execute_input":"2024-06-11T21:54:24.605022Z","iopub.status.idle":"2024-06-11T21:54:24.61088Z","shell.execute_reply.started":"2024-06-11T21:54:24.604978Z","shell.execute_reply":"2024-06-11T21:54:24.609627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#check / verify \nload_cols[44] ","metadata":{"execution":{"iopub.status.busy":"2024-06-11T21:54:27.160493Z","iopub.execute_input":"2024-06-11T21:54:27.161045Z","iopub.status.idle":"2024-06-11T21:54:27.168771Z","shell.execute_reply.started":"2024-06-11T21:54:27.161002Z","shell.execute_reply":"2024-06-11T21:54:27.167355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"reader = pl.read_csv_batched(\ntrain_csv_path, \nbatch_size = batch_size\n    )","metadata":{"execution":{"iopub.status.busy":"2024-06-11T21:54:32.667919Z","iopub.execute_input":"2024-06-11T21:54:32.6684Z","iopub.status.idle":"2024-06-11T21:54:32.748774Z","shell.execute_reply.started":"2024-06-11T21:54:32.668363Z","shell.execute_reply":"2024-06-11T21:54:32.747056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### This block goes through train csv in chunks, \n##### concatenates the batches and shuffling it \n##### and then appending the sampled data to a list.\n##### Then finally samples from the data.","metadata":{}},{"cell_type":"code","source":"%time\nrandom_sample = []\ncounter = 0\nbatches = reader.next_batches(5)\nwhile batches: \n    batches = pl.concat(batches)\n    batches = batches.sample(fraction = sample_ratio, shuffle = True)\n    random_sample.append(batches)\n    counter += len(batches)\n    batches = reader.next_batches(5)\n    \nprint(\"Processing Finished\")\n","metadata":{"execution":{"iopub.status.busy":"2024-06-11T21:54:35.175439Z","iopub.execute_input":"2024-06-11T21:54:35.175899Z","iopub.status.idle":"2024-06-11T22:16:39.48108Z","shell.execute_reply.started":"2024-06-11T21:54:35.175863Z","shell.execute_reply":"2024-06-11T22:16:39.478313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_data = pl.concat(random_sample)\n%time sample_data = sample_data.sample(1500000, shuffle=True)","metadata":{"execution":{"iopub.status.busy":"2024-06-11T22:16:39.485366Z","iopub.execute_input":"2024-06-11T22:16:39.485857Z","iopub.status.idle":"2024-06-11T22:17:00.797983Z","shell.execute_reply.started":"2024-06-11T22:16:39.48579Z","shell.execute_reply":"2024-06-11T22:17:00.795835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Optional save the parqet file if you want to work locally or dont want to build this every run. (~7.2gb)\n#%time sample_data.write_parquet(\"train_sampled_1_n_half_m.parquet\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the relevant columns from the entire dataset \nstate_sample = sample_data[load_cols]","metadata":{"execution":{"iopub.status.busy":"2024-06-11T22:17:00.800511Z","iopub.execute_input":"2024-06-11T22:17:00.802131Z","iopub.status.idle":"2024-06-11T22:17:00.835082Z","shell.execute_reply.started":"2024-06-11T22:17:00.802072Z","shell.execute_reply":"2024-06-11T22:17:00.830853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Check / Verify \nstate_sample.head(2)","metadata":{"execution":{"iopub.status.busy":"2024-06-11T22:17:00.838514Z","iopub.execute_input":"2024-06-11T22:17:00.838947Z","iopub.status.idle":"2024-06-11T22:17:00.88131Z","shell.execute_reply.started":"2024-06-11T22:17:00.838911Z","shell.execute_reply":"2024-06-11T22:17:00.879805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Clustering the selected features","metadata":{}},{"cell_type":"code","source":"#How many clusters do you want\n#Higher the number more information you get but more cpu expensive and vice versa \n#Playing around 20 was the optimal value, feel free to play around with this hyperparameter\n\nclus_n = 20","metadata":{"execution":{"iopub.status.busy":"2024-06-11T22:43:25.846522Z","iopub.execute_input":"2024-06-11T22:43:25.847768Z","iopub.status.idle":"2024-06-11T22:43:25.855718Z","shell.execute_reply.started":"2024-06-11T22:43:25.847722Z","shell.execute_reply":"2024-06-11T22:43:25.854099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Building the model (Clustering Kmeans) \n#Feel free to use other clustering algorithms such as HDSCAN, \n#please keep in mind of time complexity if you choose another algorithm \n\nstate_model = KMeans(n_clusters=clus_n)\n%time state_model.fit(state_sample)","metadata":{"execution":{"iopub.status.busy":"2024-06-11T22:43:27.870922Z","iopub.execute_input":"2024-06-11T22:43:27.871396Z","iopub.status.idle":"2024-06-11T22:47:30.852295Z","shell.execute_reply.started":"2024-06-11T22:43:27.871366Z","shell.execute_reply":"2024-06-11T22:47:30.850269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Time to Visualize ","metadata":{}},{"cell_type":"code","source":"\nfig = px.line(state_model.cluster_centers_.T,\n              template=\"plotly_dark\",\n              title=f\"Clusters for {base_col_name}\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2024-06-11T22:50:23.894405Z","iopub.execute_input":"2024-06-11T22:50:23.894911Z","iopub.status.idle":"2024-06-11T22:50:27.294487Z","shell.execute_reply.started":"2024-06-11T22:50:23.894876Z","shell.execute_reply":"2024-06-11T22:50:27.293052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#### Optional Step ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Save model if you need , will be faster to load next time (~50mb)\n#filename=\"state_clusters.sav\"\n#pickle.dump(state_model,open(filename, \"wb\"))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Please upvote if you find this useful ","metadata":{}}]}