{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **EDA Simplified: G2Net Gravitational Waves Detection**","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"## Introduction\nIn 2015, scientists detected the first class of gravitational waves and they expect that there are more of the discoveries based on that to come. And for all of this, it comes with four classes that present only signals from merging black holes and neutron stars have been detected. Among those remaining are continuous gravitational-wave signals. These are weak, but long-lasting signals sent by hastily, spinning neutron stars. But imagine the mass of our Sun, however condensed into a ball the size of a city and spinning over a thousand times a second. The extreme compactness of these stars, composed of the densest material in the universe, could allow continuous waves to be sent and then detected on Earth. There are potentially more continuous signals from neutron stars in our own galaxy and the current challenge for scientists is to make the first detection of the waves, and hopefully the skills data science can help them finish this mission. And that's how EGO (European Gravitational Observatory) and G2Net created this competition on detecting the gravitational waves. And for all of this, let's start off our data analysis on this complex competition!","metadata":{}},{"cell_type":"markdown","source":"## The Setup\nDoing all of this EDA in this compeition is straightforward. To get started, we import the pandas module as pd and the numpy module as np for data science and linear algebra. Next, we import the plotly module with the express submodule for plotting, along with the seaborn module as sns and the matplotlib module as plt.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\n\nimport plotly.express as px\nimport seaborn as sns\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:22:11.66273Z","iopub.execute_input":"2022-12-11T05:22:11.663229Z","iopub.status.idle":"2022-12-11T05:22:14.222165Z","shell.execute_reply.started":"2022-12-11T05:22:11.663131Z","shell.execute_reply":"2022-12-11T05:22:14.220412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, upon entering this competition, we observed that there are some hdf5 files in the train and test folders, so we install the h5py module via pip and then importing it.","metadata":{}},{"cell_type":"code","source":"!pip3 install h5py\nimport h5py","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:22:14.224816Z","iopub.execute_input":"2022-12-11T05:22:14.227491Z","iopub.status.idle":"2022-12-11T05:22:27.69351Z","shell.execute_reply.started":"2022-12-11T05:22:14.227437Z","shell.execute_reply":"2022-12-11T05:22:27.691979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dataframe Setup + Analysis\nAfter importing the modules, we define train_df  dataframe by assigning them to this csv file, train_labels.csv from the competition data with the read_csv function from the pd module. Finally, we display the train_df dataframe with the head function!","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(\"../input/g2net-detecting-continuous-gravitational-waves/train_labels.csv\")\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:22:27.695201Z","iopub.execute_input":"2022-12-11T05:22:27.695598Z","iopub.status.idle":"2022-12-11T05:22:27.738988Z","shell.execute_reply.started":"2022-12-11T05:22:27.695543Z","shell.execute_reply":"2022-12-11T05:22:27.737816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's proceed to find the number of data in the train_df dataframe! We use the len function to it so that we can count how many data is in there.","metadata":{}},{"cell_type":"code","source":"len(train_df)","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:22:27.741951Z","iopub.execute_input":"2022-12-11T05:22:27.742347Z","iopub.status.idle":"2022-12-11T05:22:27.749984Z","shell.execute_reply.started":"2022-12-11T05:22:27.742312Z","shell.execute_reply":"2022-12-11T05:22:27.748958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 603 data entities in the train_df dataframe since we ran this code cell above. If we print out the shape of our train_df dataframe, then it'll return 603 rows with 2 columns.","metadata":{}},{"cell_type":"markdown","source":"Now let's get into Plotly! Before that, we defined the targets variable to count the values of the target data from the train_df dataframe, setting the normalize parameter (normalization) to True. Next, we defined the fig variable to the pie function from the px module, setting the targets dataframe inside of it, along with the names parameter (names inputted) to the index of the target data from the target dataframe with the index attribute and the values parameter (values for the pie counts) to the values of the target data from the target dataframe with the values attribute. We then update our layout of the fig figure variable with the update_layout function, setting the template parameter (template specification) to ggplot2. Finally, we show off the fig figure with the show function.","metadata":{}},{"cell_type":"code","source":"targets = train_df[\"target\"].value_counts(normalize=True)\n\nfig = px.pie(targets, names=targets.index, values=targets.values)\nfig.update_layout(template=\"ggplot2\")\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:22:27.752064Z","iopub.execute_input":"2022-12-11T05:22:27.752486Z","iopub.status.idle":"2022-12-11T05:22:29.14Z","shell.execute_reply.started":"2022-12-11T05:22:27.752455Z","shell.execute_reply":"2022-12-11T05:22:29.138679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From this pie chart over the target data in the train_df dataframe, the most value counts in each of the data is 1, with 66.3% of data logged in this dataframe, while the least value is -1, with roughly 0.5% of data logged in this dataframe.\n\nNow that we have the overall basics over our analysis over the train_df dataframe, let's get into our analysis over the hdf5 files in this competition data!","metadata":{}},{"cell_type":"markdown","source":"## Data Setup & Analysis from the HDF5 Files\nIn order to start our analysis from the \"hdf5\" files in the competition dataset, we define the training_dir variable to the file directory leading to the train folder, along defining the testing_dir variable to the file directory leadning to the test folder. We also define the target_width and target_height to 360 so that we create a target width and height's 360 by 360 dataset out of the hdf5 files.","metadata":{}},{"cell_type":"code","source":"training_dir = \"../input/g2net-detecting-continuous-gravitational-waves/train/\"\ntesting_dir = \"../input/g2net-detecting-continuous-gravitational-waves/test/\"\n\ntar_width = 360\ntar_height = 360","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:22:29.141674Z","iopub.execute_input":"2022-12-11T05:22:29.142083Z","iopub.status.idle":"2022-12-11T05:22:29.147122Z","shell.execute_reply.started":"2022-12-11T05:22:29.142047Z","shell.execute_reply":"2022-12-11T05:22:29.146043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we've defined the training_dir and testing_dir variables along with the target width and heights by 360, let's create two functions to get the data types from both train and test files! But before that, we define the test_df variable to reading the csv file of the sample submission with the read_csv function from the pd module. \n\nThen we define the get_dtype_train function, containing the train_id parameter inside of it. Next in the inside of this function, we defined the file variable to reading the hdf5 files with the File function provided by the h5py module, containing the formmated string of the file directory leading to the train files of the training_dir variable and the train_id parameter, along with the 'r' string to indicate read-mode, and outside of that, the slice index of the train_id parameter. Finally in this function, we return out datatype of the file variable with the data indexes of H1 and SFTs with the dtype attribute.\n\nWe also define the get_dtype_test function, containing the test_id parameter inside of it too. And inside of this function, it's like the same as we did to the previous function, but for reading the hdf5 files, we only read the test files with the testing_dir variable and the test_id parameter.","metadata":{}},{"cell_type":"code","source":"test_df = pd.read_csv(\"../input/g2net-detecting-continuous-gravitational-waves/sample_submission.csv\")\n\ndef get_dtype_train(train_id):\n    file = h5py.File(f'{training_dir}/{train_id}.hdf5', 'r')[train_id]\n    return file['H1']['SFTs'].dtype\n\ndef get_dtype_test(test_id):\n    file = h5py.File(f'{testing_dir}/{test_id}.hdf5', 'r')[test_id]\n    return file['H1']['SFTs'].dtype","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:22:29.148453Z","iopub.execute_input":"2022-12-11T05:22:29.149776Z","iopub.status.idle":"2022-12-11T05:22:29.177756Z","shell.execute_reply.started":"2022-12-11T05:22:29.149716Z","shell.execute_reply":"2022-12-11T05:22:29.176232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After that, we define the train_df's dtype data index to the train_df dataframe's id data index and set the progress to read hdf5 files with the progress_apply function, containing the get_dtype_train function. Same goes to the test_df dataframe's dtype data index creation, but inside of the progress_apply function, the function call for that is get_dtype_test. However before that, we need to import tqdm from the tqdm module with the notebook submodule, along calling the pandas function from the tqdm module to make the module go with the progress over the pandas module.","metadata":{}},{"cell_type":"code","source":"from tqdm.notebook import tqdm\ntqdm.pandas()\n\ntrain_df[\"dtype\"] = train_df[\"id\"].progress_apply(get_dtype_train)\ntest_df[\"dtype\"] = test_df[\"id\"].progress_apply(get_dtype_test)","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:22:29.179409Z","iopub.execute_input":"2022-12-11T05:22:29.179949Z","iopub.status.idle":"2022-12-11T05:24:00.8689Z","shell.execute_reply.started":"2022-12-11T05:22:29.179896Z","shell.execute_reply":"2022-12-11T05:24:00.867383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After creating the train and test data types, let's create the train and test dimensions! We define the two empty lists, train_dimension_rows and test_dimension_rows and then create two separate for loops: looping the train_id variable to the tqdm function call of the train_df dataframe's id data index and looping the test_id variable to the tqdm function call of the test_df dataframe with the id data index.\n\nInside the first for loop, we defined the file variable to the File function from the h5py module to read the formatted string with the training_dir and train_id variables that leads to the training hdf5 files, setting it with 'r', and the slice index of the train_id variable. Next, the SFT_H and SFT_L variables is defined to the file variable with the slice index of 'H1' ('L1' for the SFT_L variable) and 'SFTs' separately. Finally, we append the dictionary containing the id key set to the train_id variable, the H_height and H_width keys set to the SFT_H's variable shape with the shape attribute having the slice index of 0 and 1 separately, and the L_height and L_width keys set to the SFT'H's variable shape with the shape attribute having the slice index of 0 and 1 separately to the train_dimension_rows list.\n\nAs for the second for loop, it's basically the same as what we did for the first for loop, but the file variable is defined to the File function reading the formatted string leading to the test directories with the testing_dir and test_id variables that leads to testing hdf5 files, setting it to 'r', and the slice index of the test_id variable.","metadata":{}},{"cell_type":"code","source":"train_dimension_rows = []\ntest_dimension_rows = []\n\nfor train_id in tqdm(train_df['id']):\n    file = h5py.File(f'{training_dir}/{train_id}.hdf5', 'r')[train_id]\n    SFT_H = file['H1']['SFTs']\n    SFT_L = file['L1']['SFTs']\n    train_dimension_rows.append({\n        'id': train_id,\n        'H_height': SFT_H.shape[0],\n        'H_width': SFT_H.shape[1],\n        'L_height': SFT_L.shape[0],\n        'L_width': SFT_L.shape[1],\n    })\n    \nfor test_id in tqdm(test_df['id']):\n    file = h5py.File(f'{testing_dir}/{test_id}.hdf5', 'r')[test_id]\n    SFT_H = file['H1']['SFTs']\n    SFT_L = file['L1']['SFTs']\n    test_dimension_rows.append({\n        'id': test_id,\n        'H_height': SFT_H.shape[0],\n        'H_width': SFT_H.shape[1],\n        'L_height': SFT_L.shape[0],\n        'L_width': SFT_L.shape[1],\n    })","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:24:00.870841Z","iopub.execute_input":"2022-12-11T05:24:00.871228Z","iopub.status.idle":"2022-12-11T05:25:36.955325Z","shell.execute_reply.started":"2022-12-11T05:24:00.871195Z","shell.execute_reply":"2022-12-11T05:25:36.953997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After another long run of this code cell above (caused by the number of test data in the test_df dataframe), let's merge our train_df dataframe with the merge function consisting of the dataframe creation of the train_dimension_rows with the DataFrame function from the pd module (and setting the on parameter for data specification to the id data index) and our test_df dataframe with the merge function again consisting of the dataframe creation of the test_dimension_rows with the DataFrame function from the pd module (and again setting the on parameter for data specification to the id data index).","metadata":{}},{"cell_type":"code","source":"train_df = train_df.merge(pd.DataFrame(train_dimension_rows), on='id')\ntest_df = test_df.merge(pd.DataFrame(test_dimension_rows), on='id')","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:36.961424Z","iopub.execute_input":"2022-12-11T05:25:36.96185Z","iopub.status.idle":"2022-12-11T05:25:37.009622Z","shell.execute_reply.started":"2022-12-11T05:25:36.961814Z","shell.execute_reply":"2022-12-11T05:25:37.008071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's take a look at the post-merged train_df and test_df dataframes with the head function!","metadata":{}},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:37.010864Z","iopub.execute_input":"2022-12-11T05:25:37.011211Z","iopub.status.idle":"2022-12-11T05:25:37.02644Z","shell.execute_reply.started":"2022-12-11T05:25:37.011182Z","shell.execute_reply":"2022-12-11T05:25:37.025087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:37.028143Z","iopub.execute_input":"2022-12-11T05:25:37.028888Z","iopub.status.idle":"2022-12-11T05:25:37.046192Z","shell.execute_reply.started":"2022-12-11T05:25:37.028854Z","shell.execute_reply":"2022-12-11T05:25:37.044677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### train_df DataFrame Analysis\n\nNow that we have the two dataframes, train_df and test_df ready, let's visualize the Handford and Livingston distributions separately in Plotly! First, we define fig to the histogram function from the px module to create a histogram distribution graph, setting the train_df inside, the x parameter (x-axis) to H_width, and the marginal parameter (graph on top of the histogram) to violin. We then update our layout of our fig figure with the update_layout function, setting the template parameter (template specification) to ggplot2. Now, let's show off our fig figure graph with the show function!","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"H_width\", marginal=\"violin\")\nfig.update_layout(template=\"ggplot2\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:37.047629Z","iopub.execute_input":"2022-12-11T05:25:37.048001Z","iopub.status.idle":"2022-12-11T05:25:37.215667Z","shell.execute_reply.started":"2022-12-11T05:25:37.047967Z","shell.execute_reply":"2022-12-11T05:25:37.214619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From between 4000 and 5000, we see a spike, in which the most data is in between 4600 and 4649, with 144 data entities, and aside from that, we see that from 0 to 1100, we see one data isolated from each other. Furthermore, as mentioned from the violin plot above from the histogram, the lower fence data is in 4381, the first quartile is in approximately 4528.25, the median is in 4586, the third quartile is in 4634, the upper fense is in 4788, and finally, the max is in 4843.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the Handford height distribution in Plotly! We define the fig variable to the histogram function from the px module, setting the train_df dataframe as our data input, the x parameter (x-axis) to H_height, and the marginal parameter (graph on top of histogram, optional) to violin. We then update our layout in the fig figure with the update_layout function, setting the template parameter (theme for the graph) to ggplot2. Lastly, we show off our fig figure variable with the show function!","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"H_height\", marginal=\"violin\")\nfig.update_layout(template=\"ggplot2\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:37.216756Z","iopub.execute_input":"2022-12-11T05:25:37.217064Z","iopub.status.idle":"2022-12-11T05:25:37.321152Z","shell.execute_reply.started":"2022-12-11T05:25:37.217036Z","shell.execute_reply":"2022-12-11T05:25:37.319823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That's weird. The histogram above for the Handford height shows that all data in the train_df's H_height is 360, with 603 data entities of it. Furthermore, the first quartile, median, third quartile, min, max, lower, and upper fence all have the 360 data counts. Thus, we foresighted the same thing to the Livingston height data if we distributed them.","metadata":{}},{"cell_type":"markdown","source":"Now let's find out the Livingston width distributions in Plotly! We define the fig variable to the histogram function from the px module, setting the x parameter (x-axis) to L_width, and the marginal parameter (graph on the histogram) to violin. We then update our layout of the graph with the update_layout function, setting the template parameter (template specification) to ggplot2. Lastly, we exhibit our fig figure graph with the show function.","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"L_width\", marginal=\"violin\")\nfig.update_layout(template=\"ggplot2\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:37.322601Z","iopub.execute_input":"2022-12-11T05:25:37.323015Z","iopub.status.idle":"2022-12-11T05:25:37.424327Z","shell.execute_reply.started":"2022-12-11T05:25:37.32298Z","shell.execute_reply":"2022-12-11T05:25:37.423049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Unlike the Handford widths, the highest number count of it is in between the range of 4550 and 4599 with 156 data entities of it. However, from 0 to 1000, we see that the isolated data entities in this graph is similar to the Handford width graph previously. Thus, the first quartile is 4531.5, the median is 4582, and the third quartile is 4636.75.","metadata":{}},{"cell_type":"markdown","source":"Now let's use the kernel density estimation method over the Handford and Livingston width distributions in seaborn! We use the kdeplot function from the sns module, setting the data parameter (dataframe input) to the train_df dataframe and the x parameter (x-axis input) to H_width data.","metadata":{}},{"cell_type":"code","source":"sns.kdeplot(data=train_df, x=\"H_width\")","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:37.425776Z","iopub.execute_input":"2022-12-11T05:25:37.426112Z","iopub.status.idle":"2022-12-11T05:25:37.704698Z","shell.execute_reply.started":"2022-12-11T05:25:37.426082Z","shell.execute_reply":"2022-12-11T05:25:37.703274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From between 4000 and 5000 in H_width data, we noticed a huge spike peak, ranging a height up to 0.0035ish in Density. Thus, we observed some tiny bumps from 0 to 1000, meaning that there were some little data around this range.","metadata":{}},{"cell_type":"markdown","source":"Now let's plot out the \"kernel density estimation\" graph in Livingston width! We apply the same thing from what we did for the Handford width distribution in \"KDE\" but the x parameter (x-axis) in the kdeplot is set to L_width.","metadata":{}},{"cell_type":"code","source":"sns.kdeplot(data=train_df, x=\"L_width\")","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:37.705939Z","iopub.execute_input":"2022-12-11T05:25:37.706266Z","iopub.status.idle":"2022-12-11T05:25:37.951358Z","shell.execute_reply.started":"2022-12-11T05:25:37.706237Z","shell.execute_reply":"2022-12-11T05:25:37.950115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just like the \"kde\" graph of the H_width data in the train_df dataframe, we visualized a tall spike in the 4000 and 5000 range in the L_width data (which means more data in this range) and saw some little bumps from 0 to 1000ish (a little data in this range).","metadata":{}},{"cell_type":"markdown","source":"Now let's create a scatter plot out of the Livingston and Handford widths combined in a scatter plot in Plotly! We define the fig variable to the scatter function from the px module to configure that we're creating a scatter diagram, setting the train_df dataframe inside, the x parameter to \"L_width\" and the y parameter to \"H_width\". Once again, we customize our format of the fig graph with the update_layout function, configuring the template parameter to ggplot2. After all of this, we exhibit the fig variable figure with the show function.","metadata":{}},{"cell_type":"code","source":"fig = px.scatter(train_df, x=\"L_width\", y=\"H_width\")\nfig.update_layout(template=\"ggplot2\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:37.953133Z","iopub.execute_input":"2022-12-11T05:25:37.953657Z","iopub.status.idle":"2022-12-11T05:25:38.053897Z","shell.execute_reply.started":"2022-12-11T05:25:37.953606Z","shell.execute_reply":"2022-12-11T05:25:38.052731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And just after that, we see a cluster of data on the top-right corner, thus seeing some three plots of data in the lower-left corner. In other words, this scatter plot diagram we created shows that there's more data ranging from between 4000 to 5000 on the Livingston width and in around 4300 to 4500 in the Handford width than the data between 718 to 1095 on the Livingston width and 718 to 1095 on the Handford width.","metadata":{}},{"cell_type":"markdown","source":"Now let's visualize the kernel density estimation of the scatter plot we plotted previously in Plotly! To do that, we use the kdeplot function from the sns module to create a kernel density estimation plot, setting the data parameter to the train_df dataframe for dataframe input, the x parameter to the L_width data column for the x-axis configuration, and the y parameter to the H_width data column for the y-axis configuration.","metadata":{}},{"cell_type":"code","source":"sns.kdeplot(data=train_df, x=\"L_width\", y=\"H_width\")","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:38.055321Z","iopub.execute_input":"2022-12-11T05:25:38.055673Z","iopub.status.idle":"2022-12-11T05:25:38.85882Z","shell.execute_reply.started":"2022-12-11T05:25:38.055643Z","shell.execute_reply":"2022-12-11T05:25:38.857644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After running the code cell just to display the kernel density estimation graph, we observed that there is a dense cluster on the top right corner. In other words, there are more data on the 4000 to 5000 range in both the H_width and L_width data column from the train_df dataframe.","metadata":{}},{"cell_type":"markdown","source":"Since we've done analyzing the train_df dataframe, let's move onwards to analyzing the test_df dataframe, and it contains the Livingston and Handford width data just like the train_df dataframe!","metadata":{}},{"cell_type":"markdown","source":"### test_df DataFrame Analysis\nPreviously, while we merged the train_df dataframe with the widths and heights of the Livingston and Handford data, we also merged them to the test_df dataframe. And just like that, let's begin the data analysis of them to the test_df dataframe!\n\nWithout further ado, let's visualize the Handford width distribution in a histogram with the Plotly way! First, we charcterize the fig variable to the histogram from the px module to configure a histogram diagram, setting the test_df dataframe as our dataframe input for the graph and the x parameter to the H_width data column for the x-axis configuration. Furthermore, we change how the way our fig variable graph look like with the update_layout function, setting the template parameter to ggplot2 so that our graph will look like a ggplot diagram. And once we finished our graph setup of the Handford distribution in a histogram, we exhibit the fig variable graph with the show function!","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(test_df, x=\"H_width\")\nfig.update_layout(template=\"ggplot2\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:38.860539Z","iopub.execute_input":"2022-12-11T05:25:38.861173Z","iopub.status.idle":"2022-12-11T05:25:38.944475Z","shell.execute_reply.started":"2022-12-11T05:25:38.861131Z","shell.execute_reply":"2022-12-11T05:25:38.94329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After the code cell above compiled the histogram distribution over the Handford widths from the test_df dataframe's H_width column, we sighted a huge spike of data distributions in the middle of our diagram we've created. In other words, the most counts in the H_width distribution is in the range of 4595 to 4599, with 212 data entities, while there's some distributions of the H_width data being counted once in the first and last parts of a diagram. Moreover to that, the massive spike in the H_width data distribution made us infer that the test_df dataframe has more data than the train_df dataframe, as it contained the data mostly of unknown gravitional waves.\n\nAnd again, as we previously distributed the H_height from the train_df dataframe, we observed that only the 360 range of it was present along with the L_height data column, so we didn't do the distribution of the H_height and L_height data from the test_df dataframe to a histogram.","metadata":{}},{"cell_type":"markdown","source":"Let's proceed to visualize the L_width (Livingston Width) data column to a histogram with Plotly! We simply define the fig variable to the histogram function from the px module for configuring a histogram diagram, setting the test_df dataframe as our dataframe input and the x parameter to the L_width data column, and then set our fig variable graph's layout with the update_layout function's template parameter to ggplot2. With that, we display our fig variable with the show function!","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(test_df, x=\"L_width\")\nfig.update_layout(template=\"ggplot2\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:38.946424Z","iopub.execute_input":"2022-12-11T05:25:38.947627Z","iopub.status.idle":"2022-12-11T05:25:39.028745Z","shell.execute_reply.started":"2022-12-11T05:25:38.947564Z","shell.execute_reply":"2022-12-11T05:25:39.027509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After the code compiled the histogram graph out from Plotly, we visualized the spike of the Livingston width data distributed from the test_df dataframe, just like the Handford width data distribution in the previous histogram graph. In other words, the highest number of L_width data in this histogram counted is in the range from 4605 to 4609, with 228 data entities, while there are some ranges of the L_width data distribution were counted as single. Furthermore, that distribution of the L_width data in the histogram made us infer that the test_df dataframe has more data than the train_df dataframe, as it contained the data mostly of unknown gravitional waves, as we told this previously while distributing the Handford width in a histogram previously.","metadata":{}},{"cell_type":"markdown","source":"Let's shift to Seaborn over the kernel density estimation of the Handford and Livingston width distribution from the test_df dataframe! Straightforwardly, we use the kdeplot from the sns module, setting the data parameter to the test_df dataframe for our dataframe input and setting the x parameter to the H_width and the L_width data columns separately from each code cell for setting up the x-axis configuration.","metadata":{}},{"cell_type":"code","source":"sns.kdeplot(data=test_df, x=\"H_width\")","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:39.030326Z","iopub.execute_input":"2022-12-11T05:25:39.030835Z","iopub.status.idle":"2022-12-11T05:25:39.288981Z","shell.execute_reply.started":"2022-12-11T05:25:39.030786Z","shell.execute_reply":"2022-12-11T05:25:39.287815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.kdeplot(data=test_df, x=\"L_width\")","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:39.290598Z","iopub.execute_input":"2022-12-11T05:25:39.291102Z","iopub.status.idle":"2022-12-11T05:25:39.550335Z","shell.execute_reply.started":"2022-12-11T05:25:39.291052Z","shell.execute_reply":"2022-12-11T05:25:39.54908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And just like that, we observed that the distribution of the H_width (Handford) and the L_width (Livingston) in the \"Kernel Density Estimation\" diagram has the crest of more data in the middle, as reinstated that the test_df dataframe has more data than the train_df dataframe because of the unknown data regarding of gravitional waves.","metadata":{}},{"cell_type":"markdown","source":"Let's visualize the overall data of Handford and Livingston width in Plotly's scatter plot! Once again, we characterize the fig variable to the scatter function from the px module for configuring our scatter plot, setting the test_df dataframe as our data input, the x parameter to the H_width data column for the x-axis configuration, and the y parameter to the L_width data column for the y-axis configuration, thus updating our way the fig variable figure looks with the update_layout function, setting  the template parameter to ggplot2. Finally, we exhibit our fig variable graph with the show function!","metadata":{}},{"cell_type":"code","source":"fig = px.scatter(test_df, x=\"H_width\", y=\"L_width\")\nfig.update_layout(template=\"ggplot2\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:39.55205Z","iopub.execute_input":"2022-12-11T05:25:39.552441Z","iopub.status.idle":"2022-12-11T05:25:39.650242Z","shell.execute_reply.started":"2022-12-11T05:25:39.552406Z","shell.execute_reply":"2022-12-11T05:25:39.649091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What we saw from the scatter plot we created is that there a cluster of data in the middle of the scatter plot, and by all of this, there are a lot of points from the H_width and L_width data from the test_df dataframe. And that shows us how the test_df dataframe contained massive data about gravitational waves, as there are undiscovered information about it currently.","metadata":{}},{"cell_type":"markdown","source":"Let's proceed to find the kernel density estimation of the overall data consisting the L_width and H_width data columns from the test_df dataframe with the \"Seaborn\" way! All we have to do is to use the sns module's kdeplot to plot our \"KDE\" chart, setting the data parameter to the test_df dataframe, the x parameter to the L_width column for the x-axis configuration, and the y parameter to the H_width column for the y-axis configuration.","metadata":{}},{"cell_type":"code","source":"sns.kdeplot(data=test_df, x=\"L_width\", y=\"H_width\")","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:39.651689Z","iopub.execute_input":"2022-12-11T05:25:39.652149Z","iopub.status.idle":"2022-12-11T05:25:44.586076Z","shell.execute_reply.started":"2022-12-11T05:25:39.652107Z","shell.execute_reply":"2022-12-11T05:25:44.584641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Wow, what an immersive look at the kernel density estimation plot! Like the scatter plot in Plotly, the kernel density estimation from the seaborn graph shows the similar, round, oval shape between the diagram. And the inner circle from this diagram shows that there are most data plots were clampered together because of the abundant data in the test_df dataframe.","metadata":{}},{"cell_type":"markdown","source":"In conclusion, we've finished the train_df and test_df dataframe analysis over the Handford and Livingston distributions one by one, in detail! Furthermore, we're are looking forward to analyze the wavy waves on the next section!","metadata":{}},{"cell_type":"markdown","source":"## Visualizing the Waves in a Spectrogram\nNow since we've mastered through analyzing the hdf5 files by Handford and Livingston width from the train_df and test_df dataframes, let's blast into the spectrogram analysis of each wave data from them! \n\nBefore that, we import the standard librosa module and the librosa module with the display submodule first for visualizing and plotting sound waves and spectrograms, followed by the matplotlib module with the pyplot submodule as plt, since we are uncertain about plotting them with Plotly. Furthermore, we need to import glob from the glob module for file management and stuff.","metadata":{}},{"cell_type":"code","source":"import librosa\nimport librosa.display\nimport matplotlib.pyplot as plt\nfrom glob import glob","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:44.587374Z","iopub.execute_input":"2022-12-11T05:25:44.587781Z","iopub.status.idle":"2022-12-11T05:25:46.230042Z","shell.execute_reply.started":"2022-12-11T05:25:44.587743Z","shell.execute_reply":"2022-12-11T05:25:46.228585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we import the two librosa modules, we define the read_data function, containing the DIR and df parameters inside. From this function, the dictionary variable called data_storage, to an empty dictionary format. After that, a try-except statement has been declared whether the directory is valid so that it can extract information from the directory.\n\nAnd from the try-except statement, the DIR parameter specified by the File function from the h5py function is read as the s variable, as the _id variable is defined to a list containing the keys of the s variable specified from the keys function by the list function with the slice index of 0 and then the data_storage dictionary's ID column is characterized to the _id variable. Following from that, the data_storage dictionary's freq column is defined to the array creation of the s variable's _id and frequency_Hz columns from the np module's array function because the function needs the data converted into the Numpy array for further processes. Later on, the Handford and Livingston dataset is generated when the data_storage's H1_SFTs, H1_timestamps, L1_SFTs, and L1_timestamps columns is created from the array of the s variable's _id, H1 (L1 for the L1_SFTs and L1_timestamps columns), and SFTs (timestamps_GPS for the second one) column by the np module's array function. Once finished creating the dataset over Handford and Livingston, the target column in the data_storage variable by defining it to the location of the df dataframe parameter that has the slice index of the df parameter's id column that is equal to the _id variable with the loc attribute and then apply the target column attribute, followed by the item function to return the first element of it as a Python scalar. At last in this with statement over reading the hdf5 files, the data_storage is returned. Meanwhile in the except part, which it has any Exception as the e variable, the e variable is printed, and then the message over the File path input check followed by returning nothing (None).","metadata":{}},{"cell_type":"code","source":"def read_data(DIR, df):\n    data_storage = {}\n    \n    try:\n        with h5py.File(DIR, 'r') as s:\n            _id = list(s.keys())[0]\n            data_storage['ID'] = _id\n            data_storage['freq'] = np.array(s[_id]['frequency_Hz']) \n            \n            # Handford\n            data_storage['H1_SFTs'] = np.array(s[_id]['H1']['SFTs'])\n            data_storage['H1_timestamps'] = np.array(s[_id]['H1']['timestamps_GPS'])\n            \n            # Livingston         \n            data_storage['L1_SFTs'] = np.array(s[_id]['L1']['SFTs'])\n            data_storage['L1_timestamps'] = np.array(s[_id]['L1']['timestamps_GPS'])\n            \n            data_storage['Target'] = df.loc[df.id==_id].target.item()\n            \n            return data_storage\n    except Exception as e:\n        print(str(e) + \"\\n\" + \"Please enter a valid path and key!!!\")\n        return None","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:46.23289Z","iopub.execute_input":"2022-12-11T05:25:46.233488Z","iopub.status.idle":"2022-12-11T05:25:46.244222Z","shell.execute_reply.started":"2022-12-11T05:25:46.233434Z","shell.execute_reply":"2022-12-11T05:25:46.243026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Note:** This code cell above was from Crookenden's notebook over plotting spectrograms, https://www.kaggle.com/code/edwardcrookenden/g2net-getting-started-eda#Plotting-Spectograms-%F0%9F%93%8A","metadata":{}},{"cell_type":"markdown","source":"After we created the read_data function, we define the glob_train and the glob_test variables to list out the training_dir and testing_dir variables and the \"/*\" string by the glob function and then define another two variables: train_data and test_data, to the function call of read_data, containing any index of the glob_train variable (glob_test for test_data) specified by a given number and the train_df dataframe (test_df dataframe for test_data).","metadata":{}},{"cell_type":"code","source":"glob_train = glob(training_dir+'/*')\nglob_test = glob(testing_dir+'/*')\n\ntrain_data = read_data(glob_train[4], train_df)\ntest_data = read_data(glob_test[4], test_df)","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:46.252261Z","iopub.execute_input":"2022-12-11T05:25:46.252718Z","iopub.status.idle":"2022-12-11T05:25:47.500005Z","shell.execute_reply.started":"2022-12-11T05:25:46.252683Z","shell.execute_reply":"2022-12-11T05:25:47.498833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we set up the data inside the train_data and test_data variables, let's plot them separately by Handford and Livingston waves category for the training data we have! Before that, we set up our figure size with the figure function from the plt module, which we configured the figsize parameter to 24 by 26. Now, we create our subplot with 4 rows, 3 columns and set their index to 1 with the plt module's subplot then title our first graph with the title function provided by the plt module, setting it to \"Handford Fourier Transforms\", then we characterize the sg_mag and sg_phase to the magphase function from the librosa module for configuring the exponent for the magnitude spectrogram, setting the train_data variable with the H1_SFTs column inside of it. Next, we create another variable, sg1, to the melspectrogram function from the librosa module and the feature attribute for configuring our spectrogram into our graph, setting the S parameter to the sg_mag variable for specifying the input for the spectrogram and once configured, we display the specshow of the sg1 variable from the librosa module that has the display attribute and the specshow function with the display function. Now for the second graph, we set up another subplot with the plt module's subplot function and its like the same as the first subplot we've created, but the index for the second subplot is set to 2. Another spectrogram graph is also plotted, but the title specified by the plt module with the title function is set to \"Livingston Fourier Transforms\" and the train_data variable with the L1_SFTs column is put in the librosa module with the magphase function to the sg_mag and sg_phase variables.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(22, 24))\nplt.subplot(4,3,1)\nplt.title(\"Handford Fourier Transforms\")\nsg_mag, sg_phase = librosa.magphase(train_data['H1_SFTs'])\nsg1 = librosa.feature.melspectrogram(S=sg_mag)\ndisplay(librosa.display.specshow(sg1))\nplt.subplot(4,3,2)\nplt.title(\"Livingston Fourier Transforms\")\nsg_mag, sg_phase = librosa.magphase(train_data['L1_SFTs'])\nsg1 = librosa.feature.melspectrogram(S=sg_mag)\ndisplay(librosa.display.specshow(sg1))","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:25:47.502868Z","iopub.execute_input":"2022-12-11T05:25:47.503311Z","iopub.status.idle":"2022-12-11T05:25:49.470805Z","shell.execute_reply.started":"2022-12-11T05:25:47.503276Z","shell.execute_reply":"2022-12-11T05:25:49.469481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's visualize the testing data by plotting them into the histogram! We vaguely follow from what we did for plotting the training data into the histogram, but we use the test_data variable's H1_SFTs and L1_SFTs columns into the magphase function from the librosa module to the sg_mag and sg_phase variables for configuring the spectrogram exponents.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(22, 24))\nplt.subplot(4,3,1)\nplt.title(\"Handford Fourier Transforms\")\nsg_mag, sg_phase = librosa.magphase(test_data['H1_SFTs'])\nsg1 = librosa.feature.melspectrogram(S=sg_mag)\ndisplay(librosa.display.specshow(sg1))\nplt.subplot(4,3,2)\nplt.title(\"Livingston Fourier Transforms\")\nsg_mag, sg_phase = librosa.magphase(test_data['L1_SFTs'])\nsg1 = librosa.feature.melspectrogram(S=sg_mag)\ndisplay(librosa.display.specshow(sg1))","metadata":{"execution":{"iopub.status.busy":"2022-12-11T05:51:59.320501Z","iopub.execute_input":"2022-12-11T05:51:59.321057Z","iopub.status.idle":"2022-12-11T05:52:01.024068Z","shell.execute_reply.started":"2022-12-11T05:51:59.321011Z","shell.execute_reply":"2022-12-11T05:52:01.022699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the spectrogram visualization in train and test data, we noticed some blurry and pixelated waves in the spectrogram, making this not clear to understand. That shows us how we can't see gravitational waves as the train and test data from the hdf5 files contained more noise once plotted, as visualizing where the gravitational waves were is an ongoing mystery for researchers at the European Gravitational Observatory.","metadata":{}},{"cell_type":"markdown","source":"And with all of the training and testing data plotted into a spectrogram with help from the librosa module, we've finished all of our data analysis on gravitational waves in our notebook!","metadata":{}},{"cell_type":"markdown","source":"## Conclusion\nWhile we surfed through our data analysis about the gravitational waves, We learned out that gravitational waves are too hard to understand, even after the scientists in the European Gravitational Observatory discovered the waves back in 2015. However, the ongoing research on the mysterious, unclear gravitational waves will become obvious in the near future as more observations will be presented in the talks from the scientists that studied the wave signals. And in other words, we can expect another research challenges on space, nature, and health after the researchers discover gravitational waves!","metadata":{}}]}