{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🤝 Intro","metadata":{}},{"cell_type":"markdown","source":"Gravitational waves are really cool. They give us key insights into the deepest workings of our reality.","metadata":{}},{"cell_type":"markdown","source":"![gravitational waves](https://www.sbs.com.au/topics/sites/sbs.com.au.topics/files/styles/body_image/public/ligo-lab-gravity-waves.gif?itok=HUT-bBQ8&mtime=1471377425)","metadata":{}},{"cell_type":"markdown","source":"For this competition, this information is stored in a hdf5 file. Let's see how we can go about and read it.","metadata":{}},{"cell_type":"markdown","source":"# 🚚 Import","metadata":{}},{"cell_type":"code","source":"import h5py  # heavy lifting hdf5 reading\n\nimport numpy as np\nimport pandas as pd","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-10-04T16:21:45.102032Z","iopub.execute_input":"2022-10-04T16:21:45.102421Z","iopub.status.idle":"2022-10-04T16:21:45.30935Z","shell.execute_reply.started":"2022-10-04T16:21:45.102389Z","shell.execute_reply":"2022-10-04T16:21:45.30804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 📁 Our HDF5 Structure","metadata":{}},{"cell_type":"markdown","source":"Let's try reading one hdf5 file from the train folder and one from the test folder.","metadata":{}},{"cell_type":"code","source":"# train\ntest_filename = '../input/g2net-detecting-continuous-gravitational-waves/test/00076c5a6.hdf5'\nwith h5py.File(test_filename, \"r\") as f:\n    for file_key in f.keys():\n        group = f[file_key]\n        print(group)\n        try:\n            for group_key in group.keys():\n                group2 = group[group_key]\n                print(f\"---->{group2}\")\n                for group_key2 in group2.keys():\n                        print(f\"--------->{group2[group_key2]}\")\n        except AttributeError:\n            pass","metadata":{"execution":{"iopub.status.busy":"2022-10-04T16:21:45.93277Z","iopub.execute_input":"2022-10-04T16:21:45.933729Z","iopub.status.idle":"2022-10-04T16:21:45.968329Z","shell.execute_reply.started":"2022-10-04T16:21:45.933673Z","shell.execute_reply":"2022-10-04T16:21:45.966779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# test\ntest_filename = '../input/g2net-detecting-continuous-gravitational-waves/test/00076c5a6.hdf5'\nwith h5py.File(test_filename, \"r\") as f:\n    for file_key in f.keys():\n        group = f[file_key]\n        print(group)\n        try:\n            for group_key in group.keys():\n                group2 = group[group_key]\n                print(f\"---->{group2}\")\n                for group_key2 in group2.keys():\n                        print(f\"--------->{group2[group_key2]}\")\n        except AttributeError:\n            pass","metadata":{"execution":{"iopub.status.busy":"2022-10-04T16:21:46.365858Z","iopub.execute_input":"2022-10-04T16:21:46.3663Z","iopub.status.idle":"2022-10-04T16:21:46.395697Z","shell.execute_reply.started":"2022-10-04T16:21:46.366266Z","shell.execute_reply":"2022-10-04T16:21:46.394267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, our hdf5 files are composed of \"groups\" (like folders) and \"datasets\" (our actual data).\nEach group contains either other groups or datasets. What we're interested in is reading the datasets.\nWith np.array we can read the datasets.","metadata":{}},{"cell_type":"markdown","source":"# 📖 Reading HDF5 Datasets","metadata":{}},{"cell_type":"code","source":"# train\ntrain_filename = '../input/g2net-detecting-continuous-gravitational-waves/train/004f23b2d.hdf5'\ndatasets = []\nwith h5py.File(train_filename, \"r\") as f:\n    for file_key in f.keys():\n        group = f[file_key]\n        if isinstance(group, h5py._hl.dataset.Dataset):\n            datasets.append(np.array(group))\n            continue\n        for group_key in group.keys():\n            group2 = group[group_key]\n            if isinstance(group2, h5py._hl.dataset.Dataset):\n                datasets.append(np.array(group2))\n                continue\n            for group_key2 in group2.keys():\n                group3 = group2[group_key2]\n                if isinstance(group3, h5py._hl.dataset.Dataset):\n                    datasets.append(np.array(group3))\n                    continue","metadata":{"execution":{"iopub.status.busy":"2022-10-04T16:21:47.520108Z","iopub.execute_input":"2022-10-04T16:21:47.520994Z","iopub.status.idle":"2022-10-04T16:21:47.879859Z","shell.execute_reply.started":"2022-10-04T16:21:47.520956Z","shell.execute_reply":"2022-10-04T16:21:47.878623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(len(datasets))\nfor dataset in datasets:\n    print(dataset)","metadata":{"execution":{"iopub.status.busy":"2022-10-04T16:21:47.888582Z","iopub.execute_input":"2022-10-04T16:21:47.888964Z","iopub.status.idle":"2022-10-04T16:21:47.903163Z","shell.execute_reply.started":"2022-10-04T16:21:47.888934Z","shell.execute_reply":"2022-10-04T16:21:47.901897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This will be the input data for our model.  \nKeep in mind, this is just from one file, so this is a single data point for our model.","metadata":{}},{"cell_type":"markdown","source":"# 🎯 Target Variable","metadata":{}},{"cell_type":"markdown","source":"We also have a .csv that associates each file with the target label (for our training hdf5 files).  \n0 means no signal.  \n1 means signal","metadata":{}},{"cell_type":"code","source":"df_labels = pd.read_csv('../input/g2net-detecting-continuous-gravitational-waves/train_labels.csv')","metadata":{"execution":{"iopub.status.busy":"2022-10-04T16:21:49.320079Z","iopub.execute_input":"2022-10-04T16:21:49.321338Z","iopub.status.idle":"2022-10-04T16:21:49.341064Z","shell.execute_reply.started":"2022-10-04T16:21:49.321291Z","shell.execute_reply":"2022-10-04T16:21:49.339876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_labels","metadata":{"execution":{"iopub.status.busy":"2022-10-04T16:21:49.744571Z","iopub.execute_input":"2022-10-04T16:21:49.745054Z","iopub.status.idle":"2022-10-04T16:21:49.774106Z","shell.execute_reply.started":"2022-10-04T16:21:49.745008Z","shell.execute_reply":"2022-10-04T16:21:49.772664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So for our example training data point we opened before","metadata":{}},{"cell_type":"code","source":"df_labels.loc[df_labels['id'] == '004f23b2d']","metadata":{"execution":{"iopub.status.busy":"2022-10-04T16:21:52.800203Z","iopub.execute_input":"2022-10-04T16:21:52.800646Z","iopub.status.idle":"2022-10-04T16:21:52.819688Z","shell.execute_reply.started":"2022-10-04T16:21:52.800582Z","shell.execute_reply":"2022-10-04T16:21:52.818007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The signal is present.","metadata":{}},{"cell_type":"markdown","source":"# 👋 Conclusion","metadata":{}},{"cell_type":"markdown","source":"Hope this helps, this is looking like a very interesting and challenging competition!  \nAny feedback is welcome, this is a work in progress!","metadata":{}}]}