{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"font-family:verdana;\"> <center>🚀 G2Net Getting Started 🚀</center> </h1>\n\n***","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#0daae3;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n        <p style=\"padding: 10px;\n              color:white;\">\n            Hopefully this notebook will give you a basic understanding of the task and data involved in this competition. Please give an upvote if you find it useful 👍\n        </p>\n    </div>\n    \n<div align = 'center'><img src= \"https://www.g2net.eu/wp-content/uploads/2019/07/2ndconference_V2-1170x600.jpg\" alt =\"Space\" style='width: 1000px;height 500px'>","metadata":{}},{"cell_type":"markdown","source":"### <span style=\"font-family:verdana; word-spacing:1.5px;\"> Contents:\n[Load in the data ⏳](#first-bullet)\n\n[What is HDF5 🤔](#second-bullet)\n    \n[Plotting Spectograms 📊](#third-bullet)   \n\n[What do the real and imaginary parts of a Short-time Fourier Transforms mean? 👻](#fourth-bullet)   \n\n[Timestamp analysis ⏱](#fith-bullet)\n\n[Generating more data ⚙️](#sixth-bullet)   \n\n[Baseline 📈](#seventh-bullry)","metadata":{"execution":{"iopub.status.busy":"2022-10-06T22:32:53.944249Z","iopub.execute_input":"2022-10-06T22:32:53.944766Z","iopub.status.idle":"2022-10-06T22:32:53.95639Z","shell.execute_reply.started":"2022-10-06T22:32:53.944726Z","shell.execute_reply":"2022-10-06T22:32:53.954342Z"}}},{"cell_type":"markdown","source":"### <span style=\"font-family:verdana; word-spacing:1.5px;\">  Task overview\n\n<span style=\"font-family:verdana; word-spacing:1.5px;\"> The goal of this competition is to find continuous gravitational-wave signals. You will develop a model sensitive enough to detect weak yet long-lasting signals emitted by rapidly-spinning neutron stars within noisy data. \n\n<span style=\"font-family:verdana; word-spacing:1.5px;\"> Each sample is comprised of a set of Short-time Fourier Transforms (SFTs) and corresponding time stamps for each interferometer. The SFTs are not always contiguous in time, since the interferometers are not continuously online.\n\n<span style=\"font-family:verdana; word-spacing:1.5px;\"> Each data sample contains either  <span style=\"color:#159364;\">real or simulated noise</span> and possibly <span style=\"color:#159364;\">a simulated continuous gravitational-wave signal (CW) </span>. The task is to <span style=\"color:#159364;\"> identify when a signal is present </span> in the data (target=1).\n","metadata":{}},{"cell_type":"markdown","source":"## Imports 🗂","metadata":{}},{"cell_type":"code","source":"%%capture\n!pip install nexusformat","metadata":{"execution":{"iopub.status.busy":"2022-10-06T23:37:31.542059Z","iopub.execute_input":"2022-10-06T23:37:31.542549Z","iopub.status.idle":"2022-10-06T23:37:42.483636Z","shell.execute_reply.started":"2022-10-06T23:37:31.542508Z","shell.execute_reply":"2022-10-06T23:37:42.481997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nfrom pathlib import Path\nimport os\nimport random\nimport numpy as np\nfrom pprint import pprint\nfrom tqdm.notebook import tqdm\n\nimport seaborn as sns\nsns.set_theme()\n\nfrom scipy import signal\nfrom scipy.fft import fftshift\nimport matplotlib.pyplot as plt\n\nimport h5py\nimport nexusformat.nexus as nx\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\nfrom IPython.display import HTML, display\ndisplay(HTML('<style>.font-family:verdana; word-spacing:1.5px;</style>'))","metadata":{"execution":{"iopub.status.busy":"2022-10-06T23:40:53.971652Z","iopub.execute_input":"2022-10-06T23:40:53.972061Z","iopub.status.idle":"2022-10-06T23:40:53.984391Z","shell.execute_reply.started":"2022-10-06T23:40:53.972026Z","shell.execute_reply":"2022-10-06T23:40:53.982718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATA_PATH = Path('../input/g2net-detecting-continuous-gravitational-waves')\nTRAIN_PATH = DATA_PATH/'train'\nTEST_PATH = DATA_PATH/'test'\ntrain_example_with_signal_path = TRAIN_PATH/'cc561e4fc.hdf5'\ntrain_example_without_signal_path = TRAIN_PATH/'fb6db0d08.hdf5'","metadata":{"execution":{"iopub.status.busy":"2022-10-06T23:45:19.174964Z","iopub.execute_input":"2022-10-06T23:45:19.176326Z","iopub.status.idle":"2022-10-06T23:45:19.182998Z","shell.execute_reply.started":"2022-10-06T23:45:19.176263Z","shell.execute_reply":"2022-10-06T23:45:19.181615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# HDF5 🤔","metadata":{}},{"cell_type":"markdown","source":"<span style=\"font-family:verdana; word-spacing:1.5px;\"> The data is stored in HDF5 files. Hierarchical Data Format version 5 (HDF5), is an open source file format that supports large, complex, heterogeneous data. HDF5 uses a \"file directory\" like structure that allows you to organize data.\n\n<span style=\"font-family:verdana; word-spacing:1.5px;\"> HDF5 simplifies the file structure to include only two major types of object:\n- <span style=\"font-family:verdana; word-spacing:1.5px;\"> Datasets, which are typed multidimensional arrays\n- <span style=\"font-family:verdana; word-spacing:1.5px;\"> Groups, which are container structures that can hold datasets and other groups\n\n<span style=\"font-family:verdana; word-spacing:1.5px;\"> We can view the structure of our data with the nexusformat package.\n\n<span style=\"font-family:verdana; word-spacing:1.5px;\"> The `h5py` package is a thin, pythonic wrapper around HDF5 we can use it to quickly load our data.","metadata":{}},{"cell_type":"markdown","source":"<div align = 'center'><img src= \"https://cdn-images-1.medium.com/max/4984/1*BXG3eNq7xZGskaXmn53kvQ.png\" alt =\"Space\" style='width: 800px;height 400px'>","metadata":{}},{"cell_type":"markdown","source":"##  <span style=\"font-family:verdana; word-spacing:1.5px;\">  We can see the structure of the training and test data below:","metadata":{}},{"cell_type":"markdown","source":"- <span style=\"font-family:verdana; word-spacing:1.5px;\">  `ID` is the top group of the HDF5 file and links the datapoint to it's label in the `train_labels` csv  (group)\n\n- <span style=\"font-family:verdana; word-spacing:1.5px;\">  `frequency_Hz` contains the range frequencies measured by the dectors (dataset)\n\n\n- <span style=\"font-family:verdana; word-spacing:1.5px;\">  `H1` contains the data for the LIGO Hanford decector (group) \n    \n    - <span style=\"font-family:verdana; word-spacing:1.5px;\">  `SFTs` is the Short-time Fourier Transforms amplitudes for each timestamp at each frequency (dataset)\n    - <span style=\"font-family:verdana; word-spacing:1.5px;\">  `timestamps` contains the timestamps for the measurement (dataset)\n\n    \n- <span style=\"font-family:verdana; word-spacing:1.5px;\">  `L1` contains the data for the LIGO Livingston decector (group) \n    \n    - <span style=\"font-family:verdana; word-spacing:1.5px;\">  `SFTs` is the Short-time Fourier Transforms amplitudes for each timestamp at each frequency (dataset)\n    - <span style=\"font-family:verdana; word-spacing:1.5px;\">  `timestamps` contains the timestamps for the measurement (dataset)    \n    \n<span style=\"font-family:verdana; word-spacing:1.5px;\"> This structure can be visualised below 👇 ","metadata":{}},{"cell_type":"markdown","source":"![](https://i.imgur.com/M6xfOri.png)","metadata":{}},{"cell_type":"markdown","source":"##  <span style=\"font-family:verdana; word-spacing:1.5px;\"> Let's look at an example","metadata":{}},{"cell_type":"code","source":"with h5py.File(train_example_with_signal_path, \"r\") as f:\n    \n    # get first object name/key; this is the data point ID\n    ID_key = list(f.keys())[0]\n    print(f\"ID: {ID_key} \\n\")\n    \n    # Retrieve the Livingston decector data\n    print(f\"- {list(f[ID_key].keys())[1]}\")\n    L1_SFTs = f[ID_key]['L1']['SFTs']\n    print(f\"-- SFTs amplitudes: {L1_SFTs.shape}\")\n    L1_ts = f[ID_key]['L1']['timestamps_GPS']\n    print(f\"-- timestamps: {L1_ts.shape} \\n\")\n    \n    # Retrieve the Hanford decector data\n    print(f\"- {list(f[ID_key].keys())[0]}\")\n    H1_SFTs = f[ID_key]['H1']['SFTs']\n    print(f\"-- SFTs amplitudes: {H1_SFTs.shape}\")\n    H1_ts = f[ID_key]['H1']['timestamps_GPS']\n    print(f\"-- timestamps: {H1_ts.shape} \\n\")\n\n    # Retrieve the frequency data\n    freq_data = np.array(f[ID_key]['frequency_Hz'])\n    print(f\"- Frequency data: {freq_data.shape} \\n\")","metadata":{"execution":{"iopub.status.busy":"2022-10-06T23:48:12.872478Z","iopub.execute_input":"2022-10-06T23:48:12.873034Z","iopub.status.idle":"2022-10-06T23:48:12.914239Z","shell.execute_reply.started":"2022-10-06T23:48:12.872986Z","shell.execute_reply":"2022-10-06T23:48:12.912666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load in the meta data ⏳","metadata":{}},{"cell_type":"code","source":"train_labels = pd.read_csv(DATA_PATH/'train_labels.csv')\ntrain_labels.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-06T23:39:32.793994Z","iopub.execute_input":"2022-10-06T23:39:32.795135Z","iopub.status.idle":"2022-10-06T23:39:32.820822Z","shell.execute_reply.started":"2022-10-06T23:39:32.795073Z","shell.execute_reply":"2022-10-06T23:39:32.818459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Plot Train Test split ###\n\nplt.figure(figsize=(6,4))\nsns.barplot(['Train', 'Test'], [len(os.listdir(TRAIN_PATH)), len(os.listdir(TEST_PATH))]);\nplt.title(f'Train test split', fontsize=12)\nplt.ylabel('Count', fontsize=12)\nplt.xlabel('Category', fontsize=12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-06T23:41:33.289273Z","iopub.execute_input":"2022-10-06T23:41:33.289689Z","iopub.status.idle":"2022-10-06T23:41:33.496846Z","shell.execute_reply.started":"2022-10-06T23:41:33.289652Z","shell.execute_reply":"2022-10-06T23:41:33.495585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"font-family:verdana; word-spacing:1.5px;\">  This is a very unusual train test split for a kaggle competetion. Normally we would see at an 80:20 split, but here it is 1:16. The competition hosts are encouraging us to generate more data ourselves ([see discussion](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/347052))","metadata":{}},{"cell_type":"markdown","source":"## <span style=\"font-family:verdana; word-spacing:1.5px;\">  Labels\n<span style=\"font-family:verdana; word-spacing:1.5px;\">  Each data sample contains either **real or simulated noise** and possibly a **simulated continuous gravitational-wave signal** (CW). The task is to identify when a signal is present in the data (target=1)\n\n<span style=\"font-family:verdana; word-spacing:1.5px;\">  The `target_labels.csv` is a file containing the target labels; 1 if the data contains the presence of a gravitational wave, 0 otherwise. (Please note the presence of a small number of files labeled -1. Physicists are currently unable to determine the status of these files.)","metadata":{}},{"cell_type":"code","source":"labels_df = pd.read_csv(DATA_PATH/'train_labels.csv')\nlabels_df.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-10-06T23:41:41.086244Z","iopub.execute_input":"2022-10-06T23:41:41.087357Z","iopub.status.idle":"2022-10-06T23:41:41.106105Z","shell.execute_reply.started":"2022-10-06T23:41:41.087301Z","shell.execute_reply":"2022-10-06T23:41:41.105018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### Plot the distribution of labels ###\nlabel_count  = labels_df['target'].value_counts()\nplt.figure(figsize=(10,8))\nsns.barplot(label_count.index, label_count.values, alpha=0.7)\nplt.title(f'Frequency of labels in training data', fontsize=12)\nplt.ylabel('Count', fontsize=12)\nplt.xlabel('label', fontsize=12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-06T23:41:41.233418Z","iopub.execute_input":"2022-10-06T23:41:41.234085Z","iopub.status.idle":"2022-10-06T23:41:41.471671Z","shell.execute_reply.started":"2022-10-06T23:41:41.234035Z","shell.execute_reply":"2022-10-06T23:41:41.470783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Sample submission:\n`sample_submission.csv` - a sample submission file in the correct format; your task is, for each file in the test/ folder, to predict the probability between [0, 1] of it it containing a continuous gravitational wave signal.","metadata":{"execution":{"iopub.status.busy":"2022-10-04T17:02:30.474744Z","iopub.execute_input":"2022-10-04T17:02:30.475143Z","iopub.status.idle":"2022-10-04T17:02:30.482773Z","shell.execute_reply.started":"2022-10-04T17:02:30.475112Z","shell.execute_reply":"2022-10-04T17:02:30.481314Z"}}},{"cell_type":"code","source":"pd.read_csv(DATA_PATH/'sample_submission.csv').head()","metadata":{"execution":{"iopub.status.busy":"2022-10-06T21:38:36.942631Z","iopub.execute_input":"2022-10-06T21:38:36.943127Z","iopub.status.idle":"2022-10-06T21:38:36.968963Z","shell.execute_reply.started":"2022-10-06T21:38:36.943078Z","shell.execute_reply":"2022-10-06T21:38:36.967392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plotting Spectograms 📊","metadata":{}},{"cell_type":"code","source":"### Helper functions to extract data and plot spectograms ###\n\ndef extract_data_from_hdf5(path, labels):\n    '''\n    Extracts data from hdf5 file and puts it into a dict. It also adds the label\n    '''\n    \n    data = {}\n    \n    with h5py.File(path, \"r\") as f:\n\n        ID_key = list(f.keys())[0]\n\n        # Retrieve the frequency data\n        data['freq'] = np.array(f[ID_key]['frequency_Hz'])\n\n        # Retrieve the Livingston decector data\n        data['L1_SFTs_amplitudes'] = np.array(f[ID_key]['L1']['SFTs'])\n        data['L1_ts'] = np.array(f[ID_key]['L1']['timestamps_GPS'])\n\n        # Retrieve the Livingston decector data\n        data['H1_SFTs_amplitudes'] = np.array(f[ID_key]['H1']['SFTs'])\n        data['H1_ts'] = np.array(f[ID_key]['H1']['timestamps_GPS'])\n        \n        # Get label from training labels if in training set\n        data['label'] = labels.loc[labels.id==ID_key].target.item()\n        \n    return data\n    \ndef plot_spectograms(data):\n    '''\n    Shows the real and imaginary amplitudes of the SFTs as spectograms for both detectors\n    '''\n    \n    fig, ax = plt.subplots(2, 2, figsize=(16, 10))\n    fig.suptitle(f\"Label {data['label']}\")\n\n    for ind, detector in enumerate(['L1', 'H1']):\n        ax[ind][0].set(xlabel=\"Timestamps [GPS]\",\n                         ylabel=\"Frequency [Hz]\",\n                         title=f\"{detector} - Real part\")\n        ax[ind][1].set(xlabel=\"Timestamps [GPS]\",\n                         ylabel=\"Frequency [Hz]\",\n                         title=f\"{detector} - Imaginary part\")\n        \n        \n        c0 = ax[ind][0].pcolormesh(data[f\"{detector}_ts\"], data['freq'],\n                                     data[f\"{detector}_SFTs_amplitudes\"].real)\n        c1 = ax[ind][1].pcolormesh(data[f\"{detector}_ts\"], data['freq'],\n                                     data[f\"{detector}_SFTs_amplitudes\"].imag)\n    \n        fig.colorbar(c0, ax=ax[ind][0])\n        fig.colorbar(c1, ax=ax[ind][1])\n        \n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-06T23:45:13.828984Z","iopub.execute_input":"2022-10-06T23:45:13.829425Z","iopub.status.idle":"2022-10-06T23:45:13.846897Z","shell.execute_reply.started":"2022-10-06T23:45:13.82939Z","shell.execute_reply":"2022-10-06T23:45:13.845076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <span style=\"font-family:verdana; word-spacing:1.5px;\"> Lets plot some spectograms. One with a simulated signal and one without!","metadata":{}},{"cell_type":"code","source":"sns.reset_orig() # Reset seaborn theme (otherwise these plots come out red 😅)\ndata = extract_data_from_hdf5(train_example_with_signal_path, labels_df)\nplot_spectograms(data)","metadata":{"execution":{"iopub.status.busy":"2022-10-06T23:45:24.50476Z","iopub.execute_input":"2022-10-06T23:45:24.505235Z","iopub.status.idle":"2022-10-06T23:45:30.034414Z","shell.execute_reply.started":"2022-10-06T23:45:24.505197Z","shell.execute_reply":"2022-10-06T23:45:30.033469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = extract_data_from_hdf5(train_example_without_signal_path, labels_df)\nplot_spectograms(data)","metadata":{"execution":{"iopub.status.busy":"2022-10-06T23:45:44.211204Z","iopub.execute_input":"2022-10-06T23:45:44.211671Z","iopub.status.idle":"2022-10-06T23:45:49.441501Z","shell.execute_reply.started":"2022-10-06T23:45:44.211634Z","shell.execute_reply":"2022-10-06T23:45:49.440014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### <span style=\"font-family:verdana; word-spacing:1.5px;\"> Very hard to notice any difference at all 🤔","metadata":{}},{"cell_type":"markdown","source":"# What do the real and imaginary parts of a Short-time Fourier Transforms mean? 👻","metadata":{}},{"cell_type":"markdown","source":"# Timestamp analysis ⏱\n\nSince the continuous gravitational wave is simulated I am not sure how the length of it is determined (it seems to be hard coded in the [generating signals nb](https://github.com/PyFstat/PyFstat/blob/ec86602bb2f93238492a7242ad90995f6654eab7/examples/tutorials/1_generating_signals.ipynb)). As a result, I assume that the timestamp data is relatively meaningless... for completness I will include a few plots :)","metadata":{}},{"cell_type":"code","source":"### Extract timestamp data from training data ###\n\nH1_timestamps, L1_timestamps, start_diff, labels, freq = ([] for i in range(5))\n\nfor p in tqdm(os.listdir(TRAIN_PATH), total=len(os.listdir(TRAIN_PATH))):\n    id_ = p.split('.')[0]\n    labels.append(labels_df.loc[labels_df.id==id_].target.item())\n    data = extract_data_from_hdf5(DATA_PATH/'train'/p, labels_df)\n    L1_timestamps.append(data['L1_ts'])\n    H1_timestamps.append(data['H1_ts'])\n    start_diff.append(data['L1_ts'][0] - data['H1_ts'][0])\n    freq.append(data['freq'])","metadata":{"execution":{"iopub.status.busy":"2022-10-06T23:52:13.941198Z","iopub.execute_input":"2022-10-06T23:52:13.941725Z","iopub.status.idle":"2022-10-06T23:54:47.497103Z","shell.execute_reply.started":"2022-10-06T23:52:13.941685Z","shell.execute_reply":"2022-10-06T23:54:47.494897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.DataFrame({'label':labels, 'L1_timestamp_length':[len(i) for i in L1_timestamps], 'H1_timestamp_length':[len(i) for i in H1_timestamps], 'Differnce in start time between detectors':start_diff})\ndf = df[df.label!=-1]","metadata":{"execution":{"iopub.status.busy":"2022-10-06T23:54:47.501637Z","iopub.execute_input":"2022-10-06T23:54:47.502766Z","iopub.status.idle":"2022-10-06T23:54:47.517987Z","shell.execute_reply.started":"2022-10-06T23:54:47.502694Z","shell.execute_reply":"2022-10-06T23:54:47.516725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set_theme()\n\nfig, ax = plt.subplots(1,3, figsize=(24,8))\nfig.suptitle(f\"In the plots the distribution of timestamps for both classes are shown; 1 indicates a simulated CW present and 0 not present\", fontsize=16)\nsns.histplot(\n        df, x=\"L1_timestamp_length\", hue=\"label\",\n        stat=\"density\", common_norm=False, bins=20, ax=ax[0], kde=True).set_title('Length of measurement for Livingston detector', fontsize=16);\n\nsns.histplot(\n        df, x=\"H1_timestamp_length\", hue=\"label\",\n        stat=\"density\", common_norm=False, bins=20, ax=ax[1], kde=True).set_title('Length of measurement for Hanford detector', fontsize=16);\n\nsns.histplot(\n        df, x=\"Differnce in start time between detectors\", hue=\"label\",\n        stat=\"density\", common_norm=False, bins=20, ax=ax[2], kde=True).set_title('Difference in starting timestamp between detectors', fontsize=16);","metadata":{"execution":{"iopub.status.busy":"2022-10-07T00:02:18.85839Z","iopub.execute_input":"2022-10-07T00:02:18.8607Z","iopub.status.idle":"2022-10-07T00:02:20.469646Z","shell.execute_reply.started":"2022-10-07T00:02:18.860636Z","shell.execute_reply":"2022-10-07T00:02:20.468273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Frequency analysis 📳\n\nDo the detectors measure the same range of frequencies each time? No. The range varies but the minimum freq measured is ~50 Hz and the maximum is ~498 Hz","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(14,6))\nsns.histplot(x=list(np.hstack(freq)), binwidth=20)\nplt.title('Histogram of the range of Frequencies detected');\nplt.xlabel('Frequency Hz')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-07T00:02:25.331006Z","iopub.execute_input":"2022-10-07T00:02:25.331973Z","iopub.status.idle":"2022-10-07T00:02:25.798074Z","shell.execute_reply.started":"2022-10-07T00:02:25.331924Z","shell.execute_reply":"2022-10-07T00:02:25.796671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Generating more data ⚙️","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Types of Noise\n- Non-stationary noise\n- Gaps\n- Narrow instrumental artifacts\n- Multiple detectors\n\nTypes of signal generation","metadata":{}},{"cell_type":"markdown","source":"# Baseline 📈","metadata":{}},{"cell_type":"markdown","source":"# Sample Submission","metadata":{}},{"cell_type":"code","source":"samp_sub = pd.read_csv(DATA_PATH/'sample_submission.csv')\nsamp_sub['target'] = 0.55\nsamp_sub.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-10-06T14:45:51.731918Z","iopub.execute_input":"2022-10-06T14:45:51.732216Z","iopub.status.idle":"2022-10-06T14:45:51.763886Z","shell.execute_reply.started":"2022-10-06T14:45:51.732188Z","shell.execute_reply":"2022-10-06T14:45:51.762784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# WIP :)","metadata":{}}]}