{
  "id": 373973,
  "title": "10 Days to go - Here's What You Need to Know",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/373973",
  "author_name": "The Devastator",
  "post_date": "2022-12-24T14:33:19.357000",
  "votes": 36,
  "comment_count": 6,
  "views": 0,
  "content": "<h3>What do we know so far?</h3>\n<h5>10 Days to go!</h5>\n<hr>\n<p>Save yourself some time and get caught up on the latest updates and important discussions that have taken place prior to the competition deadline.<br>\nThe competition ends in 10 days and there is still time to improve your solution. Here are some of the important topics that have been discussed in the Kaggle forums.</p>\n<hr>\n<h5>Background</h5>\n<ul>\n<li><p>The data is H-U-G-E and comes in hdf5 format, <a href=\"https://www.kaggle.com/chazzer\" target=\"_blank\">chazzer</a> had written a <a href=\"https://www.kaggle.com/code/chazzer/how-to-read-the-hdf5-files/\" target=\"_blank\">notebook on how to read the hdf5 files</a></p></li>\n<li><p>Shortly after <a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">JohnM</a> pushed this one step further and created <a href=\"https://www.kaggle.com/datasets/jpmiller/simplified-dataset\" target=\"_blank\">HDF files with indexed dataframes for the training set</a>. The index on each frame is the frequency array from the original file, and the columns are the timestamps.</p></li>\n</ul>\n<p><strong>Reading Example</strong></p>\n<pre><code>df_h = pd.read_hdf(, key = )\n</code></pre>\n<ul>\n<li>Great <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/357682\" target=\"_blank\">list</a> of learning resources had been published by <a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">Laura Fink</a><ul>\n<li><a href=\"https://www.gw-openscience.org/tutorials/\" target=\"_blank\">Gravitational Wave Open Science Center - Tutorials</a></li>\n<li><a href=\"https://gwpy.github.io/docs/stable/index.html\" target=\"_blank\">GWPy</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=tQ_teIUb3tE\" target=\"_blank\">Video about the Michelson Interferometer</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=FlDtXIBrAYE&amp;t=1s\" target=\"_blank\">Video about gravitational waves</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=spUNpyF58BY&amp;t=1s\" target=\"_blank\">Great tutorial video about Fourier transforms</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=s2K1JfNR7Sc\" target=\"_blank\">Video about Denoising data with Fast Fourier Transforms</a></li></ul></li>\n</ul>\n<hr>\n<h5>Best CV-LB Results (<a href=\"https://www.kaggle.com/e0xextazy\" target=\"_blank\">Mark Baushenko</a>)</h5>\n<ul>\n<li><strong>Important</strong> Comments by <a href=\"drhabib\" target=\"_blank\">DrHB</a><ul>\n<li>\"Without <code>augs</code> the gap (CV-LB) is bigger\"</li>\n<li>\"After running some experiments .. I think the biggest challenge of this competition will be to build a proper validation set\"    </li></ul></li>\n</ul>\n<hr>\n<h5>Understanding the Data (<a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">Ravi Shah</a>)</h5>\n<p>An <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/358444\" target=\"_blank\">amazing intro post</a> by <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">Ravi Shah</a> going through all important details about the data of this competition. </p>\n<p><strong>In short (check the original for more info)</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F1893bf6ae0ffb282512eec0fdce5a662%2Fstruc.PNG?generation=1665177635295503&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>L1 - data from the LIGO Livingston interferometer</li>\n<li>H1 - data from the LIGO Hanford interferometer</li>\n<li>STFs - Short-time Fourier Transforms (SFTs) - shape (360, n)</li>\n<li>timestamps - the timestamps that the STFs correspond to - shape (n,)</li>\n<li>frequency Hz - the range frequencies measured by the detectors - shape (360,)</li>\n</ul>\n<p><strong>Competition Goal</strong></p>\n<p>The goal of this competition is to create a model that can detect continuous gravitational-wave signals. These are weak yet long-lasting signals emitted by rapidly-spinning neutron stars within noisy data. </p>\n<blockquote>\n  <ul>\n  <li><strong>Target is 1 if signal is present else 0</strong></li>\n  </ul>\n</blockquote>\n<p><strong>LIGO</strong> (where the data comes from) is a detector in Livingston, Louisiana and Hanford, Washington. The detector has a L shape with equal-length arms connected to a corner station. When a gravitational wave passes by, one leg of the detector is shortened while the other is lengthened. </p>\n<p><strong>The interference makes a shift that we can analyze.</strong></p>\n<p><strong>We Get A Continuous Gravitational Waves</strong></p>\n<p>A continuous gravitational-wave signal from a Galactic neutron star will look almost perfectly constant in both frequency and amplitude.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F6b476a7c1f076bdf7c4f7234d37ace3b%2Fperfect%20cont%20wave.PNG?generation=1665181150773656&amp;alt=media\" alt=\"\"></p>\n<p><strong>But over long durations: The frequency of the signal slowly changes, because:</strong></p>\n<ul>\n<li>As the neutron star emits gravitational and electromagnetic waves, it loses energy which causes it to rotate more slowly</li>\n<li>The detector here on Earth is moving with respect to the neutron star. This changes the frequency of the gravitational waves observed in the detector.</li>\n</ul>\n<p><strong>Short-time Fourier Transforms (SFTs)</strong></p>\n<ul>\n<li>Short-time Fourier Transforms (SFTs) can be used as a way of quantifying the change of a nonstationary signal’s frequency and phase content over time.</li>\n<li>Very common type of signal preprocessing technique.</li>\n<li>This what we are given to work with (hdf5 files).</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fa8a7caafee2261f27fde68045a824202%2Fstft.PNG?generation=1665180281946268&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h5>Timestamps and Frequencies Insights (<a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">Mark Wijkhuizen</a>)</h5>\n<p><strong>There seems to be 2 differences of our data in comparison to conventional SFT's</strong></p>\n<ul>\n<li><p><strong>The time delta between each timestamp is inconsistent:</strong> Time between measurement in the SFT's can vary greatly.</p></li>\n<li><p>The frequency range is consistently ~0.2Hz, <strong>BUT the min/max frequency ranges from around ~50to ~498Hz.</strong></p></li>\n<li><p><strong>We should not interpret all spectrograms as continuous sounds:</strong> When taking a look at the timestamps they start at 2009-03-27/28 and end on 2009-07-25/26/27. The recordings are thus gathered during 4 months with often 30 minutes between the recordings. </p></li>\n</ul>\n<hr>\n<h5>Minimum float32 is 1e-38 and data**2 is 1e-44 (<a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">🐢 Jun Koda</a>)</h5>\n<p>Important <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361312\" target=\"_blank\">post</a> by <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">🐢 Jun Koda</a> sharing with us that since the smallest positive number float32 (~ 1e-38) but on the data it is ~1e-22. </p>\n<p><strong>Once we square the data, they go beyond this limit.</strong></p>\n<blockquote>\n  <p>np.float32(1.23456e-44) =&gt; 1.3e-44.</p>\n</blockquote>\n<ul>\n<li>float32 has about 8 significant digits and 1.23456e-44 will be expressed as 0.0000012e-38 with float32, only 2 digits remaining.</li>\n<li>Suggestion: that we multiply the data by 1e21 or 1e22.</li>\n</ul>\n<hr>\n<h5>PyFstat Signal Parameters (<a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">Mark Wijkhuizen</a>)</h5>\n<blockquote>\n  <p>Another great <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361562\" target=\"_blank\">post</a> by (<a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">Mark Wijkhuizen</a>)</p>\n</blockquote>\n<p>PyFstat has many parameters but it's documentation however is not very helpful for explaining them.</p>\n<p><strong><a href=\"https://www.kaggle.com/darkbitur\" target=\"_blank\">Victor Gonzalez</a></strong> Explained with an incredible comment everything about all the paramaeters. </p>\n<ul>\n<li>You can use a helper function to inject some parameters based on priors:</li>\n</ul>\n<pre><code>{**pyfstat.injection_parameters.isotropic_amplitude_priors}\n{: {: {: -, : }}, : {: {: -, : }}, : {: {: , : }}}\n</code></pre>\n<p><strong>More Parameters</strong></p>\n<ul>\n<li><strong>cosi:</strong> Cosine of the angle between the source and us. Range: [-1, 1]</li>\n<li><strong>psi and phi:</strong> polarization angle and phase. Ranges: [-pi/2, pi/2] and [0, 2*pi]</li>\n<li><strong>Alpha:</strong> Right ascension of the source's position on the sky. Range: [0, 2*pi]</li>\n<li><strong>Delta:</strong> Declination of the source's position on the sky. Range: [-pi/2, pi/2]</li>\n<li><strong>Band:</strong> This is just the width of the frequency \"slice\" to be taken. It's 0.2 hz for all files in the train and test datasets, which translates to 360 \"frequency lines\".</li>\n<li><strong>F0:</strong> This may be tricky. The generated noise will be centered around this frequency. If a signal is created, it will be also a signal with this frequency (at the source; when it arrives to the detector, doppler and other effects will shift the frequency), so it will appear more or less on the center of the data. Range used on the datasets: [50, 400] Hz.</li>\n<li><strong>F1:</strong> This is technically the first derivative of the frequency (how it changs over time). Physically, the sources (neutron stars) lose energy over time because they emit the GWs, so this should be a negative number. A reasonable range could be [-1e-9, 0] Hz but that's up to us.</li>\n<li><strong>F2:</strong> This is guess it's the second derivative, which maybe is 0 for the purposes of this competition.</li>\n<li><strong>h0:</strong> The amplitude of the signal. In the tutorials, the approach is to define it as a ratio of the sqrtSX. What makes the competition (and the real-world problem of finding continuous GWs) is that expected h0 will be between 1 and 2 orders of magnitude smaller than the amplitude of the noise. In my opinion, part of the challenge is to create your own signals with a distribution of h0 similar to the one used for the test set.</li>\n<li><strong>psi</strong> is actually defined within [-pi /4 , pi / 4] (and it's periodic, so anything beyond that folds back onto this range).</li>\n<li><strong>F2</strong> is another term describing the intrinsic evolution of a gravitational wave. As <a href=\"https://www.kaggle.com/darkbitur\" target=\"_blank\">@darkbitur</a> correctly guesses, it's set to 0 in this competition as we only focus on frequency and spindown.</li>\n<li><strong>tp, asini, period</strong> are parameters used to describe neutron stars in binary systems, and they affect the frequency evolution of the signal. You can safely ignore those in this competition (but do feel free to play around with them if you wish to!).</li>\n</ul>\n<hr>\n<h5>Can we assume the same signal trajectories between Hanford &amp; Livingston? (<a href=\"https://www.kaggle.com/ikarosilva\" target=\"_blank\">IkaroSilva</a>)</h5>\n<p>Question <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361655\" target=\"_blank\">asked</a> by <a href=\"https://www.kaggle.com/ikarosilva\" target=\"_blank\">IkaroSilva</a></p>\n<ul>\n<li>Can we assume the same signal trajectories between Hanford &amp; Livingston?</li>\n</ul>\n<p><strong>Answer (By The Hosts):</strong></p>\n<ul>\n<li><p>The frequency of a signal as seen by a detector a certain time t (what you call it S(t) in your post) does generally depend on the detector, the reason being that different detectors move \"differently\" around the Sun as they are in slightly different positions on Earth.</p></li>\n<li><p>This distinction may not be that noticeable for some cases if we use the two Advanced LIGO detectors, since they are relatively close to each other, but becomes quite apparent once you include other detectors such as Virgo, which is located in Italy.</p></li>\n</ul>\n<hr>\n<h5>What is your batch size and LB? (<a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">dragon zhang</a>)</h5>\n<p>A [good] question about the relation beween the batch-size and generalization performance by <a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">dragon zhang</a>.</p>\n<p><strong>Short Answer:</strong></p>\n<pre><code>Friends dont let friends use minibatches larger than 32.\n[Yann LeCun](https://twitter.com/ylecun/status/989610208497360896).\n</code></pre>\n<hr>\n<h5>Recap of the Top Solutions from the Previous G2Net Competition (<a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363280\" target=\"_blank\">Sinan Calisir</a>)</h5>\n<p>An amazing solutions <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363280\" target=\"_blank\">summary</a> by <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363280\" target=\"_blank\">Sinan Calisir</a> of all winning solutions of the previous competition.</p>\n<hr>\n<h5>Weird property of std(spectrogram) in train data ([Josef Slavicek</h5>\n<p>](<a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363460)\" target=\"_blank\">https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363460)</a>)</p>\n<p><strong>Wierd Observation:</strong> When computing std of spectrogram in complex64, I'm getting almost allways the same value of 1.4973568266036865e-22. If I do the same computation in complex128, the values starts do differ - and they are partly useful as indicator of target label:</p>\n<p><strong>Answer (From the Hosts):</strong></p>\n<p><strong>TL;DR:</strong> This is not a data leak, we understand the reason this is happening and it is expected. Nothing to worry about.</p>\n<blockquote>\n  <p>and they are partly useful as indicator of target label</p>\n</blockquote>\n<p>Yes, this is expected: stdev is computed from the data, and very strong signals will deviate the value away from the \"noise-only\" value. This should already tell you how \"strong\" signals in the test set are, by the way :) .</p>\n<blockquote>\n  <p>Around 80% of test data seems to be affected by this</p>\n</blockquote>\n<p>This is also well understood. We wanted to give as much real data as possible in this challenge, but we didn't have a reliable way of generating more than what the LIGO detectors produced during the run, and we did not want to run into issues about having too much correlation among different samples. For this reason, we decided to generate synthetic data using Gaussian noise, for which a PSD (or ASD, or sqrtSX, they are all essentially equivalent in this context) needs to be specified.</p>\n<blockquote>\n  <p>It would us sidechannel which is not present in real GW search data.</p>\n</blockquote>\n<p>This is not a problem in CW searches. For \"clean\" (relatively quiet) bands we are able to estimate the PSD of the noise with quite good accuracy, and that is sufficient for us to actually perform our analysis. </p>\n<hr>\n<h5>Regarding Those -1 Labels (<a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363734\" target=\"_blank\">PaulG</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363734\" target=\"_blank\">PaulG</a> <strong>Found an easter egg</strong> left by the competition Hosts! </li>\n</ul>\n<p>The dataset description for the competition states:</p>\n<pre><code>\"Please note the presence of a small number of files labeled -1.\nPhysicists are currently unable to determine the status of these files.\"\n</code></pre>\n<p>As it turns out, these are easter eggs!.<br>\nIf you plot their signal magnitudes as 2D arrays, you'll see:</p>\n<p><strong>50f09e37e</strong><br>\nA cartoon of two black holes about to collide.</p>\n<p><img src=\"https://i.ibb.co/NxnJL9M/50f09e37e.jpg\" alt=\"\"></p>\n<p><strong>62b0dd011</strong><br>\nImage that was broadcast into the universe in 1974, as part of the SETI program.<br>\n<img src=\"https://i.ibb.co/347VcjZ/62b0dd011.jpg\" alt=\"\"></p>\n<p><strong>b7666b451</strong><br>\nThe 2017 physics Nobel Prize winners (for the discovery of gravitational waves):  Rainer Weiss, Barry Barish and Kip Thorne.<br>\n<img src=\"https://i.ibb.co/S0NnW3Z/b7666b451.jpg\" alt=\"\"></p>\n<hr>\n<h5>High validation accuracy, low LB score (<a href=\"https://www.kaggle.com/simonebvr\" target=\"_blank\">Simone Bavera</a>)</h5>\n<p>An <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363810\" target=\"_blank\">interesting</a> discussion about methods for reducing the difference between the LB and the CV score. </p>\n<ul>\n<li><strong>Important</strong> Answer by <a href=\"https://www.kaggle.com/chenlin1999\" target=\"_blank\">Chen Lin</a></li>\n</ul>\n<pre><code>This competition is very different from the last one. The last contest was the signal of two black holes or neutron stars merging, which is very strong, and implementing a CNN would allow for some level of fit. But continuous gravitational waves are very weak, and even after scientists in the field have implemented various search methods and deep learning tools they have not found any real signal. If you read some of the papers you will see that even the most advanced NNs that have been tried have not reached traditional methods of search in terms of accuracy, and are usually very unsatisfactory after a SNR of less than 10.\n\nI think the huge gap between CV and LB is because 1) the training set is so small with only 600 data, there is so little variation in between that it may be difficult for the network to learn convolution methods in the face of low SNR samples, and 2) I haven't carefully checked the test set, but I think the composition of the test set must be very demanding and there must be a large number of low SNR samples.\n\nIn order to be able to bring CV and LB closer together, I think the most effective way to do this would be to construct a training set that has a large number of samples with a wide range of SNR values, rather than using 600 training samples and then trying various ML tricks.\n</code></pre>\n<hr>\n<h5>Possible instrumental artifacts in training/testing dataset (<a href=\"https://www.kaggle.com/alexz0\" target=\"_blank\">Alex Z</a>)</h5>\n<p>An <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/364854\" target=\"_blank\">interesting</a> and informative discussion about the posibility of additional instrumental artifacts (persistent line-like features) to the noise alongside the simulated signals in testing dataset.</p>\n<ul>\n<li><strong>Important</strong> Answer by <a href=\"https://www.kaggle.com/kdmitrie\" target=\"_blank\">Konstantin Dmitriev</a>:</li>\n</ul>\n<pre><code>Yes, the artifacts like persistent lines do exist alongside with the CW signal and background noise. I've done some statistical testing of the whole dataset (test+train) and found them. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Fdc5855c34dcd37e930419ec3f9039cf2%2F__results___32_1.png?generation=1668240677739409&amp;alt=media)\n</code></pre>\n<hr>\n<h5>Gap of CV and LB (<a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">Chenglu</a>)</h5>\n<p>Another <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/364169\" target=\"_blank\">discussion</a> about creating a robust CV strategy.</p>\n<p><strong>Important Note</strong> by <a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">Chenglu</a></p>\n<pre><code>With the generated dataset, I can reduce the gap to 0.1, CV ~0.77 and LB ~0.68, I think this competition is a matter of shrinking the gap, more precisely, it's generating a dataset that matches the distribution of test set.\n</code></pre>\n<hr>\n<h5>Generating samples that match the train / test distribution (<a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">Chenglu</a>)</h5>\n<p>Thank you for sharing this!!<br>\n<a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">Chenglu</a> is sharing with us a common bug that caused some problems (and you need to also be aware of it):</p>\n<p>The bug: Normalizing the data this way:</p>\n<pre><code>norm_data = data - data.() / (data.() - data.())\n</code></pre>\n<p>This is easily influenced by the value of max and min.<br>\nThe fix:</p>\n<pre><code>norm_data = data * 1e22\n</code></pre>\n<p>The fixed classifier only achieve 0.66 on classifying generating v.s. testing/training dataset.</p>\n<hr>\n<h5>What is the difference between the G2Net this year and last year? (<a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/366451\" target=\"_blank\">ForcewithMe</a>)</h5>\n<p><strong>The post lists the differences between the two competitions:</strong></p>\n<ul>\n<li><p>The raw data last year are 1d waves, whose x-axis is time, and y-axis is the amplitude of each timestamp. With CQT/CWT, the 1d wave can be converted into a 2d image(spectrum), which x-axis is time, y-axis is frequency and the intensity per unit pixel is the amplitude of a specific frequency within a specific timestamp.</p></li>\n<li><p>The raw data are 2d images(spectrum), which x-axis is time, y-axis is frequency and the intensity per unit pixel is the amplitude of a specific frequency within a specific timestamp. (Which is the same as the data after CQT last year in format? I am not sure…)</p></li>\n</ul>\n<p><strong>SNR</strong></p>\n<ul>\n<li><p>The SNR in the competition this year is much lower, which means that the signal this year is much weaker, resulting in the auc being much lower.</p></li>\n<li><p><strong>Important</strong> Addition by <a href=\"https://www.kaggle.com/alexz0\" target=\"_blank\">Alex Z</a>: Physics-wise, last year competition dealt with the waves from compact binary objects, like close orbiting black holes or white dwarfs. This year we face fast spinning asymmetric neutron stars. Last year set-up provides much more explosive waveforms, and much stronger. Continuous GW we are hunting here were never detected so far, so we have a chance to help a real scientific discovery.</p></li>\n</ul>\n<hr>\n<h5>Reverse engineering 20% of test samples with external data. (<a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">Vladimir Slaykovskiy</a>)</h5>\n<blockquote>\n  <p><strong>TL;DR:</strong> Vladimir Dropped a BOMB. </p>\n</blockquote>\n<p><strong>It might be possible to reverse engineer the injected CW signal by using raw Ligo&amp;Virgo data.</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/vslaykovsky/g2net-winning-strategy-with-external-data\" target=\"_blank\">code</a></li>\n</ul>\n<p><strong>In Short</strong></p>\n<ul>\n<li>Get raw detector data at <a href=\"https://www.gw-openscience.org/archive/O3a_4KHZ_R1/\" target=\"_blank\">https://www.gw-openscience.org/archive/O3a_4KHZ_R1/</a></li>\n<li>Generate SFTs using instructions at <a href=\"https://youtu.be/A9iWRcmG0Rs?t=2199\" target=\"_blank\">https://youtu.be/A9iWRcmG0Rs?t=2199</a> (you'll likely need to google for more details on this step). Use GPS timestamps from test set to produce SFTs.</li>\n<li>Read generated SFTs using pyfstat.utils.sft.get_sft_as_arrays</li>\n<li>Match generated SFTs with real samples from test set.</li>\n<li>Find the difference between SFTs and real samples. difference &gt; EPS -&gt; label==1; difference &lt;= EPS -&gt; label==0.</li>\n</ul>\n<blockquote>\n  <p>Wow.. well done on this!</p>\n</blockquote>\n<hr>\n<h5>Score improvement G wave detection 20% data (<a href=\"https://www.kaggle.com/igorlitvin\" target=\"_blank\">Igor Litvin</a>)</h5>\n<ul>\n<li><p><a href=\"https://www.kaggle.com/igorlitvin\" target=\"_blank\">Igor Litvin</a> proposed a method with some potential for improving your NN training.</p></li>\n<li><p>He figered out that aproximately 20% of the gravitation wave spectra are distorted.</p></li>\n<li><p>And think the distortion was made artificially <strong>because of it is exatly the same distortion for H and L shoulders of the interferometer.</strong></p></li>\n<li><p>It is possible that they multiplied by exactly the same function to H and exactly the same (but different from H) to L shoulder.</p></li>\n<li><p><a href=\"https://www.kaggle.com/code/igorlitvin/l-and-h-distortion-of-g-wave-help\" target=\"_blank\">Code here</a></p></li>\n</ul>\n<hr>\n<h5>Some Insights About Validation Strategy (<a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">Martin Kovacevic Buvinic</a>)</h5>\n<ul>\n<li><p>Great topic by <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">Martin Kovacevic Buvinic</a> about important insights about the data.</p></li>\n<li><p><strong>CV:</strong> The test data and train data are different: cv gap between validation and test is always large (~0.10). Also low correlation. - BUT! somehow a 20 folds cv model turned out to be much better. </p></li>\n<li><p><strong>Guess:</strong> Using 20 folds is only 5% of the data as validation for each fold, meaning we are only validating with 30 observations and this lead to overfitting, therefore this validation is not good, increasing the number of folds will just lead to overfitting.</p></li>\n<li><p><strong>Blending:</strong> Blending in most of the casses helps.</p></li>\n<li><p><strong>Data Generation:</strong> Adding more data usually improves the generalization of the model, if we check lb it actually improves so adding more data is helpful. The data simulated is the training data, and the best guess is that the key of the competition is to generate data that is similar to the test data.</p></li>\n<li><p><strong>The Main Insight:</strong> Using training data as validation is not optimal, because it is small and different compared to the test set. A good option would be to generate our own data that is similar to the test set and do traininig and validation with that data, also add the training data.</p></li>\n</ul>\n<hr>\n<h5>Merging generated signal into noise (<a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">Vladimir Slaykovskiy</a>)</h5>\n<p><a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">Vladimir Slaykovskiy</a> asks about the easies way to combine generated data into the a single spectrogram.</p>\n<ul>\n<li>For example is it correct to just weight-average them?<br>\n<code>sample = 0.9 * noise + 0.1 + signal</code></li>\n</ul>\n<p><strong>Answer (By The Hosts):</strong></p>\n<pre><code>If noise  signal are Fourier amplitudes (i.e. the  numbers you get  get_stft_as_arrays)  the generated SFTs have the same timestamps  frequency, then yes.\n\nNote that you can literally do whatever you want to that data. For example,  you wanted to create an instrumental artifact  a circular shape, you could create such a shape  a numpy array  the appropriate array shape  directly added to your data.\n\n`For example  it correct to just weight-average them?`\n\nThe standard way  which we do this  by generating signals the amplitude of which  already a certain fraction of your noise</code></pre>",
  "messages": [
    {
      "id": 2074716,
      "postDate": "2022-12-24T14:33:19.357Z",
      "content": "<h3>What do we know so far?</h3>\n<h5>10 Days to go!</h5>\n<hr>\n<p>Save yourself some time and get caught up on the latest updates and important discussions that have taken place prior to the competition deadline.<br>\nThe competition ends in 10 days and there is still time to improve your solution. Here are some of the important topics that have been discussed in the Kaggle forums.</p>\n<hr>\n<h5>Background</h5>\n<ul>\n<li><p>The data is H-U-G-E and comes in hdf5 format, <a href=\"https://www.kaggle.com/chazzer\" target=\"_blank\">chazzer</a> had written a <a href=\"https://www.kaggle.com/code/chazzer/how-to-read-the-hdf5-files/\" target=\"_blank\">notebook on how to read the hdf5 files</a></p></li>\n<li><p>Shortly after <a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">JohnM</a> pushed this one step further and created <a href=\"https://www.kaggle.com/datasets/jpmiller/simplified-dataset\" target=\"_blank\">HDF files with indexed dataframes for the training set</a>. The index on each frame is the frequency array from the original file, and the columns are the timestamps.</p></li>\n</ul>\n<p><strong>Reading Example</strong></p>\n<pre><code>df_h = pd.read_hdf(, key = )\n</code></pre>\n<ul>\n<li>Great <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/357682\" target=\"_blank\">list</a> of learning resources had been published by <a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">Laura Fink</a><ul>\n<li><a href=\"https://www.gw-openscience.org/tutorials/\" target=\"_blank\">Gravitational Wave Open Science Center - Tutorials</a></li>\n<li><a href=\"https://gwpy.github.io/docs/stable/index.html\" target=\"_blank\">GWPy</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=tQ_teIUb3tE\" target=\"_blank\">Video about the Michelson Interferometer</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=FlDtXIBrAYE&amp;t=1s\" target=\"_blank\">Video about gravitational waves</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=spUNpyF58BY&amp;t=1s\" target=\"_blank\">Great tutorial video about Fourier transforms</a></li>\n<li><a href=\"https://www.youtube.com/watch?v=s2K1JfNR7Sc\" target=\"_blank\">Video about Denoising data with Fast Fourier Transforms</a></li></ul></li>\n</ul>\n<hr>\n<h5>Best CV-LB Results (<a href=\"https://www.kaggle.com/e0xextazy\" target=\"_blank\">Mark Baushenko</a>)</h5>\n<ul>\n<li><strong>Important</strong> Comments by <a href=\"drhabib\" target=\"_blank\">DrHB</a><ul>\n<li>\"Without <code>augs</code> the gap (CV-LB) is bigger\"</li>\n<li>\"After running some experiments .. I think the biggest challenge of this competition will be to build a proper validation set\"    </li></ul></li>\n</ul>\n<hr>\n<h5>Understanding the Data (<a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">Ravi Shah</a>)</h5>\n<p>An <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/358444\" target=\"_blank\">amazing intro post</a> by <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">Ravi Shah</a> going through all important details about the data of this competition. </p>\n<p><strong>In short (check the original for more info)</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F1893bf6ae0ffb282512eec0fdce5a662%2Fstruc.PNG?generation=1665177635295503&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>L1 - data from the LIGO Livingston interferometer</li>\n<li>H1 - data from the LIGO Hanford interferometer</li>\n<li>STFs - Short-time Fourier Transforms (SFTs) - shape (360, n)</li>\n<li>timestamps - the timestamps that the STFs correspond to - shape (n,)</li>\n<li>frequency Hz - the range frequencies measured by the detectors - shape (360,)</li>\n</ul>\n<p><strong>Competition Goal</strong></p>\n<p>The goal of this competition is to create a model that can detect continuous gravitational-wave signals. These are weak yet long-lasting signals emitted by rapidly-spinning neutron stars within noisy data. </p>\n<blockquote>\n  <ul>\n  <li><strong>Target is 1 if signal is present else 0</strong></li>\n  </ul>\n</blockquote>\n<p><strong>LIGO</strong> (where the data comes from) is a detector in Livingston, Louisiana and Hanford, Washington. The detector has a L shape with equal-length arms connected to a corner station. When a gravitational wave passes by, one leg of the detector is shortened while the other is lengthened. </p>\n<p><strong>The interference makes a shift that we can analyze.</strong></p>\n<p><strong>We Get A Continuous Gravitational Waves</strong></p>\n<p>A continuous gravitational-wave signal from a Galactic neutron star will look almost perfectly constant in both frequency and amplitude.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F6b476a7c1f076bdf7c4f7234d37ace3b%2Fperfect%20cont%20wave.PNG?generation=1665181150773656&amp;alt=media\" alt=\"\"></p>\n<p><strong>But over long durations: The frequency of the signal slowly changes, because:</strong></p>\n<ul>\n<li>As the neutron star emits gravitational and electromagnetic waves, it loses energy which causes it to rotate more slowly</li>\n<li>The detector here on Earth is moving with respect to the neutron star. This changes the frequency of the gravitational waves observed in the detector.</li>\n</ul>\n<p><strong>Short-time Fourier Transforms (SFTs)</strong></p>\n<ul>\n<li>Short-time Fourier Transforms (SFTs) can be used as a way of quantifying the change of a nonstationary signal’s frequency and phase content over time.</li>\n<li>Very common type of signal preprocessing technique.</li>\n<li>This what we are given to work with (hdf5 files).</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fa8a7caafee2261f27fde68045a824202%2Fstft.PNG?generation=1665180281946268&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h5>Timestamps and Frequencies Insights (<a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">Mark Wijkhuizen</a>)</h5>\n<p><strong>There seems to be 2 differences of our data in comparison to conventional SFT's</strong></p>\n<ul>\n<li><p><strong>The time delta between each timestamp is inconsistent:</strong> Time between measurement in the SFT's can vary greatly.</p></li>\n<li><p>The frequency range is consistently ~0.2Hz, <strong>BUT the min/max frequency ranges from around ~50to ~498Hz.</strong></p></li>\n<li><p><strong>We should not interpret all spectrograms as continuous sounds:</strong> When taking a look at the timestamps they start at 2009-03-27/28 and end on 2009-07-25/26/27. The recordings are thus gathered during 4 months with often 30 minutes between the recordings. </p></li>\n</ul>\n<hr>\n<h5>Minimum float32 is 1e-38 and data**2 is 1e-44 (<a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">🐢 Jun Koda</a>)</h5>\n<p>Important <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361312\" target=\"_blank\">post</a> by <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">🐢 Jun Koda</a> sharing with us that since the smallest positive number float32 (~ 1e-38) but on the data it is ~1e-22. </p>\n<p><strong>Once we square the data, they go beyond this limit.</strong></p>\n<blockquote>\n  <p>np.float32(1.23456e-44) =&gt; 1.3e-44.</p>\n</blockquote>\n<ul>\n<li>float32 has about 8 significant digits and 1.23456e-44 will be expressed as 0.0000012e-38 with float32, only 2 digits remaining.</li>\n<li>Suggestion: that we multiply the data by 1e21 or 1e22.</li>\n</ul>\n<hr>\n<h5>PyFstat Signal Parameters (<a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">Mark Wijkhuizen</a>)</h5>\n<blockquote>\n  <p>Another great <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361562\" target=\"_blank\">post</a> by (<a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">Mark Wijkhuizen</a>)</p>\n</blockquote>\n<p>PyFstat has many parameters but it's documentation however is not very helpful for explaining them.</p>\n<p><strong><a href=\"https://www.kaggle.com/darkbitur\" target=\"_blank\">Victor Gonzalez</a></strong> Explained with an incredible comment everything about all the paramaeters. </p>\n<ul>\n<li>You can use a helper function to inject some parameters based on priors:</li>\n</ul>\n<pre><code>{**pyfstat.injection_parameters.isotropic_amplitude_priors}\n{: {: {: -, : }}, : {: {: -, : }}, : {: {: , : }}}\n</code></pre>\n<p><strong>More Parameters</strong></p>\n<ul>\n<li><strong>cosi:</strong> Cosine of the angle between the source and us. Range: [-1, 1]</li>\n<li><strong>psi and phi:</strong> polarization angle and phase. Ranges: [-pi/2, pi/2] and [0, 2*pi]</li>\n<li><strong>Alpha:</strong> Right ascension of the source's position on the sky. Range: [0, 2*pi]</li>\n<li><strong>Delta:</strong> Declination of the source's position on the sky. Range: [-pi/2, pi/2]</li>\n<li><strong>Band:</strong> This is just the width of the frequency \"slice\" to be taken. It's 0.2 hz for all files in the train and test datasets, which translates to 360 \"frequency lines\".</li>\n<li><strong>F0:</strong> This may be tricky. The generated noise will be centered around this frequency. If a signal is created, it will be also a signal with this frequency (at the source; when it arrives to the detector, doppler and other effects will shift the frequency), so it will appear more or less on the center of the data. Range used on the datasets: [50, 400] Hz.</li>\n<li><strong>F1:</strong> This is technically the first derivative of the frequency (how it changs over time). Physically, the sources (neutron stars) lose energy over time because they emit the GWs, so this should be a negative number. A reasonable range could be [-1e-9, 0] Hz but that's up to us.</li>\n<li><strong>F2:</strong> This is guess it's the second derivative, which maybe is 0 for the purposes of this competition.</li>\n<li><strong>h0:</strong> The amplitude of the signal. In the tutorials, the approach is to define it as a ratio of the sqrtSX. What makes the competition (and the real-world problem of finding continuous GWs) is that expected h0 will be between 1 and 2 orders of magnitude smaller than the amplitude of the noise. In my opinion, part of the challenge is to create your own signals with a distribution of h0 similar to the one used for the test set.</li>\n<li><strong>psi</strong> is actually defined within [-pi /4 , pi / 4] (and it's periodic, so anything beyond that folds back onto this range).</li>\n<li><strong>F2</strong> is another term describing the intrinsic evolution of a gravitational wave. As <a href=\"https://www.kaggle.com/darkbitur\" target=\"_blank\">@darkbitur</a> correctly guesses, it's set to 0 in this competition as we only focus on frequency and spindown.</li>\n<li><strong>tp, asini, period</strong> are parameters used to describe neutron stars in binary systems, and they affect the frequency evolution of the signal. You can safely ignore those in this competition (but do feel free to play around with them if you wish to!).</li>\n</ul>\n<hr>\n<h5>Can we assume the same signal trajectories between Hanford &amp; Livingston? (<a href=\"https://www.kaggle.com/ikarosilva\" target=\"_blank\">IkaroSilva</a>)</h5>\n<p>Question <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361655\" target=\"_blank\">asked</a> by <a href=\"https://www.kaggle.com/ikarosilva\" target=\"_blank\">IkaroSilva</a></p>\n<ul>\n<li>Can we assume the same signal trajectories between Hanford &amp; Livingston?</li>\n</ul>\n<p><strong>Answer (By The Hosts):</strong></p>\n<ul>\n<li><p>The frequency of a signal as seen by a detector a certain time t (what you call it S(t) in your post) does generally depend on the detector, the reason being that different detectors move \"differently\" around the Sun as they are in slightly different positions on Earth.</p></li>\n<li><p>This distinction may not be that noticeable for some cases if we use the two Advanced LIGO detectors, since they are relatively close to each other, but becomes quite apparent once you include other detectors such as Virgo, which is located in Italy.</p></li>\n</ul>\n<hr>\n<h5>What is your batch size and LB? (<a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">dragon zhang</a>)</h5>\n<p>A [good] question about the relation beween the batch-size and generalization performance by <a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">dragon zhang</a>.</p>\n<p><strong>Short Answer:</strong></p>\n<pre><code>Friends dont let friends use minibatches larger than 32.\n[Yann LeCun](https://twitter.com/ylecun/status/989610208497360896).\n</code></pre>\n<hr>\n<h5>Recap of the Top Solutions from the Previous G2Net Competition (<a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363280\" target=\"_blank\">Sinan Calisir</a>)</h5>\n<p>An amazing solutions <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363280\" target=\"_blank\">summary</a> by <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363280\" target=\"_blank\">Sinan Calisir</a> of all winning solutions of the previous competition.</p>\n<hr>\n<h5>Weird property of std(spectrogram) in train data ([Josef Slavicek</h5>\n<p>](<a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363460)\" target=\"_blank\">https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363460)</a>)</p>\n<p><strong>Wierd Observation:</strong> When computing std of spectrogram in complex64, I'm getting almost allways the same value of 1.4973568266036865e-22. If I do the same computation in complex128, the values starts do differ - and they are partly useful as indicator of target label:</p>\n<p><strong>Answer (From the Hosts):</strong></p>\n<p><strong>TL;DR:</strong> This is not a data leak, we understand the reason this is happening and it is expected. Nothing to worry about.</p>\n<blockquote>\n  <p>and they are partly useful as indicator of target label</p>\n</blockquote>\n<p>Yes, this is expected: stdev is computed from the data, and very strong signals will deviate the value away from the \"noise-only\" value. This should already tell you how \"strong\" signals in the test set are, by the way :) .</p>\n<blockquote>\n  <p>Around 80% of test data seems to be affected by this</p>\n</blockquote>\n<p>This is also well understood. We wanted to give as much real data as possible in this challenge, but we didn't have a reliable way of generating more than what the LIGO detectors produced during the run, and we did not want to run into issues about having too much correlation among different samples. For this reason, we decided to generate synthetic data using Gaussian noise, for which a PSD (or ASD, or sqrtSX, they are all essentially equivalent in this context) needs to be specified.</p>\n<blockquote>\n  <p>It would us sidechannel which is not present in real GW search data.</p>\n</blockquote>\n<p>This is not a problem in CW searches. For \"clean\" (relatively quiet) bands we are able to estimate the PSD of the noise with quite good accuracy, and that is sufficient for us to actually perform our analysis. </p>\n<hr>\n<h5>Regarding Those -1 Labels (<a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363734\" target=\"_blank\">PaulG</a>)</h5>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363734\" target=\"_blank\">PaulG</a> <strong>Found an easter egg</strong> left by the competition Hosts! </li>\n</ul>\n<p>The dataset description for the competition states:</p>\n<pre><code>\"Please note the presence of a small number of files labeled -1.\nPhysicists are currently unable to determine the status of these files.\"\n</code></pre>\n<p>As it turns out, these are easter eggs!.<br>\nIf you plot their signal magnitudes as 2D arrays, you'll see:</p>\n<p><strong>50f09e37e</strong><br>\nA cartoon of two black holes about to collide.</p>\n<p><img src=\"https://i.ibb.co/NxnJL9M/50f09e37e.jpg\" alt=\"\"></p>\n<p><strong>62b0dd011</strong><br>\nImage that was broadcast into the universe in 1974, as part of the SETI program.<br>\n<img src=\"https://i.ibb.co/347VcjZ/62b0dd011.jpg\" alt=\"\"></p>\n<p><strong>b7666b451</strong><br>\nThe 2017 physics Nobel Prize winners (for the discovery of gravitational waves):  Rainer Weiss, Barry Barish and Kip Thorne.<br>\n<img src=\"https://i.ibb.co/S0NnW3Z/b7666b451.jpg\" alt=\"\"></p>\n<hr>\n<h5>High validation accuracy, low LB score (<a href=\"https://www.kaggle.com/simonebvr\" target=\"_blank\">Simone Bavera</a>)</h5>\n<p>An <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363810\" target=\"_blank\">interesting</a> discussion about methods for reducing the difference between the LB and the CV score. </p>\n<ul>\n<li><strong>Important</strong> Answer by <a href=\"https://www.kaggle.com/chenlin1999\" target=\"_blank\">Chen Lin</a></li>\n</ul>\n<pre><code>This competition is very different from the last one. The last contest was the signal of two black holes or neutron stars merging, which is very strong, and implementing a CNN would allow for some level of fit. But continuous gravitational waves are very weak, and even after scientists in the field have implemented various search methods and deep learning tools they have not found any real signal. If you read some of the papers you will see that even the most advanced NNs that have been tried have not reached traditional methods of search in terms of accuracy, and are usually very unsatisfactory after a SNR of less than 10.\n\nI think the huge gap between CV and LB is because 1) the training set is so small with only 600 data, there is so little variation in between that it may be difficult for the network to learn convolution methods in the face of low SNR samples, and 2) I haven't carefully checked the test set, but I think the composition of the test set must be very demanding and there must be a large number of low SNR samples.\n\nIn order to be able to bring CV and LB closer together, I think the most effective way to do this would be to construct a training set that has a large number of samples with a wide range of SNR values, rather than using 600 training samples and then trying various ML tricks.\n</code></pre>\n<hr>\n<h5>Possible instrumental artifacts in training/testing dataset (<a href=\"https://www.kaggle.com/alexz0\" target=\"_blank\">Alex Z</a>)</h5>\n<p>An <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/364854\" target=\"_blank\">interesting</a> and informative discussion about the posibility of additional instrumental artifacts (persistent line-like features) to the noise alongside the simulated signals in testing dataset.</p>\n<ul>\n<li><strong>Important</strong> Answer by <a href=\"https://www.kaggle.com/kdmitrie\" target=\"_blank\">Konstantin Dmitriev</a>:</li>\n</ul>\n<pre><code>Yes, the artifacts like persistent lines do exist alongside with the CW signal and background noise. I've done some statistical testing of the whole dataset (test+train) and found them. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Fdc5855c34dcd37e930419ec3f9039cf2%2F__results___32_1.png?generation=1668240677739409&amp;alt=media)\n</code></pre>\n<hr>\n<h5>Gap of CV and LB (<a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">Chenglu</a>)</h5>\n<p>Another <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/364169\" target=\"_blank\">discussion</a> about creating a robust CV strategy.</p>\n<p><strong>Important Note</strong> by <a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">Chenglu</a></p>\n<pre><code>With the generated dataset, I can reduce the gap to 0.1, CV ~0.77 and LB ~0.68, I think this competition is a matter of shrinking the gap, more precisely, it's generating a dataset that matches the distribution of test set.\n</code></pre>\n<hr>\n<h5>Generating samples that match the train / test distribution (<a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">Chenglu</a>)</h5>\n<p>Thank you for sharing this!!<br>\n<a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">Chenglu</a> is sharing with us a common bug that caused some problems (and you need to also be aware of it):</p>\n<p>The bug: Normalizing the data this way:</p>\n<pre><code>norm_data = data - data.() / (data.() - data.())\n</code></pre>\n<p>This is easily influenced by the value of max and min.<br>\nThe fix:</p>\n<pre><code>norm_data = data * 1e22\n</code></pre>\n<p>The fixed classifier only achieve 0.66 on classifying generating v.s. testing/training dataset.</p>\n<hr>\n<h5>What is the difference between the G2Net this year and last year? (<a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/366451\" target=\"_blank\">ForcewithMe</a>)</h5>\n<p><strong>The post lists the differences between the two competitions:</strong></p>\n<ul>\n<li><p>The raw data last year are 1d waves, whose x-axis is time, and y-axis is the amplitude of each timestamp. With CQT/CWT, the 1d wave can be converted into a 2d image(spectrum), which x-axis is time, y-axis is frequency and the intensity per unit pixel is the amplitude of a specific frequency within a specific timestamp.</p></li>\n<li><p>The raw data are 2d images(spectrum), which x-axis is time, y-axis is frequency and the intensity per unit pixel is the amplitude of a specific frequency within a specific timestamp. (Which is the same as the data after CQT last year in format? I am not sure…)</p></li>\n</ul>\n<p><strong>SNR</strong></p>\n<ul>\n<li><p>The SNR in the competition this year is much lower, which means that the signal this year is much weaker, resulting in the auc being much lower.</p></li>\n<li><p><strong>Important</strong> Addition by <a href=\"https://www.kaggle.com/alexz0\" target=\"_blank\">Alex Z</a>: Physics-wise, last year competition dealt with the waves from compact binary objects, like close orbiting black holes or white dwarfs. This year we face fast spinning asymmetric neutron stars. Last year set-up provides much more explosive waveforms, and much stronger. Continuous GW we are hunting here were never detected so far, so we have a chance to help a real scientific discovery.</p></li>\n</ul>\n<hr>\n<h5>Reverse engineering 20% of test samples with external data. (<a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">Vladimir Slaykovskiy</a>)</h5>\n<blockquote>\n  <p><strong>TL;DR:</strong> Vladimir Dropped a BOMB. </p>\n</blockquote>\n<p><strong>It might be possible to reverse engineer the injected CW signal by using raw Ligo&amp;Virgo data.</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/vslaykovsky/g2net-winning-strategy-with-external-data\" target=\"_blank\">code</a></li>\n</ul>\n<p><strong>In Short</strong></p>\n<ul>\n<li>Get raw detector data at <a href=\"https://www.gw-openscience.org/archive/O3a_4KHZ_R1/\" target=\"_blank\">https://www.gw-openscience.org/archive/O3a_4KHZ_R1/</a></li>\n<li>Generate SFTs using instructions at <a href=\"https://youtu.be/A9iWRcmG0Rs?t=2199\" target=\"_blank\">https://youtu.be/A9iWRcmG0Rs?t=2199</a> (you'll likely need to google for more details on this step). Use GPS timestamps from test set to produce SFTs.</li>\n<li>Read generated SFTs using pyfstat.utils.sft.get_sft_as_arrays</li>\n<li>Match generated SFTs with real samples from test set.</li>\n<li>Find the difference between SFTs and real samples. difference &gt; EPS -&gt; label==1; difference &lt;= EPS -&gt; label==0.</li>\n</ul>\n<blockquote>\n  <p>Wow.. well done on this!</p>\n</blockquote>\n<hr>\n<h5>Score improvement G wave detection 20% data (<a href=\"https://www.kaggle.com/igorlitvin\" target=\"_blank\">Igor Litvin</a>)</h5>\n<ul>\n<li><p><a href=\"https://www.kaggle.com/igorlitvin\" target=\"_blank\">Igor Litvin</a> proposed a method with some potential for improving your NN training.</p></li>\n<li><p>He figered out that aproximately 20% of the gravitation wave spectra are distorted.</p></li>\n<li><p>And think the distortion was made artificially <strong>because of it is exatly the same distortion for H and L shoulders of the interferometer.</strong></p></li>\n<li><p>It is possible that they multiplied by exactly the same function to H and exactly the same (but different from H) to L shoulder.</p></li>\n<li><p><a href=\"https://www.kaggle.com/code/igorlitvin/l-and-h-distortion-of-g-wave-help\" target=\"_blank\">Code here</a></p></li>\n</ul>\n<hr>\n<h5>Some Insights About Validation Strategy (<a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">Martin Kovacevic Buvinic</a>)</h5>\n<ul>\n<li><p>Great topic by <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">Martin Kovacevic Buvinic</a> about important insights about the data.</p></li>\n<li><p><strong>CV:</strong> The test data and train data are different: cv gap between validation and test is always large (~0.10). Also low correlation. - BUT! somehow a 20 folds cv model turned out to be much better. </p></li>\n<li><p><strong>Guess:</strong> Using 20 folds is only 5% of the data as validation for each fold, meaning we are only validating with 30 observations and this lead to overfitting, therefore this validation is not good, increasing the number of folds will just lead to overfitting.</p></li>\n<li><p><strong>Blending:</strong> Blending in most of the casses helps.</p></li>\n<li><p><strong>Data Generation:</strong> Adding more data usually improves the generalization of the model, if we check lb it actually improves so adding more data is helpful. The data simulated is the training data, and the best guess is that the key of the competition is to generate data that is similar to the test data.</p></li>\n<li><p><strong>The Main Insight:</strong> Using training data as validation is not optimal, because it is small and different compared to the test set. A good option would be to generate our own data that is similar to the test set and do traininig and validation with that data, also add the training data.</p></li>\n</ul>\n<hr>\n<h5>Merging generated signal into noise (<a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">Vladimir Slaykovskiy</a>)</h5>\n<p><a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">Vladimir Slaykovskiy</a> asks about the easies way to combine generated data into the a single spectrogram.</p>\n<ul>\n<li>For example is it correct to just weight-average them?<br>\n<code>sample = 0.9 * noise + 0.1 + signal</code></li>\n</ul>\n<p><strong>Answer (By The Hosts):</strong></p>\n<pre><code>If noise  signal are Fourier amplitudes (i.e. the  numbers you get  get_stft_as_arrays)  the generated SFTs have the same timestamps  frequency, then yes.\n\nNote that you can literally do whatever you want to that data. For example,  you wanted to create an instrumental artifact  a circular shape, you could create such a shape  a numpy array  the appropriate array shape  directly added to your data.\n\n`For example  it correct to just weight-average them?`\n\nThe standard way  which we do this  by generating signals the amplitude of which  already a certain fraction of your noise</code></pre>",
      "rawMarkdown": "### What do we know so far? \n##### 10 Days to go!\n_____\nSave yourself some time and get caught up on the latest updates and important discussions that have taken place prior to the competition deadline.\nThe competition ends in 10 days and there is still time to improve your solution. Here are some of the important topics that have been discussed in the Kaggle forums.\n_____\n\n\n##### Background\n\n- The data is H-U-G-E and comes in hdf5 format, [chazzer](https://www.kaggle.com/chazzer) had written a [notebook on how to read the hdf5 files](https://www.kaggle.com/code/chazzer/how-to-read-the-hdf5-files/)\n\n- Shortly after [JohnM](https://www.kaggle.com/jpmiller) pushed this one step further and created [HDF files with indexed dataframes for the training set](https://www.kaggle.com/datasets/jpmiller/simplified-dataset). The index on each frame is the frequency array from the original file, and the columns are the timestamps.\n\n**Reading Example**\n\n```python\ndf_h = pd.read_hdf(\"../input/simplified-dataset/001121a05.h5\", key = 'h')\n```\n\n- Great [list](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/357682) of learning resources had been published by [Laura Fink](https://www.kaggle.com/allunia)\n    - [Gravitational Wave Open Science Center - Tutorials](https://www.gw-openscience.org/tutorials/)\n    - [GWPy](https://gwpy.github.io/docs/stable/index.html)\n    - [Video about the Michelson Interferometer](https://www.youtube.com/watch?v=tQ_teIUb3tE)\n    - [Video about gravitational waves](https://www.youtube.com/watch?v=FlDtXIBrAYE&t=1s)\n    - [Great tutorial video about Fourier transforms](https://www.youtube.com/watch?v=spUNpyF58BY&t=1s)\n    - [Video about Denoising data with Fast Fourier Transforms](https://www.youtube.com/watch?v=s2K1JfNR7Sc)\n \n_____\n\n##### Best CV-LB Results ([Mark Baushenko](https://www.kaggle.com/e0xextazy))\n\n- **Important** Comments by [DrHB](drhabib)\n    - \"Without `augs` the gap (CV-LB) is bigger\"\n    - \"After running some experiments .. I think the biggest challenge of this competition will be to build a proper validation set\"    \n    \n_____\n\n\n##### Understanding the Data ([Ravi Shah](https://www.kaggle.com/ravishah1))\n\nAn [amazing intro post](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/358444) by [Ravi Shah](https://www.kaggle.com/ravishah1) going through all important details about the data of this competition. \n\n**In short (check the original for more info)**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F1893bf6ae0ffb282512eec0fdce5a662%2Fstruc.PNG?generation=1665177635295503&alt=media)\n\n- L1 - data from the LIGO Livingston interferometer\n- H1 - data from the LIGO Hanford interferometer\n- STFs - Short-time Fourier Transforms (SFTs) - shape (360, n)\n- timestamps - the timestamps that the STFs correspond to - shape (n,)\n- frequency Hz - the range frequencies measured by the detectors - shape (360,)\n\n**Competition Goal**\n\nThe goal of this competition is to create a model that can detect continuous gravitational-wave signals. These are weak yet long-lasting signals emitted by rapidly-spinning neutron stars within noisy data. \n\n> - **Target is 1 if signal is present else 0**\n\n\n**LIGO** (where the data comes from) is a detector in Livingston, Louisiana and Hanford, Washington. The detector has a L shape with equal-length arms connected to a corner station. When a gravitational wave passes by, one leg of the detector is shortened while the other is lengthened. \n\n**The interference makes a shift that we can analyze.**\n\n\n**We Get A Continuous Gravitational Waves**\n\nA continuous gravitational-wave signal from a Galactic neutron star will look almost perfectly constant in both frequency and amplitude.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F6b476a7c1f076bdf7c4f7234d37ace3b%2Fperfect%20cont%20wave.PNG?generation=1665181150773656&alt=media)\n\n**But over long durations: The frequency of the signal slowly changes, because:**\n\n- As the neutron star emits gravitational and electromagnetic waves, it loses energy which causes it to rotate more slowly\n- The detector here on Earth is moving with respect to the neutron star. This changes the frequency of the gravitational waves observed in the detector.\n\n**Short-time Fourier Transforms (SFTs)**\n\n- Short-time Fourier Transforms (SFTs) can be used as a way of quantifying the change of a nonstationary signal’s frequency and phase content over time.\n- Very common type of signal preprocessing technique.\n- This what we are given to work with (hdf5 files).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fa8a7caafee2261f27fde68045a824202%2Fstft.PNG?generation=1665180281946268&alt=media)\n\n\n_____\n\n\n##### Timestamps and Frequencies Insights ([Mark Wijkhuizen](https://www.kaggle.com/markwijkhuizen))\n\n**There seems to be 2 differences of our data in comparison to conventional SFT's**\n\n- **The time delta between each timestamp is inconsistent:** Time between measurement in the SFT's can vary greatly.\n- The frequency range is consistently ~0.2Hz, **BUT the min/max frequency ranges from around ~50to ~498Hz.**\n\n- **We should not interpret all spectrograms as continuous sounds:** When taking a look at the timestamps they start at 2009-03-27/28 and end on 2009-07-25/26/27. The recordings are thus gathered during 4 months with often 30 minutes between the recordings. \n\n_____\n\n\n##### Minimum float32 is 1e-38 and data**2 is 1e-44 ([🐢 Jun Koda](https://www.kaggle.com/junkoda))\n\nImportant [post](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361312) by [🐢 Jun Koda](https://www.kaggle.com/junkoda) sharing with us that since the smallest positive number float32 (~ 1e-38) but on the data it is ~1e-22. \n\n**Once we square the data, they go beyond this limit.**\n\n> np.float32(1.23456e-44) => 1.3e-44.\n\n- float32 has about 8 significant digits and 1.23456e-44 will be expressed as 0.0000012e-38 with float32, only 2 digits remaining.\n- Suggestion: that we multiply the data by 1e21 or 1e22.\n\n_____\n\n\n##### PyFstat Signal Parameters ([Mark Wijkhuizen](https://www.kaggle.com/markwijkhuizen))\n\n> Another great [post](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361562) by ([Mark Wijkhuizen](https://www.kaggle.com/markwijkhuizen))\n\nPyFstat has many parameters but it's documentation however is not very helpful for explaining them.\n\n\n**[Victor Gonzalez](https://www.kaggle.com/darkbitur)** Explained with an incredible comment everything about all the paramaeters. \n\n- You can use a helper function to inject some parameters based on priors:\n \n```python\n{**pyfstat.injection_parameters.isotropic_amplitude_priors}\n{'cosi': {'uniform': {'low': -1.0, 'high': 1.0}}, 'psi': {'uniform': {'low': -0.7853981633974483, 'high': 0.7853981633974483}}, 'phi': {'uniform': {'low': 0, 'high': 6.283185307179586}}}\n```\n\n**More Parameters**\n\n- **cosi:** Cosine of the angle between the source and us. Range: [-1, 1]\n- **psi and phi:** polarization angle and phase. Ranges: [-pi/2, pi/2] and [0, 2*pi]\n- **Alpha:** Right ascension of the source's position on the sky. Range: [0, 2*pi]\n- **Delta:** Declination of the source's position on the sky. Range: [-pi/2, pi/2]\n- **Band:** This is just the width of the frequency \"slice\" to be taken. It's 0.2 hz for all files in the train and test datasets, which translates to 360 \"frequency lines\".\n- **F0:** This may be tricky. The generated noise will be centered around this frequency. If a signal is created, it will be also a signal with this frequency (at the source; when it arrives to the detector, doppler and other effects will shift the frequency), so it will appear more or less on the center of the data. Range used on the datasets: [50, 400] Hz.\n- **F1:** This is technically the first derivative of the frequency (how it changs over time). Physically, the sources (neutron stars) lose energy over time because they emit the GWs, so this should be a negative number. A reasonable range could be [-1e-9, 0] Hz but that's up to us.\n- **F2:** This is guess it's the second derivative, which maybe is 0 for the purposes of this competition.\n- **h0:** The amplitude of the signal. In the tutorials, the approach is to define it as a ratio of the sqrtSX. What makes the competition (and the real-world problem of finding continuous GWs) is that expected h0 will be between 1 and 2 orders of magnitude smaller than the amplitude of the noise. In my opinion, part of the challenge is to create your own signals with a distribution of h0 similar to the one used for the test set.\n- **psi** is actually defined within [-pi /4 , pi / 4] (and it's periodic, so anything beyond that folds back onto this range).\n- **F2** is another term describing the intrinsic evolution of a gravitational wave. As @darkbitur correctly guesses, it's set to 0 in this competition as we only focus on frequency and spindown.\n- **tp, asini, period** are parameters used to describe neutron stars in binary systems, and they affect the frequency evolution of the signal. You can safely ignore those in this competition (but do feel free to play around with them if you wish to!).\n\n_____\n\n\n##### Can we assume the same signal trajectories between Hanford & Livingston? ([IkaroSilva](https://www.kaggle.com/ikarosilva))\n\nQuestion [asked](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361655) by [IkaroSilva](https://www.kaggle.com/ikarosilva)\n\n- Can we assume the same signal trajectories between Hanford & Livingston?\n\n**Answer (By The Hosts):**\n\n- The frequency of a signal as seen by a detector a certain time t (what you call it S(t) in your post) does generally depend on the detector, the reason being that different detectors move \"differently\" around the Sun as they are in slightly different positions on Earth.\n\n- This distinction may not be that noticeable for some cases if we use the two Advanced LIGO detectors, since they are relatively close to each other, but becomes quite apparent once you include other detectors such as Virgo, which is located in Italy.\n\n_____\n\n\n##### What is your batch size and LB? ([dragon zhang](https://www.kaggle.com/dragonzhang))\n\nA [good] question about the relation beween the batch-size and generalization performance by [dragon zhang](https://www.kaggle.com/dragonzhang).\n\n\n**Short Answer:**\n\n```\nFriends dont let friends use minibatches larger than 32.\n[Yann LeCun](https://twitter.com/ylecun/status/989610208497360896).\n```\n\n_____\n\n\n##### Recap of the Top Solutions from the Previous G2Net Competition ([Sinan Calisir](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363280))\n\n\nAn amazing solutions [summary](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363280) by [Sinan Calisir](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363280) of all winning solutions of the previous competition.\n\n\n_____\n\n\n##### Weird property of std(spectrogram) in train data ([Josef Slavicek\n](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363460))\n\n\n**Wierd Observation:** When computing std of spectrogram in complex64, I'm getting almost allways the same value of 1.4973568266036865e-22. If I do the same computation in complex128, the values starts do differ - and they are partly useful as indicator of target label:\n\n\n**Answer (From the Hosts):**\n\n**TL;DR:** This is not a data leak, we understand the reason this is happening and it is expected. Nothing to worry about.\n\n> and they are partly useful as indicator of target label\n\nYes, this is expected: stdev is computed from the data, and very strong signals will deviate the value away from the \"noise-only\" value. This should already tell you how \"strong\" signals in the test set are, by the way :) .\n\n> Around 80% of test data seems to be affected by this\n\nThis is also well understood. We wanted to give as much real data as possible in this challenge, but we didn't have a reliable way of generating more than what the LIGO detectors produced during the run, and we did not want to run into issues about having too much correlation among different samples. For this reason, we decided to generate synthetic data using Gaussian noise, for which a PSD (or ASD, or sqrtSX, they are all essentially equivalent in this context) needs to be specified.\n\n> It would us sidechannel which is not present in real GW search data.\n\nThis is not a problem in CW searches. For \"clean\" (relatively quiet) bands we are able to estimate the PSD of the noise with quite good accuracy, and that is sufficient for us to actually perform our analysis. \n\n_____\n\n##### Regarding Those -1 Labels ([PaulG](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363734))\n\n\n- [PaulG](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363734) **Found an easter egg** left by the competition Hosts! \n\nThe dataset description for the competition states:\n``` \n\"Please note the presence of a small number of files labeled -1.\nPhysicists are currently unable to determine the status of these files.\"\n```\n\nAs it turns out, these are easter eggs!.\nIf you plot their signal magnitudes as 2D arrays, you'll see:\n\n\n**50f09e37e**\nA cartoon of two black holes about to collide.\n\n![](https://i.ibb.co/NxnJL9M/50f09e37e.jpg)\n\n**62b0dd011**\nImage that was broadcast into the universe in 1974, as part of the SETI program.\n![](https://i.ibb.co/347VcjZ/62b0dd011.jpg)\n\n**b7666b451**\nThe 2017 physics Nobel Prize winners (for the discovery of gravitational waves):  Rainer Weiss, Barry Barish and Kip Thorne.\n![](https://i.ibb.co/S0NnW3Z/b7666b451.jpg)\n\n_____\n\n\n\n##### High validation accuracy, low LB score ([Simone Bavera](https://www.kaggle.com/simonebvr))\n\nAn [interesting](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363810) discussion about methods for reducing the difference between the LB and the CV score. \n\n\n- **Important** Answer by [Chen Lin](https://www.kaggle.com/chenlin1999)\n\n\n```\nThis competition is very different from the last one. The last contest was the signal of two black holes or neutron stars merging, which is very strong, and implementing a CNN would allow for some level of fit. But continuous gravitational waves are very weak, and even after scientists in the field have implemented various search methods and deep learning tools they have not found any real signal. If you read some of the papers you will see that even the most advanced NNs that have been tried have not reached traditional methods of search in terms of accuracy, and are usually very unsatisfactory after a SNR of less than 10.\n\nI think the huge gap between CV and LB is because 1) the training set is so small with only 600 data, there is so little variation in between that it may be difficult for the network to learn convolution methods in the face of low SNR samples, and 2) I haven't carefully checked the test set, but I think the composition of the test set must be very demanding and there must be a large number of low SNR samples.\n\nIn order to be able to bring CV and LB closer together, I think the most effective way to do this would be to construct a training set that has a large number of samples with a wide range of SNR values, rather than using 600 training samples and then trying various ML tricks.\n```\n\n\n_____\n\n\n\n##### Possible instrumental artifacts in training/testing dataset ([Alex Z](https://www.kaggle.com/alexz0))\n\n\nAn [interesting](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/364854) and informative discussion about the posibility of additional instrumental artifacts (persistent line-like features) to the noise alongside the simulated signals in testing dataset.\n\n\n- **Important** Answer by [Konstantin Dmitriev](https://www.kaggle.com/kdmitrie):\n\n```\nYes, the artifacts like persistent lines do exist alongside with the CW signal and background noise. I've done some statistical testing of the whole dataset (test+train) and found them. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Fdc5855c34dcd37e930419ec3f9039cf2%2F__results___32_1.png?generation=1668240677739409&alt=media)\n\n```\n\n_____\n\n\n\n##### Gap of CV and LB ([Chenglu](https://www.kaggle.com/snaker))\n\n\nAnother [discussion](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/364169) about creating a robust CV strategy.\n\n\n**Important Note** by [Chenglu](https://www.kaggle.com/snaker)\n\n```\nWith the generated dataset, I can reduce the gap to 0.1, CV ~0.77 and LB ~0.68, I think this competition is a matter of shrinking the gap, more precisely, it's generating a dataset that matches the distribution of test set.\n```\n\n_____\n\n\n##### Generating samples that match the train / test distribution ([Chenglu](https://www.kaggle.com/snaker))\n\n\nThank you for sharing this!!\n[Chenglu](https://www.kaggle.com/snaker) is sharing with us a common bug that caused some problems (and you need to also be aware of it):\n\nThe bug: Normalizing the data this way:\n\n```python\nnorm_data = data - data.min() / (data.max() - data.min())\n```\n\nThis is easily influenced by the value of max and min.\nThe fix:\n```\nnorm_data = data * 1e22\n```\n\nThe fixed classifier only achieve 0.66 on classifying generating v.s. testing/training dataset.\n\n_____\n\n\n\n##### What is the difference between the G2Net this year and last year? ([ForcewithMe](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/366451))\n\n**The post lists the differences between the two competitions:**\n\n- The raw data last year are 1d waves, whose x-axis is time, and y-axis is the amplitude of each timestamp. With CQT/CWT, the 1d wave can be converted into a 2d image(spectrum), which x-axis is time, y-axis is frequency and the intensity per unit pixel is the amplitude of a specific frequency within a specific timestamp.\n\n- The raw data are 2d images(spectrum), which x-axis is time, y-axis is frequency and the intensity per unit pixel is the amplitude of a specific frequency within a specific timestamp. (Which is the same as the data after CQT last year in format? I am not sure…)\n\n**SNR**\n- The SNR in the competition this year is much lower, which means that the signal this year is much weaker, resulting in the auc being much lower.\n\n- **Important** Addition by [Alex Z](https://www.kaggle.com/alexz0): Physics-wise, last year competition dealt with the waves from compact binary objects, like close orbiting black holes or white dwarfs. This year we face fast spinning asymmetric neutron stars. Last year set-up provides much more explosive waveforms, and much stronger. Continuous GW we are hunting here were never detected so far, so we have a chance to help a real scientific discovery.\n\n_____\n\n\n\n##### Reverse engineering 20% of test samples with external data. ([Vladimir Slaykovskiy](https://www.kaggle.com/vslaykovsky))\n\n> **TL;DR:** Vladimir Dropped a BOMB. \n\n**It might be possible to reverse engineer the injected CW signal by using raw Ligo&Virgo data.**\n\n- [code](https://www.kaggle.com/vslaykovsky/g2net-winning-strategy-with-external-data)\n\n**In Short**\n\n- Get raw detector data at https://www.gw-openscience.org/archive/O3a_4KHZ_R1/\n- Generate SFTs using instructions at https://youtu.be/A9iWRcmG0Rs?t=2199 (you'll likely need to google for more details on this step). Use GPS timestamps from test set to produce SFTs.\n- Read generated SFTs using pyfstat.utils.sft.get_sft_as_arrays\n- Match generated SFTs with real samples from test set.\n- Find the difference between SFTs and real samples. difference > EPS -> label==1; difference <= EPS -> label==0.\n\n> Wow.. well done on this!\n\n_____\n\n\n\n##### Score improvement G wave detection 20% data ([Igor Litvin](https://www.kaggle.com/igorlitvin))\n\n- [Igor Litvin](https://www.kaggle.com/igorlitvin) proposed a method with some potential for improving your NN training.\n\n- He figered out that aproximately 20% of the gravitation wave spectra are distorted.\n- And think the distortion was made artificially **because of it is exatly the same distortion for H and L shoulders of the interferometer.**\n- It is possible that they multiplied by exactly the same function to H and exactly the same (but different from H) to L shoulder.\n\n- [Code here](https://www.kaggle.com/code/igorlitvin/l-and-h-distortion-of-g-wave-help)\n\n_____\n\n\n\n##### Some Insights About Validation Strategy ([Martin Kovacevic Buvinic](https://www.kaggle.com/ragnar123))\n\n- Great topic by [Martin Kovacevic Buvinic](https://www.kaggle.com/ragnar123) about important insights about the data.\n\n- **CV:** The test data and train data are different: cv gap between validation and test is always large (~0.10). Also low correlation. - BUT! somehow a 20 folds cv model turned out to be much better. \n- **Guess:** Using 20 folds is only 5% of the data as validation for each fold, meaning we are only validating with 30 observations and this lead to overfitting, therefore this validation is not good, increasing the number of folds will just lead to overfitting.\n- **Blending:** Blending in most of the casses helps.\n- **Data Generation:** Adding more data usually improves the generalization of the model, if we check lb it actually improves so adding more data is helpful. The data simulated is the training data, and the best guess is that the key of the competition is to generate data that is similar to the test data.\n- **The Main Insight:** Using training data as validation is not optimal, because it is small and different compared to the test set. A good option would be to generate our own data that is similar to the test set and do traininig and validation with that data, also add the training data.\n\n_____\n\n\n\n##### Merging generated signal into noise ([Vladimir Slaykovskiy](https://www.kaggle.com/vslaykovsky))\n\n[Vladimir Slaykovskiy](https://www.kaggle.com/vslaykovsky) asks about the easies way to combine generated data into the a single spectrogram.\n\n- For example is it correct to just weight-average them?\n`sample = 0.9 * noise + 0.1 + signal`\n\n**Answer (By The Hosts):**\n\n\n```python\nIf noise and signal are Fourier amplitudes (i.e. the complex numbers you get from get_stft_as_arrays) and the generated SFTs have the same timestamps and frequency, then yes.\n\nNote that you can literally do whatever you want to that data. For example, if you wanted to create an instrumental artifact with a circular shape, you could create such a shape in a numpy array with the appropriate array shape and directly added to your data.\n\n`For example is it correct to just weight-average them?`\n\nThe standard way in which we do this is by generating signals the amplitude of which is already a certain fraction of your noise's amplitude (so we can just add both arrays).\n```\n\n\n\n\n\n\n",
      "votes": 34
    },
    {
      "id": 2075382,
      "postDate": "2022-12-25T10:27:03.427Z",
      "content": "<h5>F0</h5>\n<p>F0: This may be tricky. The generated noise will be centered around this frequency. If a signal is created, it will be also a signal with this frequency (at the source; when it arrives to the detector, doppler and other effects will shift the frequency), so it will appear more or less on the center of the data. Range used on the datasets: [50, 400] Hz.</p>\n<p>There is a comment regarding the slicing in a way that F0 is not centered:</p>\n<h5>Slice out the band of interest</h5>\n<p>freqs, times, sft_data = pyfstat.utils.get_sft_as_arrays(signal_writer.sftfilepath)</p>\n<p>first_index = np.argmin(np.abs(freqs - 150.))<br>\nlast_index = np.argmin(np.abs(freqs - 150.2))</p>\n<p>freqs = freqs[first_index:last_index+1]<br>\namplitudes = {key: val[first_index:last_index + 1, :]<br>\n        for key, val in sft_data.items()}</p>\n<p>print(\"*\" * 20)<br>\nprint(\"*\" * 20)<br>\nprint(f\"These should be 150. (got {freqs[0]}) \"<br>\n      f\"and 150.2 (got {freqs[-1]}).\")</p>\n<h5>F1</h5>\n<p>F1: This is technically the first derivative of the frequency (how it changs over time). Physically, the sources (neutron stars) lose energy over time because they emit the GWs, so this should be a negative number. A reasonable range could be [-1e-9, 0] Hz but that's up to us.</p>\n<p>There is also some interest for positive values. </p>\n<h5>h0 and cosi</h5>\n<p>h0: The amplitude of the signal. In the tutorials, the approach is to define it as a ratio of the sqrtSX. What makes the competition (and the real-world problem of finding continuous GWs) is that expected h0 will be between 1 and 2 orders of magnitude smaller than the amplitude of the noise. In my opinion, part of the challenge is to create your own signals with a distribution of h0 similar to the one used for the test set.</p>\n<p>it should be mentioned SNR generation notebook and comments regarding cosi and value of h0</p>\n<p><a href=\"https://www.kaggle.com/code/chris62/generating-low-snr-gravity-waves\" target=\"_blank\">https://www.kaggle.com/code/chris62/generating-low-snr-gravity-waves</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/370202\" target=\"_blank\">https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/370202</a></p>\n<p>Rodrigo Tenorio<br>\nCOMPETITION HOST<br>\nPosted 15 days ago<br>\nThis post earned a bronze medal</p>\n<p>arrow_drop_down<br>\nmore_vert<br>\nHi <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> ,</p>\n<p>You can have a look at equation 95 of this document where the specific dependency of SNR and cosi is discussed (in this case cosi is denoted as eta).</p>\n<p>In a nutshell, cosi=+-1 gives you a stronger signal, while cosi=0 gives you a weaker signal. The specific factor is cosi^4 + 6 * cosi^2 + 1 and that essentially means around a factor 8 of difference between both extreme cases.</p>",
      "rawMarkdown": "##### F0\nF0: This may be tricky. The generated noise will be centered around this frequency. If a signal is created, it will be also a signal with this frequency (at the source; when it arrives to the detector, doppler and other effects will shift the frequency), so it will appear more or less on the center of the data. Range used on the datasets: [50, 400] Hz.\n\nThere is a comment regarding the slicing in a way that F0 is not centered:\n\n##### Slice out the band of interest\nfreqs, times, sft_data = pyfstat.utils.get_sft_as_arrays(signal_writer.sftfilepath)\n\nfirst_index = np.argmin(np.abs(freqs - 150.))\nlast_index = np.argmin(np.abs(freqs - 150.2))\n\nfreqs = freqs[first_index:last_index+1]\namplitudes = {key: val[first_index:last_index + 1, :]\n        for key, val in sft_data.items()}\n\nprint(\"*\" * 20)\nprint(\"*\" * 20)\nprint(f\"These should be 150. (got {freqs[0]}) \"\n      f\"and 150.2 (got {freqs[-1]}).\")\n##### F1\nF1: This is technically the first derivative of the frequency (how it changs over time). Physically, the sources (neutron stars) lose energy over time because they emit the GWs, so this should be a negative number. A reasonable range could be [-1e-9, 0] Hz but that's up to us.\n\nThere is also some interest for positive values. \n\n##### h0 and cosi\nh0: The amplitude of the signal. In the tutorials, the approach is to define it as a ratio of the sqrtSX. What makes the competition (and the real-world problem of finding continuous GWs) is that expected h0 will be between 1 and 2 orders of magnitude smaller than the amplitude of the noise. In my opinion, part of the challenge is to create your own signals with a distribution of h0 similar to the one used for the test set.\n\nit should be mentioned SNR generation notebook and comments regarding cosi and value of h0\n\nhttps://www.kaggle.com/code/chris62/generating-low-snr-gravity-waves\n\nhttps://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/370202\n\nRodrigo Tenorio\nCOMPETITION HOST\nPosted 15 days ago\nThis post earned a bronze medal\n\narrow_drop_down\nmore_vert\nHi @forcewithme ,\n\nYou can have a look at equation 95 of this document where the specific dependency of SNR and cosi is discussed (in this case cosi is denoted as eta).\n\nIn a nutshell, cosi=+-1 gives you a stronger signal, while cosi=0 gives you a weaker signal. The specific factor is cosi^4 + 6 * cosi^2 + 1 and that essentially means around a factor 8 of difference between both extreme cases.",
      "votes": 1
    },
    {
      "id": 2075122,
      "postDate": "2022-12-25T04:58:08.017Z",
      "content": "<p>well written.</p>",
      "rawMarkdown": "well written."
    },
    {
      "id": 2075017,
      "postDate": "2022-12-24T23:23:35.647Z",
      "content": "<p>Wow! That's a lot of precious stuff. I hope they could read it. </p>\n<p>At least, for those that didn't read on the last ending 10 days, it'll remain on the Cloud for the next Gravitational Waves competitions.</p>\n<p>Wonderful job The Devastator.</p>",
      "rawMarkdown": "Wow! That's a lot of precious stuff. I hope they could read it. \n\nAt least, for those that didn't read on the last ending 10 days, it'll remain on the Cloud for the next Gravitational Waves competitions.\n\nWonderful job The Devastator."
    },
    {
      "id": 2078389,
      "postDate": "2022-12-28T09:45:35.563Z",
      "content": "<p>This is super helpful, thanks!</p>",
      "rawMarkdown": "This is super helpful, thanks!"
    },
    {
      "id": 2077379,
      "postDate": "2022-12-27T14:21:31.120Z",
      "content": "<p>Thanks for your summary.</p>",
      "rawMarkdown": "Thanks for your summary."
    },
    {
      "id": 2076873,
      "postDate": "2022-12-27T00:44:32.893Z",
      "content": "<p>Thanks for the helpful article. </p>",
      "rawMarkdown": "Thanks for the helpful article. "
    }
  ],
  "comments": [
    {
      "id": 2075382,
      "author_name": "George Chirita",
      "author_url": "",
      "post_date": "2022-12-25T10:27:03.427000",
      "content": "<h5>F0</h5>\n<p>F0: This may be tricky. The generated noise will be centered around this frequency. If a signal is created, it will be also a signal with this frequency (at the source; when it arrives to the detector, doppler and other effects will shift the frequency), so it will appear more or less on the center of the data. Range used on the datasets: [50, 400] Hz.</p>\n<p>There is a comment regarding the slicing in a way that F0 is not centered:</p>\n<h5>Slice out the band of interest</h5>\n<p>freqs, times, sft_data = pyfstat.utils.get_sft_as_arrays(signal_writer.sftfilepath)</p>\n<p>first_index = np.argmin(np.abs(freqs - 150.))<br>\nlast_index = np.argmin(np.abs(freqs - 150.2))</p>\n<p>freqs = freqs[first_index:last_index+1]<br>\namplitudes = {key: val[first_index:last_index + 1, :]<br>\n        for key, val in sft_data.items()}</p>\n<p>print(\"*\" * 20)<br>\nprint(\"*\" * 20)<br>\nprint(f\"These should be 150. (got {freqs[0]}) \"<br>\n      f\"and 150.2 (got {freqs[-1]}).\")</p>\n<h5>F1</h5>\n<p>F1: This is technically the first derivative of the frequency (how it changs over time). Physically, the sources (neutron stars) lose energy over time because they emit the GWs, so this should be a negative number. A reasonable range could be [-1e-9, 0] Hz but that's up to us.</p>\n<p>There is also some interest for positive values. </p>\n<h5>h0 and cosi</h5>\n<p>h0: The amplitude of the signal. In the tutorials, the approach is to define it as a ratio of the sqrtSX. What makes the competition (and the real-world problem of finding continuous GWs) is that expected h0 will be between 1 and 2 orders of magnitude smaller than the amplitude of the noise. In my opinion, part of the challenge is to create your own signals with a distribution of h0 similar to the one used for the test set.</p>\n<p>it should be mentioned SNR generation notebook and comments regarding cosi and value of h0</p>\n<p><a href=\"https://www.kaggle.com/code/chris62/generating-low-snr-gravity-waves\" target=\"_blank\">https://www.kaggle.com/code/chris62/generating-low-snr-gravity-waves</a></p>\n<p><a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/370202\" target=\"_blank\">https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/370202</a></p>\n<p>Rodrigo Tenorio<br>\nCOMPETITION HOST<br>\nPosted 15 days ago<br>\nThis post earned a bronze medal</p>\n<p>arrow_drop_down<br>\nmore_vert<br>\nHi <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> ,</p>\n<p>You can have a look at equation 95 of this document where the specific dependency of SNR and cosi is discussed (in this case cosi is denoted as eta).</p>\n<p>In a nutshell, cosi=+-1 gives you a stronger signal, while cosi=0 gives you a weaker signal. The specific factor is cosi^4 + 6 * cosi^2 + 1 and that essentially means around a factor 8 of difference between both extreme cases.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2075122,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2022-12-25T04:58:08.017000",
      "content": "<p>well written.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2075017,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2022-12-24T23:23:35.647000",
      "content": "<p>Wow! That's a lot of precious stuff. I hope they could read it. </p>\n<p>At least, for those that didn't read on the last ending 10 days, it'll remain on the Cloud for the next Gravitational Waves competitions.</p>\n<p>Wonderful job The Devastator.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2078389,
      "author_name": "tainangao",
      "author_url": "",
      "post_date": "2022-12-28T09:45:35.563000",
      "content": "<p>This is super helpful, thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2077379,
      "author_name": "Ursula Lu",
      "author_url": "",
      "post_date": "2022-12-27T14:21:31.120000",
      "content": "<p>Thanks for your summary.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2076873,
      "author_name": "plasmabio",
      "author_url": "",
      "post_date": "2022-12-27T00:44:32.893000",
      "content": "<p>Thanks for the helpful article. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2074716": "### What do we know so far? \n##### 10 Days to go!\n_____\nSave yourself some time and get caught up on the latest updates and important discussions that have taken place prior to the competition deadline.\nThe competition ends in 10 days and there is still time to improve your solution. Here are some of the important topics that have been discussed in the Kaggle forums.\n_____\n\n\n##### Background\n\n- The data is H-U-G-E and comes in hdf5 format, [chazzer](https://www.kaggle.com/chazzer) had written a [notebook on how to read the hdf5 files](https://www.kaggle.com/code/chazzer/how-to-read-the-hdf5-files/)\n\n- Shortly after [JohnM](https://www.kaggle.com/jpmiller) pushed this one step further and created [HDF files with indexed dataframes for the training set](https://www.kaggle.com/datasets/jpmiller/simplified-dataset). The index on each frame is the frequency array from the original file, and the columns are the timestamps.\n\n**Reading Example**\n\n```python\ndf_h = pd.read_hdf(\"../input/simplified-dataset/001121a05.h5\", key = 'h')\n```\n\n- Great [list](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/357682) of learning resources had been published by [Laura Fink](https://www.kaggle.com/allunia)\n    - [Gravitational Wave Open Science Center - Tutorials](https://www.gw-openscience.org/tutorials/)\n    - [GWPy](https://gwpy.github.io/docs/stable/index.html)\n    - [Video about the Michelson Interferometer](https://www.youtube.com/watch?v=tQ_teIUb3tE)\n    - [Video about gravitational waves](https://www.youtube.com/watch?v=FlDtXIBrAYE&t=1s)\n    - [Great tutorial video about Fourier transforms](https://www.youtube.com/watch?v=spUNpyF58BY&t=1s)\n    - [Video about Denoising data with Fast Fourier Transforms](https://www.youtube.com/watch?v=s2K1JfNR7Sc)\n \n_____\n\n##### Best CV-LB Results ([Mark Baushenko](https://www.kaggle.com/e0xextazy))\n\n- **Important** Comments by [DrHB](drhabib)\n    - \"Without `augs` the gap (CV-LB) is bigger\"\n    - \"After running some experiments .. I think the biggest challenge of this competition will be to build a proper validation set\"    \n    \n_____\n\n\n##### Understanding the Data ([Ravi Shah](https://www.kaggle.com/ravishah1))\n\nAn [amazing intro post](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/358444) by [Ravi Shah](https://www.kaggle.com/ravishah1) going through all important details about the data of this competition. \n\n**In short (check the original for more info)**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F1893bf6ae0ffb282512eec0fdce5a662%2Fstruc.PNG?generation=1665177635295503&alt=media)\n\n- L1 - data from the LIGO Livingston interferometer\n- H1 - data from the LIGO Hanford interferometer\n- STFs - Short-time Fourier Transforms (SFTs) - shape (360, n)\n- timestamps - the timestamps that the STFs correspond to - shape (n,)\n- frequency Hz - the range frequencies measured by the detectors - shape (360,)\n\n**Competition Goal**\n\nThe goal of this competition is to create a model that can detect continuous gravitational-wave signals. These are weak yet long-lasting signals emitted by rapidly-spinning neutron stars within noisy data. \n\n> - **Target is 1 if signal is present else 0**\n\n\n**LIGO** (where the data comes from) is a detector in Livingston, Louisiana and Hanford, Washington. The detector has a L shape with equal-length arms connected to a corner station. When a gravitational wave passes by, one leg of the detector is shortened while the other is lengthened. \n\n**The interference makes a shift that we can analyze.**\n\n\n**We Get A Continuous Gravitational Waves**\n\nA continuous gravitational-wave signal from a Galactic neutron star will look almost perfectly constant in both frequency and amplitude.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F6b476a7c1f076bdf7c4f7234d37ace3b%2Fperfect%20cont%20wave.PNG?generation=1665181150773656&alt=media)\n\n**But over long durations: The frequency of the signal slowly changes, because:**\n\n- As the neutron star emits gravitational and electromagnetic waves, it loses energy which causes it to rotate more slowly\n- The detector here on Earth is moving with respect to the neutron star. This changes the frequency of the gravitational waves observed in the detector.\n\n**Short-time Fourier Transforms (SFTs)**\n\n- Short-time Fourier Transforms (SFTs) can be used as a way of quantifying the change of a nonstationary signal’s frequency and phase content over time.\n- Very common type of signal preprocessing technique.\n- This what we are given to work with (hdf5 files).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fa8a7caafee2261f27fde68045a824202%2Fstft.PNG?generation=1665180281946268&alt=media)\n\n\n_____\n\n\n##### Timestamps and Frequencies Insights ([Mark Wijkhuizen](https://www.kaggle.com/markwijkhuizen))\n\n**There seems to be 2 differences of our data in comparison to conventional SFT's**\n\n- **The time delta between each timestamp is inconsistent:** Time between measurement in the SFT's can vary greatly.\n- The frequency range is consistently ~0.2Hz, **BUT the min/max frequency ranges from around ~50to ~498Hz.**\n\n- **We should not interpret all spectrograms as continuous sounds:** When taking a look at the timestamps they start at 2009-03-27/28 and end on 2009-07-25/26/27. The recordings are thus gathered during 4 months with often 30 minutes between the recordings. \n\n_____\n\n\n##### Minimum float32 is 1e-38 and data**2 is 1e-44 ([🐢 Jun Koda](https://www.kaggle.com/junkoda))\n\nImportant [post](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361312) by [🐢 Jun Koda](https://www.kaggle.com/junkoda) sharing with us that since the smallest positive number float32 (~ 1e-38) but on the data it is ~1e-22. \n\n**Once we square the data, they go beyond this limit.**\n\n> np.float32(1.23456e-44) => 1.3e-44.\n\n- float32 has about 8 significant digits and 1.23456e-44 will be expressed as 0.0000012e-38 with float32, only 2 digits remaining.\n- Suggestion: that we multiply the data by 1e21 or 1e22.\n\n_____\n\n\n##### PyFstat Signal Parameters ([Mark Wijkhuizen](https://www.kaggle.com/markwijkhuizen))\n\n> Another great [post](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361562) by ([Mark Wijkhuizen](https://www.kaggle.com/markwijkhuizen))\n\nPyFstat has many parameters but it's documentation however is not very helpful for explaining them.\n\n\n**[Victor Gonzalez](https://www.kaggle.com/darkbitur)** Explained with an incredible comment everything about all the paramaeters. \n\n- You can use a helper function to inject some parameters based on priors:\n \n```python\n{**pyfstat.injection_parameters.isotropic_amplitude_priors}\n{'cosi': {'uniform': {'low': -1.0, 'high': 1.0}}, 'psi': {'uniform': {'low': -0.7853981633974483, 'high': 0.7853981633974483}}, 'phi': {'uniform': {'low': 0, 'high': 6.283185307179586}}}\n```\n\n**More Parameters**\n\n- **cosi:** Cosine of the angle between the source and us. Range: [-1, 1]\n- **psi and phi:** polarization angle and phase. Ranges: [-pi/2, pi/2] and [0, 2*pi]\n- **Alpha:** Right ascension of the source's position on the sky. Range: [0, 2*pi]\n- **Delta:** Declination of the source's position on the sky. Range: [-pi/2, pi/2]\n- **Band:** This is just the width of the frequency \"slice\" to be taken. It's 0.2 hz for all files in the train and test datasets, which translates to 360 \"frequency lines\".\n- **F0:** This may be tricky. The generated noise will be centered around this frequency. If a signal is created, it will be also a signal with this frequency (at the source; when it arrives to the detector, doppler and other effects will shift the frequency), so it will appear more or less on the center of the data. Range used on the datasets: [50, 400] Hz.\n- **F1:** This is technically the first derivative of the frequency (how it changs over time). Physically, the sources (neutron stars) lose energy over time because they emit the GWs, so this should be a negative number. A reasonable range could be [-1e-9, 0] Hz but that's up to us.\n- **F2:** This is guess it's the second derivative, which maybe is 0 for the purposes of this competition.\n- **h0:** The amplitude of the signal. In the tutorials, the approach is to define it as a ratio of the sqrtSX. What makes the competition (and the real-world problem of finding continuous GWs) is that expected h0 will be between 1 and 2 orders of magnitude smaller than the amplitude of the noise. In my opinion, part of the challenge is to create your own signals with a distribution of h0 similar to the one used for the test set.\n- **psi** is actually defined within [-pi /4 , pi / 4] (and it's periodic, so anything beyond that folds back onto this range).\n- **F2** is another term describing the intrinsic evolution of a gravitational wave. As @darkbitur correctly guesses, it's set to 0 in this competition as we only focus on frequency and spindown.\n- **tp, asini, period** are parameters used to describe neutron stars in binary systems, and they affect the frequency evolution of the signal. You can safely ignore those in this competition (but do feel free to play around with them if you wish to!).\n\n_____\n\n\n##### Can we assume the same signal trajectories between Hanford & Livingston? ([IkaroSilva](https://www.kaggle.com/ikarosilva))\n\nQuestion [asked](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361655) by [IkaroSilva](https://www.kaggle.com/ikarosilva)\n\n- Can we assume the same signal trajectories between Hanford & Livingston?\n\n**Answer (By The Hosts):**\n\n- The frequency of a signal as seen by a detector a certain time t (what you call it S(t) in your post) does generally depend on the detector, the reason being that different detectors move \"differently\" around the Sun as they are in slightly different positions on Earth.\n\n- This distinction may not be that noticeable for some cases if we use the two Advanced LIGO detectors, since they are relatively close to each other, but becomes quite apparent once you include other detectors such as Virgo, which is located in Italy.\n\n_____\n\n\n##### What is your batch size and LB? ([dragon zhang](https://www.kaggle.com/dragonzhang))\n\nA [good] question about the relation beween the batch-size and generalization performance by [dragon zhang](https://www.kaggle.com/dragonzhang).\n\n\n**Short Answer:**\n\n```\nFriends dont let friends use minibatches larger than 32.\n[Yann LeCun](https://twitter.com/ylecun/status/989610208497360896).\n```\n\n_____\n\n\n##### Recap of the Top Solutions from the Previous G2Net Competition ([Sinan Calisir](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363280))\n\n\nAn amazing solutions [summary](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363280) by [Sinan Calisir](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363280) of all winning solutions of the previous competition.\n\n\n_____\n\n\n##### Weird property of std(spectrogram) in train data ([Josef Slavicek\n](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363460))\n\n\n**Wierd Observation:** When computing std of spectrogram in complex64, I'm getting almost allways the same value of 1.4973568266036865e-22. If I do the same computation in complex128, the values starts do differ - and they are partly useful as indicator of target label:\n\n\n**Answer (From the Hosts):**\n\n**TL;DR:** This is not a data leak, we understand the reason this is happening and it is expected. Nothing to worry about.\n\n> and they are partly useful as indicator of target label\n\nYes, this is expected: stdev is computed from the data, and very strong signals will deviate the value away from the \"noise-only\" value. This should already tell you how \"strong\" signals in the test set are, by the way :) .\n\n> Around 80% of test data seems to be affected by this\n\nThis is also well understood. We wanted to give as much real data as possible in this challenge, but we didn't have a reliable way of generating more than what the LIGO detectors produced during the run, and we did not want to run into issues about having too much correlation among different samples. For this reason, we decided to generate synthetic data using Gaussian noise, for which a PSD (or ASD, or sqrtSX, they are all essentially equivalent in this context) needs to be specified.\n\n> It would us sidechannel which is not present in real GW search data.\n\nThis is not a problem in CW searches. For \"clean\" (relatively quiet) bands we are able to estimate the PSD of the noise with quite good accuracy, and that is sufficient for us to actually perform our analysis. \n\n_____\n\n##### Regarding Those -1 Labels ([PaulG](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363734))\n\n\n- [PaulG](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363734) **Found an easter egg** left by the competition Hosts! \n\nThe dataset description for the competition states:\n``` \n\"Please note the presence of a small number of files labeled -1.\nPhysicists are currently unable to determine the status of these files.\"\n```\n\nAs it turns out, these are easter eggs!.\nIf you plot their signal magnitudes as 2D arrays, you'll see:\n\n\n**50f09e37e**\nA cartoon of two black holes about to collide.\n\n![](https://i.ibb.co/NxnJL9M/50f09e37e.jpg)\n\n**62b0dd011**\nImage that was broadcast into the universe in 1974, as part of the SETI program.\n![](https://i.ibb.co/347VcjZ/62b0dd011.jpg)\n\n**b7666b451**\nThe 2017 physics Nobel Prize winners (for the discovery of gravitational waves):  Rainer Weiss, Barry Barish and Kip Thorne.\n![](https://i.ibb.co/S0NnW3Z/b7666b451.jpg)\n\n_____\n\n\n\n##### High validation accuracy, low LB score ([Simone Bavera](https://www.kaggle.com/simonebvr))\n\nAn [interesting](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363810) discussion about methods for reducing the difference between the LB and the CV score. \n\n\n- **Important** Answer by [Chen Lin](https://www.kaggle.com/chenlin1999)\n\n\n```\nThis competition is very different from the last one. The last contest was the signal of two black holes or neutron stars merging, which is very strong, and implementing a CNN would allow for some level of fit. But continuous gravitational waves are very weak, and even after scientists in the field have implemented various search methods and deep learning tools they have not found any real signal. If you read some of the papers you will see that even the most advanced NNs that have been tried have not reached traditional methods of search in terms of accuracy, and are usually very unsatisfactory after a SNR of less than 10.\n\nI think the huge gap between CV and LB is because 1) the training set is so small with only 600 data, there is so little variation in between that it may be difficult for the network to learn convolution methods in the face of low SNR samples, and 2) I haven't carefully checked the test set, but I think the composition of the test set must be very demanding and there must be a large number of low SNR samples.\n\nIn order to be able to bring CV and LB closer together, I think the most effective way to do this would be to construct a training set that has a large number of samples with a wide range of SNR values, rather than using 600 training samples and then trying various ML tricks.\n```\n\n\n_____\n\n\n\n##### Possible instrumental artifacts in training/testing dataset ([Alex Z](https://www.kaggle.com/alexz0))\n\n\nAn [interesting](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/364854) and informative discussion about the posibility of additional instrumental artifacts (persistent line-like features) to the noise alongside the simulated signals in testing dataset.\n\n\n- **Important** Answer by [Konstantin Dmitriev](https://www.kaggle.com/kdmitrie):\n\n```\nYes, the artifacts like persistent lines do exist alongside with the CW signal and background noise. I've done some statistical testing of the whole dataset (test+train) and found them. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4308868%2Fdc5855c34dcd37e930419ec3f9039cf2%2F__results___32_1.png?generation=1668240677739409&alt=media)\n\n```\n\n_____\n\n\n\n##### Gap of CV and LB ([Chenglu](https://www.kaggle.com/snaker))\n\n\nAnother [discussion](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/364169) about creating a robust CV strategy.\n\n\n**Important Note** by [Chenglu](https://www.kaggle.com/snaker)\n\n```\nWith the generated dataset, I can reduce the gap to 0.1, CV ~0.77 and LB ~0.68, I think this competition is a matter of shrinking the gap, more precisely, it's generating a dataset that matches the distribution of test set.\n```\n\n_____\n\n\n##### Generating samples that match the train / test distribution ([Chenglu](https://www.kaggle.com/snaker))\n\n\nThank you for sharing this!!\n[Chenglu](https://www.kaggle.com/snaker) is sharing with us a common bug that caused some problems (and you need to also be aware of it):\n\nThe bug: Normalizing the data this way:\n\n```python\nnorm_data = data - data.min() / (data.max() - data.min())\n```\n\nThis is easily influenced by the value of max and min.\nThe fix:\n```\nnorm_data = data * 1e22\n```\n\nThe fixed classifier only achieve 0.66 on classifying generating v.s. testing/training dataset.\n\n_____\n\n\n\n##### What is the difference between the G2Net this year and last year? ([ForcewithMe](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/366451))\n\n**The post lists the differences between the two competitions:**\n\n- The raw data last year are 1d waves, whose x-axis is time, and y-axis is the amplitude of each timestamp. With CQT/CWT, the 1d wave can be converted into a 2d image(spectrum), which x-axis is time, y-axis is frequency and the intensity per unit pixel is the amplitude of a specific frequency within a specific timestamp.\n\n- The raw data are 2d images(spectrum), which x-axis is time, y-axis is frequency and the intensity per unit pixel is the amplitude of a specific frequency within a specific timestamp. (Which is the same as the data after CQT last year in format? I am not sure…)\n\n**SNR**\n- The SNR in the competition this year is much lower, which means that the signal this year is much weaker, resulting in the auc being much lower.\n\n- **Important** Addition by [Alex Z](https://www.kaggle.com/alexz0): Physics-wise, last year competition dealt with the waves from compact binary objects, like close orbiting black holes or white dwarfs. This year we face fast spinning asymmetric neutron stars. Last year set-up provides much more explosive waveforms, and much stronger. Continuous GW we are hunting here were never detected so far, so we have a chance to help a real scientific discovery.\n\n_____\n\n\n\n##### Reverse engineering 20% of test samples with external data. ([Vladimir Slaykovskiy](https://www.kaggle.com/vslaykovsky))\n\n> **TL;DR:** Vladimir Dropped a BOMB. \n\n**It might be possible to reverse engineer the injected CW signal by using raw Ligo&Virgo data.**\n\n- [code](https://www.kaggle.com/vslaykovsky/g2net-winning-strategy-with-external-data)\n\n**In Short**\n\n- Get raw detector data at https://www.gw-openscience.org/archive/O3a_4KHZ_R1/\n- Generate SFTs using instructions at https://youtu.be/A9iWRcmG0Rs?t=2199 (you'll likely need to google for more details on this step). Use GPS timestamps from test set to produce SFTs.\n- Read generated SFTs using pyfstat.utils.sft.get_sft_as_arrays\n- Match generated SFTs with real samples from test set.\n- Find the difference between SFTs and real samples. difference > EPS -> label==1; difference <= EPS -> label==0.\n\n> Wow.. well done on this!\n\n_____\n\n\n\n##### Score improvement G wave detection 20% data ([Igor Litvin](https://www.kaggle.com/igorlitvin))\n\n- [Igor Litvin](https://www.kaggle.com/igorlitvin) proposed a method with some potential for improving your NN training.\n\n- He figered out that aproximately 20% of the gravitation wave spectra are distorted.\n- And think the distortion was made artificially **because of it is exatly the same distortion for H and L shoulders of the interferometer.**\n- It is possible that they multiplied by exactly the same function to H and exactly the same (but different from H) to L shoulder.\n\n- [Code here](https://www.kaggle.com/code/igorlitvin/l-and-h-distortion-of-g-wave-help)\n\n_____\n\n\n\n##### Some Insights About Validation Strategy ([Martin Kovacevic Buvinic](https://www.kaggle.com/ragnar123))\n\n- Great topic by [Martin Kovacevic Buvinic](https://www.kaggle.com/ragnar123) about important insights about the data.\n\n- **CV:** The test data and train data are different: cv gap between validation and test is always large (~0.10). Also low correlation. - BUT! somehow a 20 folds cv model turned out to be much better. \n- **Guess:** Using 20 folds is only 5% of the data as validation for each fold, meaning we are only validating with 30 observations and this lead to overfitting, therefore this validation is not good, increasing the number of folds will just lead to overfitting.\n- **Blending:** Blending in most of the casses helps.\n- **Data Generation:** Adding more data usually improves the generalization of the model, if we check lb it actually improves so adding more data is helpful. The data simulated is the training data, and the best guess is that the key of the competition is to generate data that is similar to the test data.\n- **The Main Insight:** Using training data as validation is not optimal, because it is small and different compared to the test set. A good option would be to generate our own data that is similar to the test set and do traininig and validation with that data, also add the training data.\n\n_____\n\n\n\n##### Merging generated signal into noise ([Vladimir Slaykovskiy](https://www.kaggle.com/vslaykovsky))\n\n[Vladimir Slaykovskiy](https://www.kaggle.com/vslaykovsky) asks about the easies way to combine generated data into the a single spectrogram.\n\n- For example is it correct to just weight-average them?\n`sample = 0.9 * noise + 0.1 + signal`\n\n**Answer (By The Hosts):**\n\n\n```python\nIf noise and signal are Fourier amplitudes (i.e. the complex numbers you get from get_stft_as_arrays) and the generated SFTs have the same timestamps and frequency, then yes.\n\nNote that you can literally do whatever you want to that data. For example, if you wanted to create an instrumental artifact with a circular shape, you could create such a shape in a numpy array with the appropriate array shape and directly added to your data.\n\n`For example is it correct to just weight-average them?`\n\nThe standard way in which we do this is by generating signals the amplitude of which is already a certain fraction of your noise's amplitude (so we can just add both arrays).\n```\n\n\n\n\n\n\n",
    "2075382": "##### F0\nF0: This may be tricky. The generated noise will be centered around this frequency. If a signal is created, it will be also a signal with this frequency (at the source; when it arrives to the detector, doppler and other effects will shift the frequency), so it will appear more or less on the center of the data. Range used on the datasets: [50, 400] Hz.\n\nThere is a comment regarding the slicing in a way that F0 is not centered:\n\n##### Slice out the band of interest\nfreqs, times, sft_data = pyfstat.utils.get_sft_as_arrays(signal_writer.sftfilepath)\n\nfirst_index = np.argmin(np.abs(freqs - 150.))\nlast_index = np.argmin(np.abs(freqs - 150.2))\n\nfreqs = freqs[first_index:last_index+1]\namplitudes = {key: val[first_index:last_index + 1, :]\n        for key, val in sft_data.items()}\n\nprint(\"*\" * 20)\nprint(\"*\" * 20)\nprint(f\"These should be 150. (got {freqs[0]}) \"\n      f\"and 150.2 (got {freqs[-1]}).\")\n##### F1\nF1: This is technically the first derivative of the frequency (how it changs over time). Physically, the sources (neutron stars) lose energy over time because they emit the GWs, so this should be a negative number. A reasonable range could be [-1e-9, 0] Hz but that's up to us.\n\nThere is also some interest for positive values. \n\n##### h0 and cosi\nh0: The amplitude of the signal. In the tutorials, the approach is to define it as a ratio of the sqrtSX. What makes the competition (and the real-world problem of finding continuous GWs) is that expected h0 will be between 1 and 2 orders of magnitude smaller than the amplitude of the noise. In my opinion, part of the challenge is to create your own signals with a distribution of h0 similar to the one used for the test set.\n\nit should be mentioned SNR generation notebook and comments regarding cosi and value of h0\n\nhttps://www.kaggle.com/code/chris62/generating-low-snr-gravity-waves\n\nhttps://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/370202\n\nRodrigo Tenorio\nCOMPETITION HOST\nPosted 15 days ago\nThis post earned a bronze medal\n\narrow_drop_down\nmore_vert\nHi @forcewithme ,\n\nYou can have a look at equation 95 of this document where the specific dependency of SNR and cosi is discussed (in this case cosi is denoted as eta).\n\nIn a nutshell, cosi=+-1 gives you a stronger signal, while cosi=0 gives you a weaker signal. The specific factor is cosi^4 + 6 * cosi^2 + 1 and that essentially means around a factor 8 of difference between both extreme cases.",
    "2075122": "well written.",
    "2075017": "Wow! That's a lot of precious stuff. I hope they could read it. \n\nAt least, for those that didn't read on the last ending 10 days, it'll remain on the Cloud for the next Gravitational Waves competitions.\n\nWonderful job The Devastator.",
    "2078389": "This is super helpful, thanks!",
    "2077379": "Thanks for your summary.",
    "2076873": "Thanks for the helpful article. "
  }
}