{
  "id": 373461,
  "title": "Augmentations For This Competition",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/373461",
  "author_name": "The Devastator",
  "post_date": "2022-12-21T15:38:28.770000",
  "votes": 13,
  "comment_count": 0,
  "views": 0,
  "content": "<h3>Augmentations</h3>\n<p>There are many options for augmentation we can apply to this competition some on the spectral level while others on the signal itself. In this post I try to go over the most promising ones:</p>\n<ul>\n<li>Used last year by the winning solutions </li>\n<li>And some that might also be intuitive to try</li>\n</ul>\n<h5>Last Year <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275476\" target=\"_blank\">1st</a> <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275507\" target=\"_blank\">place</a> solution</h5>\n<p>Used the following augmentations (Implementations - My own. No code released for this solution)</p>\n<p><strong>Time Masking</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Masking the time of the input signal by a certain amount. This can be done by multiplying the time domain representation of the audio signal by a masking function.</p>\n</blockquote>\n<pre><code> ():    \n     nn.Sequential(torchaudio.transforms.TimeMasking(time_mask_param = time_mask_param),)(torch_img)\n</code></pre>\n<p><strong>Frequency Masking:</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Masking the frequency of the input audio signal by a certain amount. This can be done by multiplying the Fourier transform of the audio signal by a masking function. </p>\n</blockquote>\n<pre><code> ():\n     nn.Sequential(torchaudio.transforms.FrequencyMasking(freq_mask_param = freq_mask_param),)(torch_img)\n</code></pre>\n<blockquote>\n  <p>The first place writeup also used minor time shifts between channels</p>\n</blockquote>\n<p><strong>Combined Time And Freq: SpecAug</strong> (Not used by the 1st place solution)</p>\n<blockquote>\n  <p>**Paper:<a href=\"https://arxiv.org/abs/1904.08779\" target=\"_blank\">SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. Park, et al.</a></p>\n</blockquote>\n<pre><code> ():\n     nn.Sequential(torchaudio.transforms.FrequencyMasking(freq_mask_param = freq_mask_param),\n                         torchaudio.transforms.TimeMasking(time_mask_param = time_mask_param))(torch_img)\n</code></pre>\n<h5>Last Year's <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275341\" target=\"_blank\">2nd place solution</a></h5>\n<blockquote>\n  <p>Source: <a href=\"https://github.com/analokmaus/kaggle-g2net-public\" target=\"_blank\">code release</a></p>\n</blockquote>\n<ul>\n<li><p>Used 2 augmentations. </p></li>\n<li><p>Gaussian noise for 2d-CNN networks</p></li>\n</ul>\n<pre><code> ():\n     ():\n        ().__init__(always_apply, p)\n        self.min_snr = min_snr\n        self.max_snr = max_snr\n\n     ():\n        snr = np.random.uniform(self.min_snr, self.max_snr)\n        white_noise = np.random.randn((y))\n         add_noise_snr(y, white_noise, snr)\n</code></pre>\n<ul>\n<li>Flipped wave amplitude for 1d-CNN </li>\n</ul>\n<pre><code> ():\n     (): ().__init__(always_apply, p)\n     ():  y * -        \n</code></pre>\n<h5>Last Year's <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275331\" target=\"_blank\">4th place solution</a></h5>\n<ul>\n<li>Used 2 augmentations: Shifts and colored noise.</li>\n</ul>\n<p><strong>Vertical Shifting</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Shifting the image vertically by a certain amount.</p>\n</blockquote>\n<pre><code> ():\n    img[] = np.roll(img[], vshift, axis = )\n    img[] = np.roll(img[], vshift, axis = )\n     img\n</code></pre>\n<p><strong>Colored Noise Augmentation</strong><br>\nColored noise augmentation was done channel-wise (No further information was provided).</p>\n<h5>Last Year's <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275343\" target=\"_blank\">6th place solution</a></h5>\n<ul>\n<li>Used 2 augmentations: Random shifts and Random change phase.</li>\n</ul>\n<p><strong>Random Phase Augmentation</strong> (Implementation: My own. Might be inaccurate)</p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Randomly adding and subtracting phase to the spectral representation of the audio signal (Fourier basis and ifft). </p>\n</blockquote>\n<pre><code>  ():\n    rand_phase = np.exp( * np.random.uniform(,  * np.pi, Z.shape))\n     np.(Z) * rand_phase\n</code></pre>\n<h5>Last Year's <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275349\" target=\"_blank\">7th place solution</a></h5>\n<ul>\n<li>Used the following augmentations: Roll +/-500 points (p=0.5), Scale 0.85 - 1.15 (p=0.5), -1 * Multiply (p=0.5)</li>\n</ul>\n<h5>Last Year's <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275335\" target=\"_blank\">8th place solution</a></h5>\n<ul>\n<li>Used the following 4 augmentations: Zebra Mixup, Negative Flip, Swap with Other Negative, LIGO Swap</li>\n</ul>\n<p><strong>Zebra Mixup</strong></p>\n<p>They report that this augmentation works really well.</p>\n<pre><code>wave_new[:, ::] = wave1[:, ::]\nwave_new[:, ::] = wave2[:, ::]\n</code></pre>\n<p><strong>Negative Flip</strong><br>\nThey report that this worked for only negative sample</p>\n<pre><code>wave = wave[:, ::-].copy()\n</code></pre>\n<p><strong>Swap with Other Negative</strong></p>\n<p>Negative samples swapped with other negative samples</p>\n<pre><code> np.random.uniform() &gt; : wave1[] = wave2[]\n np.random.uniform() &gt; : wave1[] = wave2[]\n np.random.uniform() &gt; : wave1[] = wave2[]\n</code></pre>\n<p><strong>LIGO Swap</strong></p>\n<p>Since both logos are completely the same - we can swap them.</p>\n<pre><code>w0 = wave1[].copy()\nwave1[] = wave1[]\nwave1[] = w0\n</code></pre>\n<h5>Last Year's <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275405\" target=\"_blank\">10th place solution</a></h5>\n<ul>\n<li>Used the following augmentations: Random shift, Random shift each channel, Random turning off of a single channel, Random swapping channel of the two waves (target=0 only)</li>\n</ul>\n<p><strong>Random shift each channel</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Randomly shifting each channel in the time domain representation of the audio signal by a certain amount.</p>\n</blockquote>\n<pre><code> ():\n\n    shift = np.random.randint(-max_shift, max_shift, )\n\n    img[][:, ] = np.roll(img[][:, ], shift[], axis = )\n    img[][:, ] = np.roll(img[][:, ], shift[], axis = )\n    img[][:, ] = np.roll(img[][:, ], shift[], axis = )\n    img[][:, ] = np.roll(img[][:, ], shift[], axis = )\n     img\n</code></pre>\n<p><strong>Random turning off of a single channel</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Randomly turning off a channel of the audio signal.</p>\n</blockquote>\n<pre><code> ():\n    channel = np.random.randint()\n    img[][:, channel] = np.zeros(img[].shape[])\n    img[][:, channel] = np.zeros(img[].shape[])\n     img\n</code></pre>\n<p><strong>Random swapping channel of the two waves (target=0 only)</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Randomly swapping channels of the audio signal.</p>\n</blockquote>\n<pre><code> ():\n    channel = np.random.randint()\n    img_swap = img[][:, channel].copy()\n    img[][:, channel] = img[][:, channel]\n    img[][:, channel] = img_swap\n     img\n</code></pre>\n<h5>Last Year's <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275405\" target=\"_blank\">11th place solution</a></h5>\n<ul>\n<li>Used 2 augmentations: before converting the signal to a spectrogram: MixUp and Randomly rolling time shifts</li>\n</ul>\n<p><strong>MixUp</strong></p>\n<blockquote>\n  <p><strong>Paper: <a href=\"https://arxiv.org/abs/1710.09412\" target=\"_blank\">Mixup: Beyond Empirical Risk Minimization</a></strong></p>\n  <p><strong>TL;DR:</strong> Mixing the signals together.</p>\n</blockquote>\n<pre><code> ():\n    lam = np.random.beta(alpha, alpha)\n    img = lam * img1 + ( - lam) * img2\n     img\n</code></pre>\n<p><strong>Randomly rolling time shifts</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Shifting the signal by a random amount.</p>\n</blockquote>\n<pre><code> ():\n\n    time_shift = np.random.randint(-time_shift, time_shift, )[]\n\n    img[] = np.roll(img[], time_shift, axis = )\n    img[] = np.roll(img[], time_shift, axis = )\n     img\n</code></pre>\n<h4>Other Augmentations To Try</h4>\n<p>The followings are more augmentation functions that based on the above, also might work well on this competition.</p>\n<p><strong>Horizontal Flipping</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Horizontal flipping of the image. </p>\n</blockquote>\n<pre><code> ():\n    img[] = np.fliplr(img[])\n    img[] = np.fliplr(img[])\n     img\n</code></pre>\n<p><strong>Vertical Flipping</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Vertical flipping of the image.</p>\n</blockquote>\n<pre><code> ():\n    img[] = np.flipud(img[])\n    img[] = np.flipud(img[])\n     img\n</code></pre>\n<p><strong>Time stretching</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Time stretching of the input audio signal by a certain amount. This can be done by interpolating the audio signal.</p>\n</blockquote>\n<pre><code> ():\n     nn.Sequential(torchaudio.transforms.TimeStretch(rate=),)(torch_img)\n</code></pre>\n<p><strong>Pitch shifting</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Shifting the pitch of the input audio signal by a certain amount. This can be done by interpolating the audio signal.</p>\n</blockquote>\n<pre><code> ():\n     nn.Sequential(torchaudio.transforms.PitchShift(n_steps=),)(torch_img)\n</code></pre>\n<p><strong>Noise injection</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Add noise to the signal by a certain amount in relation to SNR.</p>\n</blockquote>\n<p>$$SNR = \\frac{P_{signal}}{P_{noise}}$$</p>\n<p>Noise can come in different flavors depending on what we are modeling, a good start is Additive White Gaussian Noise (AWGN). To model AWGN we need to add a zero-mean gaussian random variable to our original signal. The variance of that random variable will affect the average noise power.</p>\n<p>For a Gaussian random variable X, the average power E[X^2], also known as the second moment:</p>\n<p>$$E[X^2] = \\sigma_x^2 + \\mu_x^2$$</p>\n<pre><code>\n ():\n    signal_power = (torch_img**).mean()\n    noise_power = signal_power/(**(snr/))\n    noise = (torch.randn(torch_img.size())*(noise_power**)).(torch.cuda.FloatTensor)\n     (torch_img + noise)\n</code></pre>",
  "messages": [
    {
      "id": 2072001,
      "postDate": "2022-12-21T15:38:28.770Z",
      "content": "<h3>Augmentations</h3>\n<p>There are many options for augmentation we can apply to this competition some on the spectral level while others on the signal itself. In this post I try to go over the most promising ones:</p>\n<ul>\n<li>Used last year by the winning solutions </li>\n<li>And some that might also be intuitive to try</li>\n</ul>\n<h5>Last Year <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275476\" target=\"_blank\">1st</a> <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275507\" target=\"_blank\">place</a> solution</h5>\n<p>Used the following augmentations (Implementations - My own. No code released for this solution)</p>\n<p><strong>Time Masking</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Masking the time of the input signal by a certain amount. This can be done by multiplying the time domain representation of the audio signal by a masking function.</p>\n</blockquote>\n<pre><code> ():    \n     nn.Sequential(torchaudio.transforms.TimeMasking(time_mask_param = time_mask_param),)(torch_img)\n</code></pre>\n<p><strong>Frequency Masking:</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Masking the frequency of the input audio signal by a certain amount. This can be done by multiplying the Fourier transform of the audio signal by a masking function. </p>\n</blockquote>\n<pre><code> ():\n     nn.Sequential(torchaudio.transforms.FrequencyMasking(freq_mask_param = freq_mask_param),)(torch_img)\n</code></pre>\n<blockquote>\n  <p>The first place writeup also used minor time shifts between channels</p>\n</blockquote>\n<p><strong>Combined Time And Freq: SpecAug</strong> (Not used by the 1st place solution)</p>\n<blockquote>\n  <p>**Paper:<a href=\"https://arxiv.org/abs/1904.08779\" target=\"_blank\">SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. Park, et al.</a></p>\n</blockquote>\n<pre><code> ():\n     nn.Sequential(torchaudio.transforms.FrequencyMasking(freq_mask_param = freq_mask_param),\n                         torchaudio.transforms.TimeMasking(time_mask_param = time_mask_param))(torch_img)\n</code></pre>\n<h5>Last Year's <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275341\" target=\"_blank\">2nd place solution</a></h5>\n<blockquote>\n  <p>Source: <a href=\"https://github.com/analokmaus/kaggle-g2net-public\" target=\"_blank\">code release</a></p>\n</blockquote>\n<ul>\n<li><p>Used 2 augmentations. </p></li>\n<li><p>Gaussian noise for 2d-CNN networks</p></li>\n</ul>\n<pre><code> ():\n     ():\n        ().__init__(always_apply, p)\n        self.min_snr = min_snr\n        self.max_snr = max_snr\n\n     ():\n        snr = np.random.uniform(self.min_snr, self.max_snr)\n        white_noise = np.random.randn((y))\n         add_noise_snr(y, white_noise, snr)\n</code></pre>\n<ul>\n<li>Flipped wave amplitude for 1d-CNN </li>\n</ul>\n<pre><code> ():\n     (): ().__init__(always_apply, p)\n     ():  y * -        \n</code></pre>\n<h5>Last Year's <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275331\" target=\"_blank\">4th place solution</a></h5>\n<ul>\n<li>Used 2 augmentations: Shifts and colored noise.</li>\n</ul>\n<p><strong>Vertical Shifting</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Shifting the image vertically by a certain amount.</p>\n</blockquote>\n<pre><code> ():\n    img[] = np.roll(img[], vshift, axis = )\n    img[] = np.roll(img[], vshift, axis = )\n     img\n</code></pre>\n<p><strong>Colored Noise Augmentation</strong><br>\nColored noise augmentation was done channel-wise (No further information was provided).</p>\n<h5>Last Year's <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275343\" target=\"_blank\">6th place solution</a></h5>\n<ul>\n<li>Used 2 augmentations: Random shifts and Random change phase.</li>\n</ul>\n<p><strong>Random Phase Augmentation</strong> (Implementation: My own. Might be inaccurate)</p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Randomly adding and subtracting phase to the spectral representation of the audio signal (Fourier basis and ifft). </p>\n</blockquote>\n<pre><code>  ():\n    rand_phase = np.exp( * np.random.uniform(,  * np.pi, Z.shape))\n     np.(Z) * rand_phase\n</code></pre>\n<h5>Last Year's <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275349\" target=\"_blank\">7th place solution</a></h5>\n<ul>\n<li>Used the following augmentations: Roll +/-500 points (p=0.5), Scale 0.85 - 1.15 (p=0.5), -1 * Multiply (p=0.5)</li>\n</ul>\n<h5>Last Year's <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275335\" target=\"_blank\">8th place solution</a></h5>\n<ul>\n<li>Used the following 4 augmentations: Zebra Mixup, Negative Flip, Swap with Other Negative, LIGO Swap</li>\n</ul>\n<p><strong>Zebra Mixup</strong></p>\n<p>They report that this augmentation works really well.</p>\n<pre><code>wave_new[:, ::] = wave1[:, ::]\nwave_new[:, ::] = wave2[:, ::]\n</code></pre>\n<p><strong>Negative Flip</strong><br>\nThey report that this worked for only negative sample</p>\n<pre><code>wave = wave[:, ::-].copy()\n</code></pre>\n<p><strong>Swap with Other Negative</strong></p>\n<p>Negative samples swapped with other negative samples</p>\n<pre><code> np.random.uniform() &gt; : wave1[] = wave2[]\n np.random.uniform() &gt; : wave1[] = wave2[]\n np.random.uniform() &gt; : wave1[] = wave2[]\n</code></pre>\n<p><strong>LIGO Swap</strong></p>\n<p>Since both logos are completely the same - we can swap them.</p>\n<pre><code>w0 = wave1[].copy()\nwave1[] = wave1[]\nwave1[] = w0\n</code></pre>\n<h5>Last Year's <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275405\" target=\"_blank\">10th place solution</a></h5>\n<ul>\n<li>Used the following augmentations: Random shift, Random shift each channel, Random turning off of a single channel, Random swapping channel of the two waves (target=0 only)</li>\n</ul>\n<p><strong>Random shift each channel</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Randomly shifting each channel in the time domain representation of the audio signal by a certain amount.</p>\n</blockquote>\n<pre><code> ():\n\n    shift = np.random.randint(-max_shift, max_shift, )\n\n    img[][:, ] = np.roll(img[][:, ], shift[], axis = )\n    img[][:, ] = np.roll(img[][:, ], shift[], axis = )\n    img[][:, ] = np.roll(img[][:, ], shift[], axis = )\n    img[][:, ] = np.roll(img[][:, ], shift[], axis = )\n     img\n</code></pre>\n<p><strong>Random turning off of a single channel</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Randomly turning off a channel of the audio signal.</p>\n</blockquote>\n<pre><code> ():\n    channel = np.random.randint()\n    img[][:, channel] = np.zeros(img[].shape[])\n    img[][:, channel] = np.zeros(img[].shape[])\n     img\n</code></pre>\n<p><strong>Random swapping channel of the two waves (target=0 only)</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Randomly swapping channels of the audio signal.</p>\n</blockquote>\n<pre><code> ():\n    channel = np.random.randint()\n    img_swap = img[][:, channel].copy()\n    img[][:, channel] = img[][:, channel]\n    img[][:, channel] = img_swap\n     img\n</code></pre>\n<h5>Last Year's <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275405\" target=\"_blank\">11th place solution</a></h5>\n<ul>\n<li>Used 2 augmentations: before converting the signal to a spectrogram: MixUp and Randomly rolling time shifts</li>\n</ul>\n<p><strong>MixUp</strong></p>\n<blockquote>\n  <p><strong>Paper: <a href=\"https://arxiv.org/abs/1710.09412\" target=\"_blank\">Mixup: Beyond Empirical Risk Minimization</a></strong></p>\n  <p><strong>TL;DR:</strong> Mixing the signals together.</p>\n</blockquote>\n<pre><code> ():\n    lam = np.random.beta(alpha, alpha)\n    img = lam * img1 + ( - lam) * img2\n     img\n</code></pre>\n<p><strong>Randomly rolling time shifts</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Shifting the signal by a random amount.</p>\n</blockquote>\n<pre><code> ():\n\n    time_shift = np.random.randint(-time_shift, time_shift, )[]\n\n    img[] = np.roll(img[], time_shift, axis = )\n    img[] = np.roll(img[], time_shift, axis = )\n     img\n</code></pre>\n<h4>Other Augmentations To Try</h4>\n<p>The followings are more augmentation functions that based on the above, also might work well on this competition.</p>\n<p><strong>Horizontal Flipping</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Horizontal flipping of the image. </p>\n</blockquote>\n<pre><code> ():\n    img[] = np.fliplr(img[])\n    img[] = np.fliplr(img[])\n     img\n</code></pre>\n<p><strong>Vertical Flipping</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Vertical flipping of the image.</p>\n</blockquote>\n<pre><code> ():\n    img[] = np.flipud(img[])\n    img[] = np.flipud(img[])\n     img\n</code></pre>\n<p><strong>Time stretching</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Time stretching of the input audio signal by a certain amount. This can be done by interpolating the audio signal.</p>\n</blockquote>\n<pre><code> ():\n     nn.Sequential(torchaudio.transforms.TimeStretch(rate=),)(torch_img)\n</code></pre>\n<p><strong>Pitch shifting</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Shifting the pitch of the input audio signal by a certain amount. This can be done by interpolating the audio signal.</p>\n</blockquote>\n<pre><code> ():\n     nn.Sequential(torchaudio.transforms.PitchShift(n_steps=),)(torch_img)\n</code></pre>\n<p><strong>Noise injection</strong></p>\n<blockquote>\n  <p><strong>TL;DR:</strong> Add noise to the signal by a certain amount in relation to SNR.</p>\n</blockquote>\n<p>$$SNR = \\frac{P_{signal}}{P_{noise}}$$</p>\n<p>Noise can come in different flavors depending on what we are modeling, a good start is Additive White Gaussian Noise (AWGN). To model AWGN we need to add a zero-mean gaussian random variable to our original signal. The variance of that random variable will affect the average noise power.</p>\n<p>For a Gaussian random variable X, the average power E[X^2], also known as the second moment:</p>\n<p>$$E[X^2] = \\sigma_x^2 + \\mu_x^2$$</p>\n<pre><code>\n ():\n    signal_power = (torch_img**).mean()\n    noise_power = signal_power/(**(snr/))\n    noise = (torch.randn(torch_img.size())*(noise_power**)).(torch.cuda.FloatTensor)\n     (torch_img + noise)\n</code></pre>",
      "rawMarkdown": "### Augmentations\n\nThere are many options for augmentation we can apply to this competition some on the spectral level while others on the signal itself. In this post I try to go over the most promising ones:\n\n- Used last year by the winning solutions \n- And some that might also be intuitive to try\n\n##### Last Year [1st](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275476) [place](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275507) solution\n\nUsed the following augmentations (Implementations - My own. No code released for this solution)\n\n**Time Masking**\n\n> **TL;DR:** Masking the time of the input signal by a certain amount. This can be done by multiplying the time domain representation of the audio signal by a masking function.\n\n```python\ndef augment_time_mask(torch_img, time_mask_param = 10):    \n    return nn.Sequential(torchaudio.transforms.TimeMasking(time_mask_param = time_mask_param),)(torch_img)\n```\n\n**Frequency Masking:**\n\n> **TL;DR:** Masking the frequency of the input audio signal by a certain amount. This can be done by multiplying the Fourier transform of the audio signal by a masking function. \n\n```python\ndef augment_freq_mask(torch_img, freq_mask_param = 10):\n    return nn.Sequential(torchaudio.transforms.FrequencyMasking(freq_mask_param = freq_mask_param),)(torch_img)\n```\n\n> The first place writeup also used minor time shifts between channels\n\n**Combined Time And Freq: SpecAug** (Not used by the 1st place solution)\n\n> **Paper:[SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. Park, et al.](https://arxiv.org/abs/1904.08779)\n\n```python\ndef augment_spec_augment(torch_img, time_mask_param = 20, freq_mask_param = 20):\n    return nn.Sequential(torchaudio.transforms.FrequencyMasking(freq_mask_param = freq_mask_param),\n                         torchaudio.transforms.TimeMasking(time_mask_param = time_mask_param))(torch_img)\n```\n\n\n##### Last Year's [2nd place solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275341)\n\n> Source: [code release](https://github.com/analokmaus/kaggle-g2net-public)\n\n- Used 2 augmentations. \n\n- Gaussian noise for 2d-CNN networks\n\n```python\nclass GaussianNoiseSNR(AudioTransformPerChannel):\n    def __init__(self, always_apply=False, p=0.5, min_snr=5.0, max_snr=20.0, **kwargs):\n        super().__init__(always_apply, p)\n        self.min_snr = min_snr\n        self.max_snr = max_snr\n\n    def apply(self, y: np.ndarray):\n        snr = np.random.uniform(self.min_snr, self.max_snr)\n        white_noise = np.random.randn(len(y))\n        return add_noise_snr(y, white_noise, snr)\n```\n\n- Flipped wave amplitude for 1d-CNN \n        \n```python\nclass FlipWave(AudioTransform):\n    def __init__(self, always_apply = False, p = 0.5): super().__init__(always_apply, p)\n    def apply(self, y: np.ndarray): return y * -1        \n```\n\n\n##### Last Year's [4th place solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275331)\n\n- Used 2 augmentations: Shifts and colored noise.\n\n**Vertical Shifting**\n\n> **TL;DR:** Shifting the image vertically by a certain amount.\n\n```python\ndef augment_vertical_shift(img, vshift):\n    img[0] = np.roll(img[0], vshift, axis = 0)\n    img[1] = np.roll(img[1], vshift, axis = 0)\n    return img\n\n```\n\n**Colored Noise Augmentation**\nColored noise augmentation was done channel-wise (No further information was provided).\n\n\n##### Last Year's [6th place solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275343)\n\n- Used 2 augmentations: Random shifts and Random change phase.\n\n**Random Phase Augmentation** (Implementation: My own. Might be inaccurate)\n\n> **TL;DR:** Randomly adding and subtracting phase to the spectral representation of the audio signal (Fourier basis and ifft). \n\n```python\n def add_random_phase(Z):\n    rand_phase = np.exp(1j * np.random.uniform(0, 2 * np.pi, Z.shape))\n    return np.abs(Z) * rand_phase\n```\n\n\n##### Last Year's [7th place solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275349)\n\n- Used the following augmentations: Roll +/-500 points (p=0.5), Scale 0.85 - 1.15 (p=0.5), -1 * Multiply (p=0.5)\n\n\n\n##### Last Year's [8th place solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275335)\n\n- Used the following 4 augmentations: Zebra Mixup, Negative Flip, Swap with Other Negative, LIGO Swap\n\n**Zebra Mixup**\n\nThey report that this augmentation works really well.\n\n```python\nwave_new[:, 0:4096:2] = wave1[:, 0:4096:2]\nwave_new[:, 1:4096:2] = wave2[:, 1:4096:2]\n```\n\n**Negative Flip**\nThey report that this worked for only negative sample\n\n```python\nwave = wave[:, ::-1].copy()\n```\n\n**Swap with Other Negative**\n\nNegative samples swapped with other negative samples\n\n```python\nif np.random.uniform() > 0.5: wave1[0] = wave2[0]\nif np.random.uniform() > 0.5: wave1[1] = wave2[1]\nif np.random.uniform() > 0.5: wave1[2] = wave2[2]\n```\n\n**LIGO Swap**\n\nSince both logos are completely the same - we can swap them.\n\n```python\nw0 = wave1[0].copy()\nwave1[0] = wave1[1]\nwave1[1] = w0\n```\n\n##### Last Year's [10th place solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275405)\n\n- Used the following augmentations: Random shift, Random shift each channel, Random turning off of a single channel, Random swapping channel of the two waves (target=0 only)\n\n**Random shift each channel**\n\n> **TL;DR:** Randomly shifting each channel in the time domain representation of the audio signal by a certain amount.\n\n```python\ndef augment_random_shift_each_channel(img, max_shift = 20):\n    \n    shift = np.random.randint(-max_shift, max_shift, 2)\n    \n    img[0][:, 0] = np.roll(img[0][:, 0], shift[0], axis = 0)\n    img[0][:, 1] = np.roll(img[0][:, 1], shift[1], axis = 0)\n    img[1][:, 0] = np.roll(img[1][:, 0], shift[0], axis = 0)\n    img[1][:, 1] = np.roll(img[1][:, 1], shift[1], axis = 0)\n    return img\n```\n\n**Random turning off of a single channel**\n\n> **TL;DR:** Randomly turning off a channel of the audio signal.\n\n```python\ndef augment_random_channel_drop(img):\n    channel = np.random.randint(2)\n    img[0][:, channel] = np.zeros(img[0].shape[0])\n    img[1][:, channel] = np.zeros(img[1].shape[0])\n    return img\n```\n\n**Random swapping channel of the two waves (target=0 only)**\n\n> **TL;DR:** Randomly swapping channels of the audio signal.\n\n```python\ndef augment_random_channel_swap(img):\n    channel = np.random.randint(2)\n    img_swap = img[0][:, channel].copy()\n    img[0][:, channel] = img[1][:, channel]\n    img[1][:, channel] = img_swap\n    return img\n```\n\n\n##### Last Year's [11th place solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275405)\n\n- Used 2 augmentations: before converting the signal to a spectrogram: MixUp and Randomly rolling time shifts\n\n**MixUp**\n\n> **Paper: [Mixup: Beyond Empirical Risk Minimization](https://arxiv.org/abs/1710.09412)**\n\n> **TL;DR:** Mixing the signals together.\n\n```python\ndef augment_mixup(img1, img2, alpha = 1.0):\n    lam = np.random.beta(alpha, alpha)\n    img = lam * img1 + (1 - lam) * img2\n    return img\n```\n\n**Randomly rolling time shifts**\n\n> **TL;DR:** Shifting the signal by a random amount.\n\n```python\ndef augment_random_time_shift(img, time_shift = 0):\n    \n    time_shift = np.random.randint(-time_shift, time_shift, 1)[0]\n    \n    img[0] = np.roll(img[0], time_shift, axis = 0)\n    img[1] = np.roll(img[1], time_shift, axis = 0)\n    return img\n```\n\n\n\n#### Other Augmentations To Try\n\nThe followings are more augmentation functions that based on the above, also might work well on this competition.\n\n**Horizontal Flipping**\n\n> **TL;DR:** Horizontal flipping of the image. \n\n```python\ndef augment_horizontal_flipping(img):\n    img[0] = np.fliplr(img[0])\n    img[1] = np.fliplr(img[1])\n    return img\n```\n\n**Vertical Flipping**\n\n> **TL;DR:** Vertical flipping of the image.\n\n\n```python\ndef augment_vertical_flipping(img):\n    img[0] = np.flipud(img[0])\n    img[1] = np.flipud(img[1])\n    return img\n```\n\n**Time stretching**\n\n> **TL;DR:** Time stretching of the input audio signal by a certain amount. This can be done by interpolating the audio signal.\n\n```python\ndef augment_time_stretch(torch_img):\n    return nn.Sequential(torchaudio.transforms.TimeStretch(rate=1.5),)(torch_img)\n```\n\n**Pitch shifting**\n\n> **TL;DR:** Shifting the pitch of the input audio signal by a certain amount. This can be done by interpolating the audio signal.\n\n```python\ndef augment_pitch_shift(torch_img):\n    return nn.Sequential(torchaudio.transforms.PitchShift(n_steps=2),)(torch_img)\n```\n\n**Noise injection**\n\n> **TL;DR:** Add noise to the signal by a certain amount in relation to SNR.\n\n$$SNR = \\frac{P_{signal}}{P_{noise}}$$\n\nNoise can come in different flavors depending on what we are modeling, a good start is Additive White Gaussian Noise (AWGN). To model AWGN we need to add a zero-mean gaussian random variable to our original signal. The variance of that random variable will affect the average noise power.\n\nFor a Gaussian random variable X, the average power E[X^2], also known as the second moment:\n\n$$E[X^2] = \\sigma_x^2 + \\mu_x^2$$\n\n\n```python\n# Adding noise using target SNR\ndef augment_noise_injection(torch_img, snr=40):\n    signal_power = (torch_img**2).mean()\n    noise_power = signal_power/(10**(snr/10))\n    noise = (torch.randn(torch_img.size())*(noise_power**0.5)).type(torch.cuda.FloatTensor)\n    return (torch_img + noise)\n```\n\n",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2072001": "### Augmentations\n\nThere are many options for augmentation we can apply to this competition some on the spectral level while others on the signal itself. In this post I try to go over the most promising ones:\n\n- Used last year by the winning solutions \n- And some that might also be intuitive to try\n\n##### Last Year [1st](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275476) [place](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275507) solution\n\nUsed the following augmentations (Implementations - My own. No code released for this solution)\n\n**Time Masking**\n\n> **TL;DR:** Masking the time of the input signal by a certain amount. This can be done by multiplying the time domain representation of the audio signal by a masking function.\n\n```python\ndef augment_time_mask(torch_img, time_mask_param = 10):    \n    return nn.Sequential(torchaudio.transforms.TimeMasking(time_mask_param = time_mask_param),)(torch_img)\n```\n\n**Frequency Masking:**\n\n> **TL;DR:** Masking the frequency of the input audio signal by a certain amount. This can be done by multiplying the Fourier transform of the audio signal by a masking function. \n\n```python\ndef augment_freq_mask(torch_img, freq_mask_param = 10):\n    return nn.Sequential(torchaudio.transforms.FrequencyMasking(freq_mask_param = freq_mask_param),)(torch_img)\n```\n\n> The first place writeup also used minor time shifts between channels\n\n**Combined Time And Freq: SpecAug** (Not used by the 1st place solution)\n\n> **Paper:[SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. Park, et al.](https://arxiv.org/abs/1904.08779)\n\n```python\ndef augment_spec_augment(torch_img, time_mask_param = 20, freq_mask_param = 20):\n    return nn.Sequential(torchaudio.transforms.FrequencyMasking(freq_mask_param = freq_mask_param),\n                         torchaudio.transforms.TimeMasking(time_mask_param = time_mask_param))(torch_img)\n```\n\n\n##### Last Year's [2nd place solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275341)\n\n> Source: [code release](https://github.com/analokmaus/kaggle-g2net-public)\n\n- Used 2 augmentations. \n\n- Gaussian noise for 2d-CNN networks\n\n```python\nclass GaussianNoiseSNR(AudioTransformPerChannel):\n    def __init__(self, always_apply=False, p=0.5, min_snr=5.0, max_snr=20.0, **kwargs):\n        super().__init__(always_apply, p)\n        self.min_snr = min_snr\n        self.max_snr = max_snr\n\n    def apply(self, y: np.ndarray):\n        snr = np.random.uniform(self.min_snr, self.max_snr)\n        white_noise = np.random.randn(len(y))\n        return add_noise_snr(y, white_noise, snr)\n```\n\n- Flipped wave amplitude for 1d-CNN \n        \n```python\nclass FlipWave(AudioTransform):\n    def __init__(self, always_apply = False, p = 0.5): super().__init__(always_apply, p)\n    def apply(self, y: np.ndarray): return y * -1        \n```\n\n\n##### Last Year's [4th place solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275331)\n\n- Used 2 augmentations: Shifts and colored noise.\n\n**Vertical Shifting**\n\n> **TL;DR:** Shifting the image vertically by a certain amount.\n\n```python\ndef augment_vertical_shift(img, vshift):\n    img[0] = np.roll(img[0], vshift, axis = 0)\n    img[1] = np.roll(img[1], vshift, axis = 0)\n    return img\n\n```\n\n**Colored Noise Augmentation**\nColored noise augmentation was done channel-wise (No further information was provided).\n\n\n##### Last Year's [6th place solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275343)\n\n- Used 2 augmentations: Random shifts and Random change phase.\n\n**Random Phase Augmentation** (Implementation: My own. Might be inaccurate)\n\n> **TL;DR:** Randomly adding and subtracting phase to the spectral representation of the audio signal (Fourier basis and ifft). \n\n```python\n def add_random_phase(Z):\n    rand_phase = np.exp(1j * np.random.uniform(0, 2 * np.pi, Z.shape))\n    return np.abs(Z) * rand_phase\n```\n\n\n##### Last Year's [7th place solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275349)\n\n- Used the following augmentations: Roll +/-500 points (p=0.5), Scale 0.85 - 1.15 (p=0.5), -1 * Multiply (p=0.5)\n\n\n\n##### Last Year's [8th place solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275335)\n\n- Used the following 4 augmentations: Zebra Mixup, Negative Flip, Swap with Other Negative, LIGO Swap\n\n**Zebra Mixup**\n\nThey report that this augmentation works really well.\n\n```python\nwave_new[:, 0:4096:2] = wave1[:, 0:4096:2]\nwave_new[:, 1:4096:2] = wave2[:, 1:4096:2]\n```\n\n**Negative Flip**\nThey report that this worked for only negative sample\n\n```python\nwave = wave[:, ::-1].copy()\n```\n\n**Swap with Other Negative**\n\nNegative samples swapped with other negative samples\n\n```python\nif np.random.uniform() > 0.5: wave1[0] = wave2[0]\nif np.random.uniform() > 0.5: wave1[1] = wave2[1]\nif np.random.uniform() > 0.5: wave1[2] = wave2[2]\n```\n\n**LIGO Swap**\n\nSince both logos are completely the same - we can swap them.\n\n```python\nw0 = wave1[0].copy()\nwave1[0] = wave1[1]\nwave1[1] = w0\n```\n\n##### Last Year's [10th place solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275405)\n\n- Used the following augmentations: Random shift, Random shift each channel, Random turning off of a single channel, Random swapping channel of the two waves (target=0 only)\n\n**Random shift each channel**\n\n> **TL;DR:** Randomly shifting each channel in the time domain representation of the audio signal by a certain amount.\n\n```python\ndef augment_random_shift_each_channel(img, max_shift = 20):\n    \n    shift = np.random.randint(-max_shift, max_shift, 2)\n    \n    img[0][:, 0] = np.roll(img[0][:, 0], shift[0], axis = 0)\n    img[0][:, 1] = np.roll(img[0][:, 1], shift[1], axis = 0)\n    img[1][:, 0] = np.roll(img[1][:, 0], shift[0], axis = 0)\n    img[1][:, 1] = np.roll(img[1][:, 1], shift[1], axis = 0)\n    return img\n```\n\n**Random turning off of a single channel**\n\n> **TL;DR:** Randomly turning off a channel of the audio signal.\n\n```python\ndef augment_random_channel_drop(img):\n    channel = np.random.randint(2)\n    img[0][:, channel] = np.zeros(img[0].shape[0])\n    img[1][:, channel] = np.zeros(img[1].shape[0])\n    return img\n```\n\n**Random swapping channel of the two waves (target=0 only)**\n\n> **TL;DR:** Randomly swapping channels of the audio signal.\n\n```python\ndef augment_random_channel_swap(img):\n    channel = np.random.randint(2)\n    img_swap = img[0][:, channel].copy()\n    img[0][:, channel] = img[1][:, channel]\n    img[1][:, channel] = img_swap\n    return img\n```\n\n\n##### Last Year's [11th place solution](https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275405)\n\n- Used 2 augmentations: before converting the signal to a spectrogram: MixUp and Randomly rolling time shifts\n\n**MixUp**\n\n> **Paper: [Mixup: Beyond Empirical Risk Minimization](https://arxiv.org/abs/1710.09412)**\n\n> **TL;DR:** Mixing the signals together.\n\n```python\ndef augment_mixup(img1, img2, alpha = 1.0):\n    lam = np.random.beta(alpha, alpha)\n    img = lam * img1 + (1 - lam) * img2\n    return img\n```\n\n**Randomly rolling time shifts**\n\n> **TL;DR:** Shifting the signal by a random amount.\n\n```python\ndef augment_random_time_shift(img, time_shift = 0):\n    \n    time_shift = np.random.randint(-time_shift, time_shift, 1)[0]\n    \n    img[0] = np.roll(img[0], time_shift, axis = 0)\n    img[1] = np.roll(img[1], time_shift, axis = 0)\n    return img\n```\n\n\n\n#### Other Augmentations To Try\n\nThe followings are more augmentation functions that based on the above, also might work well on this competition.\n\n**Horizontal Flipping**\n\n> **TL;DR:** Horizontal flipping of the image. \n\n```python\ndef augment_horizontal_flipping(img):\n    img[0] = np.fliplr(img[0])\n    img[1] = np.fliplr(img[1])\n    return img\n```\n\n**Vertical Flipping**\n\n> **TL;DR:** Vertical flipping of the image.\n\n\n```python\ndef augment_vertical_flipping(img):\n    img[0] = np.flipud(img[0])\n    img[1] = np.flipud(img[1])\n    return img\n```\n\n**Time stretching**\n\n> **TL;DR:** Time stretching of the input audio signal by a certain amount. This can be done by interpolating the audio signal.\n\n```python\ndef augment_time_stretch(torch_img):\n    return nn.Sequential(torchaudio.transforms.TimeStretch(rate=1.5),)(torch_img)\n```\n\n**Pitch shifting**\n\n> **TL;DR:** Shifting the pitch of the input audio signal by a certain amount. This can be done by interpolating the audio signal.\n\n```python\ndef augment_pitch_shift(torch_img):\n    return nn.Sequential(torchaudio.transforms.PitchShift(n_steps=2),)(torch_img)\n```\n\n**Noise injection**\n\n> **TL;DR:** Add noise to the signal by a certain amount in relation to SNR.\n\n$$SNR = \\frac{P_{signal}}{P_{noise}}$$\n\nNoise can come in different flavors depending on what we are modeling, a good start is Additive White Gaussian Noise (AWGN). To model AWGN we need to add a zero-mean gaussian random variable to our original signal. The variance of that random variable will affect the average noise power.\n\nFor a Gaussian random variable X, the average power E[X^2], also known as the second moment:\n\n$$E[X^2] = \\sigma_x^2 + \\mu_x^2$$\n\n\n```python\n# Adding noise using target SNR\ndef augment_noise_injection(torch_img, snr=40):\n    signal_power = (torch_img**2).mean()\n    noise_power = signal_power/(10**(snr/10))\n    noise = (torch.randn(torch_img.size())*(noise_power**0.5)).type(torch.cuda.FloatTensor)\n    return (torch_img + noise)\n```\n\n"
  }
}