{
  "id": 370202,
  "title": "Low SNR Experiments",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/370202",
  "author_name": "chris",
  "post_date": "2022-12-03T16:08:53.897000",
  "votes": 56,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I took a step back to do some experiments to see how SNR affects the ability of networks (CNN and efficientnet) to detect signals. Here are some results (nothing earth shattering I think, but still interesting):</p>\n<h2>Step 1. Generate Waves</h2>\n<p>I modified the tutorial code in order to generate waves of varying strengths. I made 3 different simulated datasets, one at 5% signal to noise (SNR) (1:20), 2% (1:50) and 1% (1:100). The code to make those sets is here (I ran it 3 different times at 3 different h0 values): <a href=\"https://www.kaggle.com/chris62/generating-low-snr-gravity-waves\" target=\"_blank\">https://www.kaggle.com/chris62/generating-low-snr-gravity-waves</a></p>\n<p>Note that I only am making 7.5 day long signals, which result in a 360x360 image - this is much shorter than the full data in the competition (&gt; 4000 px wide), but was much easier to manage and use for my experiments. For my experiments below, I generated 1,000 train images and 1,000 validation images for each dataset, with 50% signals and 50% noise only.</p>\n<p>I also generated examples at 1:1 SNR, 1:10 and 1:100, just to see what the results looked like, and converted them into images using the \"power\" calculation for H1 and L1, and then subtracted the mean and divided by the standard deviation:</p>\n<pre><code> ():\n    data = np.load(filename)\n\n    img = np.zeros((,,IMGSIZE))\n\n    img[] = data[,:,:]** + data[,:,:]**\n    img[] = data[,:,:]** + data[,:,:]**\n\n    img = img - np.mean(img)\n    img = img / np.std(img)\n\n     img.transpose(,,)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fe3d51a2fa6ab78c425c43c7feef0a52a%2Fsnr.png?generation=1670081815216592&amp;alt=media\" alt=\"\"></p>\n<p>The 100% SNR signal totally overwhelms the noise, and after normalization, that's all you can see… how easy the competition would be if that was the case! 😆 </p>\n<p>Unfortunately, the competition signals are 1 to 2 orders of magnitude weaker than the noise, which means the signals will be somewhere between the 10% and 1% SNR image above. At 1% SNR you can't see the signal at all - so we have our work cut out for us! </p>\n<h2>Step 2. Simple CNN architecture</h2>\n<p>I wanted to test the 5%, 2% and 1% dataset on a non-pre-trained, simple architecture, so I created a CNN model that would take the 2 channel 360x360 images, and return a single value (1 for signal, 0 for noise):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fae8379131e73a6a2cabd575e24439106%2FScreen%20Shot%202022-12-03%20at%2010.42.06%20AM.png?generation=1670082226219825&amp;alt=media\" alt=\"\"></p>\n<p>The only really interesting part of that is the first convolution, which is 5x31. My theory was that a very wide first convolution would be helpful in detecting the mostly horizontal waves, and it seemed to work fairly well in my experiments.</p>\n<p>My first tests were with just 100 training and validation samples across a wide range of SNR values - from 1.0 all the way down to 0.01.  I trained that model for 25 epochs using the AdamW optimizer with a learning rate of 1e-4 using binary cross entropy with logits loss.</p>\n<p>From 1.0 down to 0.1 (100% SNR to 10%), the model was very easily to get close to 1.0 AUC, which is totally expected given the how easy it is to distinguish the signals above:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fb16259ff469c0e4bc397bc95679ed0ba%2FScreen%20Shot%202022-12-03%20at%2010.48.15%20AM.png?generation=1670082510716865&amp;alt=media\" alt=\"\"></p>\n<p>Then I tested from 10% down to 1% SNR (1:100), and the AUC started going down - crossing 0.5 between 1% and 2% (0.5 AUC means total guessing):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2F9a100c472a502e0aa56af8ac5a4b1474%2FScreen%20Shot%202022-12-03%20at%2010.48.55%20AM.png?generation=1670082576318418&amp;alt=media\" alt=\"\"></p>\n<p>So from that we can take that 1. there is some hope, since AUC was &gt; 0.5 all the way down to 2%, BUT it obviously becomes increasingly difficult to detect signals at lower and lower SNRs. (Again, not earth shattering, but still interesting I think).</p>\n<p>The next step was to try the CNN on my 5%, 2% and 1% datasets, but I had a ton of trouble actually training - the network continuously overfit:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2F23f584f3d2201cec84bd85b1271478cb%2FScreen%20Shot%202022-12-03%20at%2010.50.55%20AM.png?generation=1670082698310513&amp;alt=media\" alt=\"\"></p>\n<p>Eventually I was able to get it to train by adding 0.3 dropout and data augmentations (same as <a href=\"https://www.kaggle.com/code/leolu1998/g2net-basic-audio-data-augmentation-inference\" target=\"_blank\">this notebook</a> ) but ONLY got good results for the 5% dataset!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fa02b81a9324d43e2399c3adea20356a1%2FScreen%20Shot%202022-12-03%20at%2010.54.02%20AM.png?generation=1670082857103133&amp;alt=media\" alt=\"\"></p>\n<p>I was finally able to achieve a validation loss of 0.344 with an AUC of 0.814 for the 1:20 (5% SNR) dataset. That 0.814 looks good - but! since it's only on the 5% dataset, it's actually not as good as it seemed, and I wasn't able to get the 2% or 1% dataset to really train at all.</p>\n<p>Learning from this step: Getting all the way to 1:100 SNR wasn't possible for me with a simple architecture. Time to bring out the big network!</p>\n<h2>Step 3. Pre-trained efficientnet</h2>\n<p>Next I switch to a tf_efficientnet_b7_ns pretrained backbone, with a custom classifier head similar to <a href=\"https://www.kaggle.com/code/leolu1998/g2net-basic-audio-data-augmentation-inference\" target=\"_blank\">this notebook</a> and trained it on the 5%, 2% and 1% datasets using an lr of 1e-3, weight decay of 5e-6, a batch size of 8, and BCE loss and (after some difficulty), got these results:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fe54661518e8f7ec8328c738262412b89%2FScreen%20Shot%202022-12-03%20at%2010.57.46%20AM.png?generation=1670083093142300&amp;alt=media\" alt=\"\"></p>\n<p>With the AUCs of:</p>\n<p>5% dataset: 0.872<br>\n2% dataset: 0.605<br>\n1% dataset: 0.5 (no learning)</p>\n<p>So what can we learn from that? Together with the CNN results, it's clear that signals as low as 1:20 SNR (5%) are really easy to detect - even with suboptimal data (I was only using 7.5 day signals remember).</p>\n<p>The 2% dataset only <em>just</em> started learning, even with the giant efficientnet, so that's what I'm thinking of as the lower bound of possible learning without a lot of tricks.</p>\n<p>The 1% dataset <em>still</em> didn't learn at all however - which tells me that signals that are 2 orders of magnitude weaker than the noise are going to be very, very tricky to detect. </p>\n<h2>Conclusions</h2>\n<p>So, how does that help us in this competition? Well, it tells me that there is some fundamental difference in the noise and signals at around 1:50 SNR since it becomes extremely difficult to detect. I think my next experiment might be to try to make different models for \"super easy / high SNR signals\", and \"very low SNR\" signals - it's possible that a single model can't be used to detect both kinds, since the high SNR signals seem to be very different than the low SNR ones.</p>\n<p>It also tells me that even simple CNN architectures can be helpful for high SNR signals - perhaps there is some room for simple models working along side more complex models to differentiate.</p>\n<p>Also, it tells me that if current gravitational wave detectors have theoretical signals at about 1:100 SNR then any work to reduce that noise floor could go a <em>long</em> way towards detecting waves, since there seems to be pretty hard cutoff in how easy signals are to detect at those levels.</p>\n<p>So - were all these experiments really helpful for the competition? I'm not actually sure 😂 but it was interesting for me to go on this exploration, and hopefully it prompts some interesting ideas for you.</p>",
  "messages": [
    {
      "id": 2053796,
      "postDate": "2022-12-03T16:08:53.897Z",
      "content": "<p>I took a step back to do some experiments to see how SNR affects the ability of networks (CNN and efficientnet) to detect signals. Here are some results (nothing earth shattering I think, but still interesting):</p>\n<h2>Step 1. Generate Waves</h2>\n<p>I modified the tutorial code in order to generate waves of varying strengths. I made 3 different simulated datasets, one at 5% signal to noise (SNR) (1:20), 2% (1:50) and 1% (1:100). The code to make those sets is here (I ran it 3 different times at 3 different h0 values): <a href=\"https://www.kaggle.com/chris62/generating-low-snr-gravity-waves\" target=\"_blank\">https://www.kaggle.com/chris62/generating-low-snr-gravity-waves</a></p>\n<p>Note that I only am making 7.5 day long signals, which result in a 360x360 image - this is much shorter than the full data in the competition (&gt; 4000 px wide), but was much easier to manage and use for my experiments. For my experiments below, I generated 1,000 train images and 1,000 validation images for each dataset, with 50% signals and 50% noise only.</p>\n<p>I also generated examples at 1:1 SNR, 1:10 and 1:100, just to see what the results looked like, and converted them into images using the \"power\" calculation for H1 and L1, and then subtracted the mean and divided by the standard deviation:</p>\n<pre><code> ():\n    data = np.load(filename)\n\n    img = np.zeros((,,IMGSIZE))\n\n    img[] = data[,:,:]** + data[,:,:]**\n    img[] = data[,:,:]** + data[,:,:]**\n\n    img = img - np.mean(img)\n    img = img / np.std(img)\n\n     img.transpose(,,)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fe3d51a2fa6ab78c425c43c7feef0a52a%2Fsnr.png?generation=1670081815216592&amp;alt=media\" alt=\"\"></p>\n<p>The 100% SNR signal totally overwhelms the noise, and after normalization, that's all you can see… how easy the competition would be if that was the case! 😆 </p>\n<p>Unfortunately, the competition signals are 1 to 2 orders of magnitude weaker than the noise, which means the signals will be somewhere between the 10% and 1% SNR image above. At 1% SNR you can't see the signal at all - so we have our work cut out for us! </p>\n<h2>Step 2. Simple CNN architecture</h2>\n<p>I wanted to test the 5%, 2% and 1% dataset on a non-pre-trained, simple architecture, so I created a CNN model that would take the 2 channel 360x360 images, and return a single value (1 for signal, 0 for noise):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fae8379131e73a6a2cabd575e24439106%2FScreen%20Shot%202022-12-03%20at%2010.42.06%20AM.png?generation=1670082226219825&amp;alt=media\" alt=\"\"></p>\n<p>The only really interesting part of that is the first convolution, which is 5x31. My theory was that a very wide first convolution would be helpful in detecting the mostly horizontal waves, and it seemed to work fairly well in my experiments.</p>\n<p>My first tests were with just 100 training and validation samples across a wide range of SNR values - from 1.0 all the way down to 0.01.  I trained that model for 25 epochs using the AdamW optimizer with a learning rate of 1e-4 using binary cross entropy with logits loss.</p>\n<p>From 1.0 down to 0.1 (100% SNR to 10%), the model was very easily to get close to 1.0 AUC, which is totally expected given the how easy it is to distinguish the signals above:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fb16259ff469c0e4bc397bc95679ed0ba%2FScreen%20Shot%202022-12-03%20at%2010.48.15%20AM.png?generation=1670082510716865&amp;alt=media\" alt=\"\"></p>\n<p>Then I tested from 10% down to 1% SNR (1:100), and the AUC started going down - crossing 0.5 between 1% and 2% (0.5 AUC means total guessing):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2F9a100c472a502e0aa56af8ac5a4b1474%2FScreen%20Shot%202022-12-03%20at%2010.48.55%20AM.png?generation=1670082576318418&amp;alt=media\" alt=\"\"></p>\n<p>So from that we can take that 1. there is some hope, since AUC was &gt; 0.5 all the way down to 2%, BUT it obviously becomes increasingly difficult to detect signals at lower and lower SNRs. (Again, not earth shattering, but still interesting I think).</p>\n<p>The next step was to try the CNN on my 5%, 2% and 1% datasets, but I had a ton of trouble actually training - the network continuously overfit:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2F23f584f3d2201cec84bd85b1271478cb%2FScreen%20Shot%202022-12-03%20at%2010.50.55%20AM.png?generation=1670082698310513&amp;alt=media\" alt=\"\"></p>\n<p>Eventually I was able to get it to train by adding 0.3 dropout and data augmentations (same as <a href=\"https://www.kaggle.com/code/leolu1998/g2net-basic-audio-data-augmentation-inference\" target=\"_blank\">this notebook</a> ) but ONLY got good results for the 5% dataset!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fa02b81a9324d43e2399c3adea20356a1%2FScreen%20Shot%202022-12-03%20at%2010.54.02%20AM.png?generation=1670082857103133&amp;alt=media\" alt=\"\"></p>\n<p>I was finally able to achieve a validation loss of 0.344 with an AUC of 0.814 for the 1:20 (5% SNR) dataset. That 0.814 looks good - but! since it's only on the 5% dataset, it's actually not as good as it seemed, and I wasn't able to get the 2% or 1% dataset to really train at all.</p>\n<p>Learning from this step: Getting all the way to 1:100 SNR wasn't possible for me with a simple architecture. Time to bring out the big network!</p>\n<h2>Step 3. Pre-trained efficientnet</h2>\n<p>Next I switch to a tf_efficientnet_b7_ns pretrained backbone, with a custom classifier head similar to <a href=\"https://www.kaggle.com/code/leolu1998/g2net-basic-audio-data-augmentation-inference\" target=\"_blank\">this notebook</a> and trained it on the 5%, 2% and 1% datasets using an lr of 1e-3, weight decay of 5e-6, a batch size of 8, and BCE loss and (after some difficulty), got these results:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fe54661518e8f7ec8328c738262412b89%2FScreen%20Shot%202022-12-03%20at%2010.57.46%20AM.png?generation=1670083093142300&amp;alt=media\" alt=\"\"></p>\n<p>With the AUCs of:</p>\n<p>5% dataset: 0.872<br>\n2% dataset: 0.605<br>\n1% dataset: 0.5 (no learning)</p>\n<p>So what can we learn from that? Together with the CNN results, it's clear that signals as low as 1:20 SNR (5%) are really easy to detect - even with suboptimal data (I was only using 7.5 day signals remember).</p>\n<p>The 2% dataset only <em>just</em> started learning, even with the giant efficientnet, so that's what I'm thinking of as the lower bound of possible learning without a lot of tricks.</p>\n<p>The 1% dataset <em>still</em> didn't learn at all however - which tells me that signals that are 2 orders of magnitude weaker than the noise are going to be very, very tricky to detect. </p>\n<h2>Conclusions</h2>\n<p>So, how does that help us in this competition? Well, it tells me that there is some fundamental difference in the noise and signals at around 1:50 SNR since it becomes extremely difficult to detect. I think my next experiment might be to try to make different models for \"super easy / high SNR signals\", and \"very low SNR\" signals - it's possible that a single model can't be used to detect both kinds, since the high SNR signals seem to be very different than the low SNR ones.</p>\n<p>It also tells me that even simple CNN architectures can be helpful for high SNR signals - perhaps there is some room for simple models working along side more complex models to differentiate.</p>\n<p>Also, it tells me that if current gravitational wave detectors have theoretical signals at about 1:100 SNR then any work to reduce that noise floor could go a <em>long</em> way towards detecting waves, since there seems to be pretty hard cutoff in how easy signals are to detect at those levels.</p>\n<p>So - were all these experiments really helpful for the competition? I'm not actually sure 😂 but it was interesting for me to go on this exploration, and hopefully it prompts some interesting ideas for you.</p>",
      "rawMarkdown": "I took a step back to do some experiments to see how SNR affects the ability of networks (CNN and efficientnet) to detect signals. Here are some results (nothing earth shattering I think, but still interesting):\n\n## Step 1. Generate Waves\n\nI modified the tutorial code in order to generate waves of varying strengths. I made 3 different simulated datasets, one at 5% signal to noise (SNR) (1:20), 2% (1:50) and 1% (1:100). The code to make those sets is here (I ran it 3 different times at 3 different h0 values): [https://www.kaggle.com/chris62/generating-low-snr-gravity-waves](https://www.kaggle.com/chris62/generating-low-snr-gravity-waves)\n\nNote that I only am making 7.5 day long signals, which result in a 360x360 image - this is much shorter than the full data in the competition (> 4000 px wide), but was much easier to manage and use for my experiments. For my experiments below, I generated 1,000 train images and 1,000 validation images for each dataset, with 50% signals and 50% noise only.\n\nI also generated examples at 1:1 SNR, 1:10 and 1:100, just to see what the results looked like, and converted them into images using the \"power\" calculation for H1 and L1, and then subtracted the mean and divided by the standard deviation:\n\n```python\ndef npy_to_img(filename):\n    data = np.load(filename)\n\n    img = np.zeros((3,360,IMGSIZE))\n\n    img[1] = data[0,:,:]**2 + data[1,:,:]**2\n    img[2] = data[2,:,:]**2 + data[3,:,:]**2\n\n    img = img - np.mean(img)\n    img = img / np.std(img)\n    \n    return img.transpose(1,2,0)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fe3d51a2fa6ab78c425c43c7feef0a52a%2Fsnr.png?generation=1670081815216592&alt=media)\n\nThe 100% SNR signal totally overwhelms the noise, and after normalization, that's all you can see... how easy the competition would be if that was the case! 😆 \n\nUnfortunately, the competition signals are 1 to 2 orders of magnitude weaker than the noise, which means the signals will be somewhere between the 10% and 1% SNR image above. At 1% SNR you can't see the signal at all - so we have our work cut out for us! \n\n## Step 2. Simple CNN architecture\n\nI wanted to test the 5%, 2% and 1% dataset on a non-pre-trained, simple architecture, so I created a CNN model that would take the 2 channel 360x360 images, and return a single value (1 for signal, 0 for noise):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fae8379131e73a6a2cabd575e24439106%2FScreen%20Shot%202022-12-03%20at%2010.42.06%20AM.png?generation=1670082226219825&alt=media)\n\nThe only really interesting part of that is the first convolution, which is 5x31. My theory was that a very wide first convolution would be helpful in detecting the mostly horizontal waves, and it seemed to work fairly well in my experiments.\n\nMy first tests were with just 100 training and validation samples across a wide range of SNR values - from 1.0 all the way down to 0.01.  I trained that model for 25 epochs using the AdamW optimizer with a learning rate of 1e-4 using binary cross entropy with logits loss.\n\nFrom 1.0 down to 0.1 (100% SNR to 10%), the model was very easily to get close to 1.0 AUC, which is totally expected given the how easy it is to distinguish the signals above:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fb16259ff469c0e4bc397bc95679ed0ba%2FScreen%20Shot%202022-12-03%20at%2010.48.15%20AM.png?generation=1670082510716865&alt=media)\n\nThen I tested from 10% down to 1% SNR (1:100), and the AUC started going down - crossing 0.5 between 1% and 2% (0.5 AUC means total guessing):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2F9a100c472a502e0aa56af8ac5a4b1474%2FScreen%20Shot%202022-12-03%20at%2010.48.55%20AM.png?generation=1670082576318418&alt=media)\n\nSo from that we can take that 1. there is some hope, since AUC was > 0.5 all the way down to 2%, BUT it obviously becomes increasingly difficult to detect signals at lower and lower SNRs. (Again, not earth shattering, but still interesting I think).\n\nThe next step was to try the CNN on my 5%, 2% and 1% datasets, but I had a ton of trouble actually training - the network continuously overfit:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2F23f584f3d2201cec84bd85b1271478cb%2FScreen%20Shot%202022-12-03%20at%2010.50.55%20AM.png?generation=1670082698310513&alt=media)\n\nEventually I was able to get it to train by adding 0.3 dropout and data augmentations (same as [this notebook](https://www.kaggle.com/code/leolu1998/g2net-basic-audio-data-augmentation-inference) ) but ONLY got good results for the 5% dataset!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fa02b81a9324d43e2399c3adea20356a1%2FScreen%20Shot%202022-12-03%20at%2010.54.02%20AM.png?generation=1670082857103133&alt=media)\n\nI was finally able to achieve a validation loss of 0.344 with an AUC of 0.814 for the 1:20 (5% SNR) dataset. That 0.814 looks good - but! since it's only on the 5% dataset, it's actually not as good as it seemed, and I wasn't able to get the 2% or 1% dataset to really train at all.\n\nLearning from this step: Getting all the way to 1:100 SNR wasn't possible for me with a simple architecture. Time to bring out the big network!\n\n## Step 3. Pre-trained efficientnet\n\nNext I switch to a tf_efficientnet_b7_ns pretrained backbone, with a custom classifier head similar to [this notebook](https://www.kaggle.com/code/leolu1998/g2net-basic-audio-data-augmentation-inference) and trained it on the 5%, 2% and 1% datasets using an lr of 1e-3, weight decay of 5e-6, a batch size of 8, and BCE loss and (after some difficulty), got these results:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fe54661518e8f7ec8328c738262412b89%2FScreen%20Shot%202022-12-03%20at%2010.57.46%20AM.png?generation=1670083093142300&alt=media)\n\nWith the AUCs of:\n\n5% dataset: 0.872\n2% dataset: 0.605\n1% dataset: 0.5 (no learning)\n\nSo what can we learn from that? Together with the CNN results, it's clear that signals as low as 1:20 SNR (5%) are really easy to detect - even with suboptimal data (I was only using 7.5 day signals remember).\n\nThe 2% dataset only _just_ started learning, even with the giant efficientnet, so that's what I'm thinking of as the lower bound of possible learning without a lot of tricks.\n\nThe 1% dataset _still_ didn't learn at all however - which tells me that signals that are 2 orders of magnitude weaker than the noise are going to be very, very tricky to detect. \n\n## Conclusions\n\nSo, how does that help us in this competition? Well, it tells me that there is some fundamental difference in the noise and signals at around 1:50 SNR since it becomes extremely difficult to detect. I think my next experiment might be to try to make different models for \"super easy / high SNR signals\", and \"very low SNR\" signals - it's possible that a single model can't be used to detect both kinds, since the high SNR signals seem to be very different than the low SNR ones.\n\nIt also tells me that even simple CNN architectures can be helpful for high SNR signals - perhaps there is some room for simple models working along side more complex models to differentiate.\n\nAlso, it tells me that if current gravitational wave detectors have theoretical signals at about 1:100 SNR then any work to reduce that noise floor could go a _long_ way towards detecting waves, since there seems to be pretty hard cutoff in how easy signals are to detect at those levels.\n\nSo - were all these experiments really helpful for the competition? I'm not actually sure 😂 but it was interesting for me to go on this exploration, and hopefully it prompts some interesting ideas for you.\n\n",
      "votes": 55
    },
    {
      "id": 2059290,
      "postDate": "2022-12-08T17:13:33.587Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> , this is really cool! Thanks for sharing.</p>\n<p>Question: What do you mean by \"SNR\"? Do you mean sqrtSX / h0 ? <br>\nShould I read \"1:100 SNR\" as <code>sqrtSX / h0 = 100</code>?</p>",
      "rawMarkdown": "Hi @chris62 , this is really cool! Thanks for sharing.\n\nQuestion: What do you mean by \"SNR\"? Do you mean sqrtSX / h0 ? \nShould I read \"1:100 SNR\" as `sqrtSX / h0 = 100`?",
      "votes": 1,
      "replies": [
        {
          "id": 2059297,
          "postDate": "2022-12-08T17:25:50.470Z",
          "content": "<p>Yeah, I am a bit careless with terminology here; I mean the ratio (or as a percent) of the sqrtSX compared to the h0 param - so a signal with sqrtSX / 100 is what I'm calling 1:100 (or 1%); it's not perfect (because things also change with f0 and f1, etc), but it seemed to still lead to interesting results. Does that answer your question?</p>",
          "rawMarkdown": "Yeah, I am a bit careless with terminology here; I mean the ratio (or as a percent) of the sqrtSX compared to the h0 param - so a signal with sqrtSX / 100 is what I'm calling 1:100 (or 1%); it's not perfect (because things also change with f0 and f1, etc), but it seemed to still lead to interesting results. Does that answer your question?",
          "votes": 1
        },
        {
          "id": 2059318,
          "postDate": "2022-12-08T17:49:59.610Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> ,</p>\n<p>Yup, it does. In some cases (this may be one of them) I like to think in practically the same terms.<br>\nThe problem is that these units don't take <code>cosi</code> into account, which may make things vary by an<br>\norder of magnitude or so, but as long as one is aware of this, it should be Ok.</p>\n<p>Thanks!</p>",
          "rawMarkdown": "Hi @chris62 ,\n\nYup, it does. In some cases (this may be one of them) I like to think in practically the same terms.\nThe problem is that these units don't take `cosi` into account, which may make things vary by an\norder of magnitude or so, but as long as one is aware of this, it should be Ok.\n\nThanks!",
          "votes": 2
        },
        {
          "id": 2059321,
          "postDate": "2022-12-08T17:52:45.567Z",
          "content": "<p>Yeah; I actually didn't realize that when I put this together, but after more experiments I have come to think that I should have measured the actual amplitude of the resulting signal compared to the mean of the noise, and that would have been better probably :) </p>",
          "rawMarkdown": "Yeah; I actually didn't realize that when I put this together, but after more experiments I have come to think that I should have measured the actual amplitude of the resulting signal compared to the mean of the noise, and that would have been better probably :) ",
          "votes": 1
        },
        {
          "id": 2060570,
          "postDate": "2022-12-10T06:01:23.340Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rodrigotenorio\" target=\"_blank\">@rodrigotenorio</a> , how does <code>cosi</code> influence  the magnitude. I can't find the formula of <code>cosi</code> and magnitude or something else in your notebook. All I know is <code>cosi</code> is \"inclination angle of the source\". </p>",
          "rawMarkdown": "Hi @rodrigotenorio , how does `cosi` influence  the magnitude. I can't find the formula of `cosi` and magnitude or something else in your notebook. All I know is `cosi` is \"inclination angle of the source\". "
        },
        {
          "id": 2060718,
          "postDate": "2022-12-10T10:42:55.313Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> ,</p>\n<p>You can have a look at equation 95 of <a href=\"https://dcc.ligo.org/LIGO-T0900149\" target=\"_blank\">this document</a> where the specific dependency of SNR and cosi is discussed (in this case <code>cosi</code> is denoted as <code>eta</code>).</p>\n<p>In a nutshell, <code>cosi=+-1</code> gives you a stronger signal, while <code>cosi=0</code> gives you a weaker signal. The specific factor is <code>cosi^4 + 6 * cosi^2 + 1</code> and that essentially means around a factor 8 of difference between both extreme cases.</p>",
          "rawMarkdown": "Hi @forcewithme ,\n\nYou can have a look at equation 95 of [this document](https://dcc.ligo.org/LIGO-T0900149) where the specific dependency of SNR and cosi is discussed (in this case `cosi` is denoted as `eta`).\n\nIn a nutshell, `cosi=+-1` gives you a stronger signal, while `cosi=0` gives you a weaker signal. The specific factor is `cosi^4 + 6 * cosi^2 + 1` and that essentially means around a factor 8 of difference between both extreme cases.",
          "votes": 4
        },
        {
          "id": 2060885,
          "postDate": "2022-12-10T14:30:49.147Z",
          "content": "<p>Amazing, thank you for clarification! <a href=\"https://www.kaggle.com/rodrigotenorio\" target=\"_blank\">@rodrigotenorio</a> </p>",
          "rawMarkdown": "Amazing, thank you for clarification! @rodrigotenorio ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2056024,
      "postDate": "2022-12-05T16:51:40.247Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> , do you use generated data in your current LB solution?</p>",
      "rawMarkdown": "Hi @chris62 , do you use generated data in your current LB solution?",
      "replies": [
        {
          "id": 2056040,
          "postDate": "2022-12-05T17:10:57.850Z",
          "content": "<p>Yes; if it's possible to get that high without generated data then I'm not sure how 😜</p>\n<p>Have you used generated data yet?</p>",
          "rawMarkdown": "Yes; if it's possible to get that high without generated data then I'm not sure how 😜\n\nHave you used generated data yet?",
          "votes": 1
        },
        {
          "id": 2056334,
          "postDate": "2022-12-06T02:50:06.203Z",
          "content": "<p>No, I haven't made my generated data work.☹️</p>",
          "rawMarkdown": "No, I haven't made my generated data work.☹️",
          "votes": 2
        }
      ]
    },
    {
      "id": 2061427,
      "postDate": "2022-12-11T06:06:06.757Z",
      "content": "<p>Thanks for your sharing</p>",
      "rawMarkdown": "Thanks for your sharing"
    }
  ],
  "comments": [
    {
      "id": 2059290,
      "author_name": "Rodrigo Tenorio",
      "author_url": "",
      "post_date": "2022-12-08T17:13:33.587000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> , this is really cool! Thanks for sharing.</p>\n<p>Question: What do you mean by \"SNR\"? Do you mean sqrtSX / h0 ? <br>\nShould I read \"1:100 SNR\" as <code>sqrtSX / h0 = 100</code>?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2059297,
          "author_name": "chris",
          "author_url": "",
          "post_date": "2022-12-08T17:25:50.470000",
          "content": "<p>Yeah, I am a bit careless with terminology here; I mean the ratio (or as a percent) of the sqrtSX compared to the h0 param - so a signal with sqrtSX / 100 is what I'm calling 1:100 (or 1%); it's not perfect (because things also change with f0 and f1, etc), but it seemed to still lead to interesting results. Does that answer your question?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2059318,
          "author_name": "Rodrigo Tenorio",
          "author_url": "",
          "post_date": "2022-12-08T17:49:59.610000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> ,</p>\n<p>Yup, it does. In some cases (this may be one of them) I like to think in practically the same terms.<br>\nThe problem is that these units don't take <code>cosi</code> into account, which may make things vary by an<br>\norder of magnitude or so, but as long as one is aware of this, it should be Ok.</p>\n<p>Thanks!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2059321,
          "author_name": "chris",
          "author_url": "",
          "post_date": "2022-12-08T17:52:45.567000",
          "content": "<p>Yeah; I actually didn't realize that when I put this together, but after more experiments I have come to think that I should have measured the actual amplitude of the resulting signal compared to the mean of the noise, and that would have been better probably :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2060570,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2022-12-10T06:01:23.340000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rodrigotenorio\" target=\"_blank\">@rodrigotenorio</a> , how does <code>cosi</code> influence  the magnitude. I can't find the formula of <code>cosi</code> and magnitude or something else in your notebook. All I know is <code>cosi</code> is \"inclination angle of the source\". </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2060718,
          "author_name": "Rodrigo Tenorio",
          "author_url": "",
          "post_date": "2022-12-10T10:42:55.313000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a> ,</p>\n<p>You can have a look at equation 95 of <a href=\"https://dcc.ligo.org/LIGO-T0900149\" target=\"_blank\">this document</a> where the specific dependency of SNR and cosi is discussed (in this case <code>cosi</code> is denoted as <code>eta</code>).</p>\n<p>In a nutshell, <code>cosi=+-1</code> gives you a stronger signal, while <code>cosi=0</code> gives you a weaker signal. The specific factor is <code>cosi^4 + 6 * cosi^2 + 1</code> and that essentially means around a factor 8 of difference between both extreme cases.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2060885,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2022-12-10T14:30:49.147000",
          "content": "<p>Amazing, thank you for clarification! <a href=\"https://www.kaggle.com/rodrigotenorio\" target=\"_blank\">@rodrigotenorio</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2056024,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2022-12-05T16:51:40.247000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chris62\" target=\"_blank\">@chris62</a> , do you use generated data in your current LB solution?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2056040,
          "author_name": "chris",
          "author_url": "",
          "post_date": "2022-12-05T17:10:57.850000",
          "content": "<p>Yes; if it's possible to get that high without generated data then I'm not sure how 😜</p>\n<p>Have you used generated data yet?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2056334,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2022-12-06T02:50:06.203000",
          "content": "<p>No, I haven't made my generated data work.☹️</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2061427,
      "author_name": "bill",
      "author_url": "",
      "post_date": "2022-12-11T06:06:06.757000",
      "content": "<p>Thanks for your sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2053796": "I took a step back to do some experiments to see how SNR affects the ability of networks (CNN and efficientnet) to detect signals. Here are some results (nothing earth shattering I think, but still interesting):\n\n## Step 1. Generate Waves\n\nI modified the tutorial code in order to generate waves of varying strengths. I made 3 different simulated datasets, one at 5% signal to noise (SNR) (1:20), 2% (1:50) and 1% (1:100). The code to make those sets is here (I ran it 3 different times at 3 different h0 values): [https://www.kaggle.com/chris62/generating-low-snr-gravity-waves](https://www.kaggle.com/chris62/generating-low-snr-gravity-waves)\n\nNote that I only am making 7.5 day long signals, which result in a 360x360 image - this is much shorter than the full data in the competition (> 4000 px wide), but was much easier to manage and use for my experiments. For my experiments below, I generated 1,000 train images and 1,000 validation images for each dataset, with 50% signals and 50% noise only.\n\nI also generated examples at 1:1 SNR, 1:10 and 1:100, just to see what the results looked like, and converted them into images using the \"power\" calculation for H1 and L1, and then subtracted the mean and divided by the standard deviation:\n\n```python\ndef npy_to_img(filename):\n    data = np.load(filename)\n\n    img = np.zeros((3,360,IMGSIZE))\n\n    img[1] = data[0,:,:]**2 + data[1,:,:]**2\n    img[2] = data[2,:,:]**2 + data[3,:,:]**2\n\n    img = img - np.mean(img)\n    img = img / np.std(img)\n    \n    return img.transpose(1,2,0)\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fe3d51a2fa6ab78c425c43c7feef0a52a%2Fsnr.png?generation=1670081815216592&alt=media)\n\nThe 100% SNR signal totally overwhelms the noise, and after normalization, that's all you can see... how easy the competition would be if that was the case! 😆 \n\nUnfortunately, the competition signals are 1 to 2 orders of magnitude weaker than the noise, which means the signals will be somewhere between the 10% and 1% SNR image above. At 1% SNR you can't see the signal at all - so we have our work cut out for us! \n\n## Step 2. Simple CNN architecture\n\nI wanted to test the 5%, 2% and 1% dataset on a non-pre-trained, simple architecture, so I created a CNN model that would take the 2 channel 360x360 images, and return a single value (1 for signal, 0 for noise):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fae8379131e73a6a2cabd575e24439106%2FScreen%20Shot%202022-12-03%20at%2010.42.06%20AM.png?generation=1670082226219825&alt=media)\n\nThe only really interesting part of that is the first convolution, which is 5x31. My theory was that a very wide first convolution would be helpful in detecting the mostly horizontal waves, and it seemed to work fairly well in my experiments.\n\nMy first tests were with just 100 training and validation samples across a wide range of SNR values - from 1.0 all the way down to 0.01.  I trained that model for 25 epochs using the AdamW optimizer with a learning rate of 1e-4 using binary cross entropy with logits loss.\n\nFrom 1.0 down to 0.1 (100% SNR to 10%), the model was very easily to get close to 1.0 AUC, which is totally expected given the how easy it is to distinguish the signals above:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fb16259ff469c0e4bc397bc95679ed0ba%2FScreen%20Shot%202022-12-03%20at%2010.48.15%20AM.png?generation=1670082510716865&alt=media)\n\nThen I tested from 10% down to 1% SNR (1:100), and the AUC started going down - crossing 0.5 between 1% and 2% (0.5 AUC means total guessing):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2F9a100c472a502e0aa56af8ac5a4b1474%2FScreen%20Shot%202022-12-03%20at%2010.48.55%20AM.png?generation=1670082576318418&alt=media)\n\nSo from that we can take that 1. there is some hope, since AUC was > 0.5 all the way down to 2%, BUT it obviously becomes increasingly difficult to detect signals at lower and lower SNRs. (Again, not earth shattering, but still interesting I think).\n\nThe next step was to try the CNN on my 5%, 2% and 1% datasets, but I had a ton of trouble actually training - the network continuously overfit:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2F23f584f3d2201cec84bd85b1271478cb%2FScreen%20Shot%202022-12-03%20at%2010.50.55%20AM.png?generation=1670082698310513&alt=media)\n\nEventually I was able to get it to train by adding 0.3 dropout and data augmentations (same as [this notebook](https://www.kaggle.com/code/leolu1998/g2net-basic-audio-data-augmentation-inference) ) but ONLY got good results for the 5% dataset!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fa02b81a9324d43e2399c3adea20356a1%2FScreen%20Shot%202022-12-03%20at%2010.54.02%20AM.png?generation=1670082857103133&alt=media)\n\nI was finally able to achieve a validation loss of 0.344 with an AUC of 0.814 for the 1:20 (5% SNR) dataset. That 0.814 looks good - but! since it's only on the 5% dataset, it's actually not as good as it seemed, and I wasn't able to get the 2% or 1% dataset to really train at all.\n\nLearning from this step: Getting all the way to 1:100 SNR wasn't possible for me with a simple architecture. Time to bring out the big network!\n\n## Step 3. Pre-trained efficientnet\n\nNext I switch to a tf_efficientnet_b7_ns pretrained backbone, with a custom classifier head similar to [this notebook](https://www.kaggle.com/code/leolu1998/g2net-basic-audio-data-augmentation-inference) and trained it on the 5%, 2% and 1% datasets using an lr of 1e-3, weight decay of 5e-6, a batch size of 8, and BCE loss and (after some difficulty), got these results:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F40584%2Fe54661518e8f7ec8328c738262412b89%2FScreen%20Shot%202022-12-03%20at%2010.57.46%20AM.png?generation=1670083093142300&alt=media)\n\nWith the AUCs of:\n\n5% dataset: 0.872\n2% dataset: 0.605\n1% dataset: 0.5 (no learning)\n\nSo what can we learn from that? Together with the CNN results, it's clear that signals as low as 1:20 SNR (5%) are really easy to detect - even with suboptimal data (I was only using 7.5 day signals remember).\n\nThe 2% dataset only _just_ started learning, even with the giant efficientnet, so that's what I'm thinking of as the lower bound of possible learning without a lot of tricks.\n\nThe 1% dataset _still_ didn't learn at all however - which tells me that signals that are 2 orders of magnitude weaker than the noise are going to be very, very tricky to detect. \n\n## Conclusions\n\nSo, how does that help us in this competition? Well, it tells me that there is some fundamental difference in the noise and signals at around 1:50 SNR since it becomes extremely difficult to detect. I think my next experiment might be to try to make different models for \"super easy / high SNR signals\", and \"very low SNR\" signals - it's possible that a single model can't be used to detect both kinds, since the high SNR signals seem to be very different than the low SNR ones.\n\nIt also tells me that even simple CNN architectures can be helpful for high SNR signals - perhaps there is some room for simple models working along side more complex models to differentiate.\n\nAlso, it tells me that if current gravitational wave detectors have theoretical signals at about 1:100 SNR then any work to reduce that noise floor could go a _long_ way towards detecting waves, since there seems to be pretty hard cutoff in how easy signals are to detect at those levels.\n\nSo - were all these experiments really helpful for the competition? I'm not actually sure 😂 but it was interesting for me to go on this exploration, and hopefully it prompts some interesting ideas for you.\n\n",
    "2059290": "Hi @chris62 , this is really cool! Thanks for sharing.\n\nQuestion: What do you mean by \"SNR\"? Do you mean sqrtSX / h0 ? \nShould I read \"1:100 SNR\" as `sqrtSX / h0 = 100`?",
    "2056024": "Hi @chris62 , do you use generated data in your current LB solution?",
    "2061427": "Thanks for your sharing"
  }
}