{
  "id": 375957,
  "title": "20th place solution (how I spent lots of time on things that didn't work)",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/375957",
  "author_name": "Jonathan McKinney",
  "post_date": "2023-01-04T07:06:27.866000",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thanks to organizers and making a fun physics-based competition.</p>\n<p>I'm very impressed by the creativity (and simplicity) of the solutions (so far) posted from <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> and <a href=\"https://www.kaggle.com/shunrcn\" target=\"_blank\">@shunrcn</a>, I'm sure there are more to come.</p>\n<p>Also thanks to the kaggle community for being such great competitors.</p>\n<h1>Solution:</h1>\n<h2>Train</h2>\n<ul>\n<li>pyfstat to generate 20k signals with test-data-based gaps</li>\n<li>Same portion of real vs. fake data as in test set (fake being roughly constant standard deviation)</li>\n<li>All real type data generated by sampling from test data (e.g.start times, durations, standard deviations vs. time)</li>\n<li>Gaussian random complex noise added to complex signal, as well as direct real noise to absolute magnitude.</li>\n<li>Ignore h0, just add signal to noise in normalized way based upon total power in signal vs. noise along signal.  Signal power went from just barely human visible at 360x360 to 10x lower.</li>\n<li>About 200k training samples, Adam, Cosine scheduler, batch=32*4</li>\n</ul>\n<h2>Train and Test</h2>\n<ul>\n<li>Fill-out time data to 5760, binning data to be uniformly spaced</li>\n<li>Cleaned data by removing all line or other spot noise from test data (quite effectively and fast)</li>\n<li>Fill-in gaps of all data with surrounding level of noise</li>\n<li>Compute absolute value as well as H1 times complex conjugate of L1 (equiv of cross-correction in FFT space)</li>\n<li>Reduced by np.mean (like avgpool) down to 360x360 to help reduce noise</li>\n<li>Normalize</li>\n</ul>\n<h2>Model</h2>\n<ul>\n<li>tf_efficientnet_b5_ns, 20-30 epochs, 5 folds</li>\n<li>No Aug</li>\n<li>Mixup (was required to avoid overfitting, is that what others saw?)</li>\n<li>TTA of various seeds for filling-in the noise between gaps (worried model would overfit on clusters of noise)</li>\n<li>Blended best single model, 15 5-fold TTAs of that best, large kernel</li>\n<li>Nicely, CV score matched public very well.</li>\n</ul>\n<h1>Things I wrongly disregarded as not helpful</h1>\n<ul>\n<li>Normalizing per time point.  I disregarded it too quickly, thinking the signal would vary in strength and that would make things harder.  I was simply wrong to think that.  Thanks <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> </li>\n</ul>\n<h1>Things that failed or didn't help:</h1>\n<ul>\n<li>Treating as regression on power fraction (didn't help)</li>\n<li>replknet large kernel model (didn't help)</li>\n<li>Larger image sizes, 360x720 and 360x1440 (didn't help)</li>\n<li>Denoising VAE with noise+signal given, signal as goal (failed totally to learn, even at small resolutions)</li>\n<li>Stable diffusion (VAE they trained fails totally on our noise level)</li>\n<li>resnet18 (failed to work at all for me)</li>\n<li>mixnet_m (does ok)</li>\n<li>mixnet_l (does ok)</li>\n<li>convmixer_768_32 (does ok, but slow)</li>\n<li>vit_base_patch16_224 and vit_relpos_small_patch16_224 (though ViT would pick up on sequence of patches, but failed totally on me, probably bug in usage)</li>\n<li>large_kernel (<a href=\"https://www.kaggle.com/code/laeyoung/g2net-large-kernel-inference/data\" target=\"_blank\">https://www.kaggle.com/code/laeyoung/g2net-large-kernel-inference/data</a>) but dropping avgpool since not needed.  Tried various kernel sizes, nothing really helped. Ok model but weaker than efficientnet</li>\n<li>Training on fake and real data separately (only did bit worse)</li>\n<li>Adding noise as Aug (didn't help)</li>\n<li>BiLSTM + Conv approach like done for BH-BH case and other competitions (isft -&gt; sft is not perfect, and even matching pyfstat sft, the human-visible signals get easily lost by the isft process, so time series approach would likely fail, so gave up)</li>\n<li>Downloading real data to find segments that match in real test data, as <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">@vslaykovsky</a> said.  I did that, but quickly realized the data did not match for gaps etc.  So disregarded it.</li>\n<li>Use pyfstat signal detector itself, both MCMC and grid search.  I had to hack the code to convert hd5f -&gt; sft.  Does very poorly even on easy signals when range of search is even a tiny bigger than actual signal.  So failed to work.  </li>\n<li>Look at test set overlap.  Form bounding boxes for every test data set, and find matches.  There are about 1000 matches, with varying overlaps.  I used those to remove the noise.  Clean-up standard deviation and mean at interfaces left over.  But didn't help, probably because subtractions made data less clean.</li>\n</ul>\n<h1>What I don't understand:</h1>\n<p>1) I always got boost from blending with good amount of top kernel with 0.761 that was partially based upon large kernel.  I couldn't get independently high enough.  My blend used half my stuff and half of that kernel in the end that did best on private.  The large kernel model, part of that 0.761 score, seems to be doing better on fake data for weaker signals.  But I never could reproduce that with own large kernel or replknet models.<br>\n2) Why did VAE fail so bad?  I never got out of AUC~0.5 domain.</p>",
  "messages": [
    {
      "id": 2085453,
      "postDate": "2023-01-04T07:06:27.867Z",
      "content": "<p>Thanks to organizers and making a fun physics-based competition.</p>\n<p>I'm very impressed by the creativity (and simplicity) of the solutions (so far) posted from <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> and <a href=\"https://www.kaggle.com/shunrcn\" target=\"_blank\">@shunrcn</a>, I'm sure there are more to come.</p>\n<p>Also thanks to the kaggle community for being such great competitors.</p>\n<h1>Solution:</h1>\n<h2>Train</h2>\n<ul>\n<li>pyfstat to generate 20k signals with test-data-based gaps</li>\n<li>Same portion of real vs. fake data as in test set (fake being roughly constant standard deviation)</li>\n<li>All real type data generated by sampling from test data (e.g.start times, durations, standard deviations vs. time)</li>\n<li>Gaussian random complex noise added to complex signal, as well as direct real noise to absolute magnitude.</li>\n<li>Ignore h0, just add signal to noise in normalized way based upon total power in signal vs. noise along signal.  Signal power went from just barely human visible at 360x360 to 10x lower.</li>\n<li>About 200k training samples, Adam, Cosine scheduler, batch=32*4</li>\n</ul>\n<h2>Train and Test</h2>\n<ul>\n<li>Fill-out time data to 5760, binning data to be uniformly spaced</li>\n<li>Cleaned data by removing all line or other spot noise from test data (quite effectively and fast)</li>\n<li>Fill-in gaps of all data with surrounding level of noise</li>\n<li>Compute absolute value as well as H1 times complex conjugate of L1 (equiv of cross-correction in FFT space)</li>\n<li>Reduced by np.mean (like avgpool) down to 360x360 to help reduce noise</li>\n<li>Normalize</li>\n</ul>\n<h2>Model</h2>\n<ul>\n<li>tf_efficientnet_b5_ns, 20-30 epochs, 5 folds</li>\n<li>No Aug</li>\n<li>Mixup (was required to avoid overfitting, is that what others saw?)</li>\n<li>TTA of various seeds for filling-in the noise between gaps (worried model would overfit on clusters of noise)</li>\n<li>Blended best single model, 15 5-fold TTAs of that best, large kernel</li>\n<li>Nicely, CV score matched public very well.</li>\n</ul>\n<h1>Things I wrongly disregarded as not helpful</h1>\n<ul>\n<li>Normalizing per time point.  I disregarded it too quickly, thinking the signal would vary in strength and that would make things harder.  I was simply wrong to think that.  Thanks <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> </li>\n</ul>\n<h1>Things that failed or didn't help:</h1>\n<ul>\n<li>Treating as regression on power fraction (didn't help)</li>\n<li>replknet large kernel model (didn't help)</li>\n<li>Larger image sizes, 360x720 and 360x1440 (didn't help)</li>\n<li>Denoising VAE with noise+signal given, signal as goal (failed totally to learn, even at small resolutions)</li>\n<li>Stable diffusion (VAE they trained fails totally on our noise level)</li>\n<li>resnet18 (failed to work at all for me)</li>\n<li>mixnet_m (does ok)</li>\n<li>mixnet_l (does ok)</li>\n<li>convmixer_768_32 (does ok, but slow)</li>\n<li>vit_base_patch16_224 and vit_relpos_small_patch16_224 (though ViT would pick up on sequence of patches, but failed totally on me, probably bug in usage)</li>\n<li>large_kernel (<a href=\"https://www.kaggle.com/code/laeyoung/g2net-large-kernel-inference/data\" target=\"_blank\">https://www.kaggle.com/code/laeyoung/g2net-large-kernel-inference/data</a>) but dropping avgpool since not needed.  Tried various kernel sizes, nothing really helped. Ok model but weaker than efficientnet</li>\n<li>Training on fake and real data separately (only did bit worse)</li>\n<li>Adding noise as Aug (didn't help)</li>\n<li>BiLSTM + Conv approach like done for BH-BH case and other competitions (isft -&gt; sft is not perfect, and even matching pyfstat sft, the human-visible signals get easily lost by the isft process, so time series approach would likely fail, so gave up)</li>\n<li>Downloading real data to find segments that match in real test data, as <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">@vslaykovsky</a> said.  I did that, but quickly realized the data did not match for gaps etc.  So disregarded it.</li>\n<li>Use pyfstat signal detector itself, both MCMC and grid search.  I had to hack the code to convert hd5f -&gt; sft.  Does very poorly even on easy signals when range of search is even a tiny bigger than actual signal.  So failed to work.  </li>\n<li>Look at test set overlap.  Form bounding boxes for every test data set, and find matches.  There are about 1000 matches, with varying overlaps.  I used those to remove the noise.  Clean-up standard deviation and mean at interfaces left over.  But didn't help, probably because subtractions made data less clean.</li>\n</ul>\n<h1>What I don't understand:</h1>\n<p>1) I always got boost from blending with good amount of top kernel with 0.761 that was partially based upon large kernel.  I couldn't get independently high enough.  My blend used half my stuff and half of that kernel in the end that did best on private.  The large kernel model, part of that 0.761 score, seems to be doing better on fake data for weaker signals.  But I never could reproduce that with own large kernel or replknet models.<br>\n2) Why did VAE fail so bad?  I never got out of AUC~0.5 domain.</p>",
      "rawMarkdown": "Thanks to organizers and making a fun physics-based competition.\n\nI'm very impressed by the creativity (and simplicity) of the solutions (so far) posted from @junkoda and @shunrcn, I'm sure there are more to come.\n\nAlso thanks to the kaggle community for being such great competitors.\n\n# Solution:\n\n## Train\n- pyfstat to generate 20k signals with test-data-based gaps\n- Same portion of real vs. fake data as in test set (fake being roughly constant standard deviation)\n- All real type data generated by sampling from test data (e.g.start times, durations, standard deviations vs. time)\n- Gaussian random complex noise added to complex signal, as well as direct real noise to absolute magnitude.\n- Ignore h0, just add signal to noise in normalized way based upon total power in signal vs. noise along signal.  Signal power went from just barely human visible at 360x360 to 10x lower.\n- About 200k training samples, Adam, Cosine scheduler, batch=32*4\n\n## Train and Test\n- Fill-out time data to 5760, binning data to be uniformly spaced\n- Cleaned data by removing all line or other spot noise from test data (quite effectively and fast)\n- Fill-in gaps of all data with surrounding level of noise\n- Compute absolute value as well as H1 times complex conjugate of L1 (equiv of cross-correction in FFT space)\n- Reduced by np.mean (like avgpool) down to 360x360 to help reduce noise\n- Normalize\n\n## Model\n- tf_efficientnet_b5_ns, 20-30 epochs, 5 folds\n- No Aug\n- Mixup (was required to avoid overfitting, is that what others saw?)\n- TTA of various seeds for filling-in the noise between gaps (worried model would overfit on clusters of noise)\n- Blended best single model, 15 5-fold TTAs of that best, large kernel\n- Nicely, CV score matched public very well.\n\n# Things I wrongly disregarded as not helpful\n- Normalizing per time point.  I disregarded it too quickly, thinking the signal would vary in strength and that would make things harder.  I was simply wrong to think that.  Thanks @ren4yu \n\n# Things that failed or didn't help:\n- Treating as regression on power fraction (didn't help)\n- replknet large kernel model (didn't help)\n- Larger image sizes, 360x720 and 360x1440 (didn't help)\n- Denoising VAE with noise+signal given, signal as goal (failed totally to learn, even at small resolutions)\n- Stable diffusion (VAE they trained fails totally on our noise level)\n- resnet18 (failed to work at all for me)\n- mixnet_m (does ok)\n- mixnet_l (does ok)\n- convmixer_768_32 (does ok, but slow)\n- vit_base_patch16_224 and vit_relpos_small_patch16_224 (though ViT would pick up on sequence of patches, but failed totally on me, probably bug in usage)\n- large_kernel (https://www.kaggle.com/code/laeyoung/g2net-large-kernel-inference/data) but dropping avgpool since not needed.  Tried various kernel sizes, nothing really helped. Ok model but weaker than efficientnet\n- Training on fake and real data separately (only did bit worse)\n- Adding noise as Aug (didn't help)\n- BiLSTM + Conv approach like done for BH-BH case and other competitions (isft -> sft is not perfect, and even matching pyfstat sft, the human-visible signals get easily lost by the isft process, so time series approach would likely fail, so gave up)\n- Downloading real data to find segments that match in real test data, as @vslaykovsky said.  I did that, but quickly realized the data did not match for gaps etc.  So disregarded it.\n- Use pyfstat signal detector itself, both MCMC and grid search.  I had to hack the code to convert hd5f -> sft.  Does very poorly even on easy signals when range of search is even a tiny bigger than actual signal.  So failed to work.  \n- Look at test set overlap.  Form bounding boxes for every test data set, and find matches.  There are about 1000 matches, with varying overlaps.  I used those to remove the noise.  Clean-up standard deviation and mean at interfaces left over.  But didn't help, probably because subtractions made data less clean.\n\n# What I don't understand:\n1) I always got boost from blending with good amount of top kernel with 0.761 that was partially based upon large kernel.  I couldn't get independently high enough.  My blend used half my stuff and half of that kernel in the end that did best on private.  The large kernel model, part of that 0.761 score, seems to be doing better on fake data for weaker signals.  But I never could reproduce that with own large kernel or replknet models.\n2) Why did VAE fail so bad?  I never got out of AUC~0.5 domain.",
      "votes": 12
    },
    {
      "id": 2085479,
      "postDate": "2023-01-04T07:30:26.107Z",
      "content": "<p>Congrats Jonathan McKinney！Very excellent solution! </p>",
      "rawMarkdown": "Congrats Jonathan McKinney！Very excellent solution! ",
      "votes": 1
    },
    {
      "id": 2086342,
      "postDate": "2023-01-04T18:06:16.263Z",
      "content": "<p>Excellent solution! I can tell you put a lot of work into this competition. Congrats <a href=\"https://www.kaggle.com/pseudotensor\" target=\"_blank\">@pseudotensor</a>!</p>",
      "rawMarkdown": "Excellent solution! I can tell you put a lot of work into this competition. Congrats @pseudotensor!"
    }
  ],
  "comments": [
    {
      "id": 2085479,
      "author_name": "BarryZhou",
      "author_url": "",
      "post_date": "2023-01-04T07:30:26.107000",
      "content": "<p>Congrats Jonathan McKinney！Very excellent solution! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2086342,
      "author_name": "Ravi Shah",
      "author_url": "",
      "post_date": "2023-01-04T18:06:16.263000",
      "content": "<p>Excellent solution! I can tell you put a lot of work into this competition. Congrats <a href=\"https://www.kaggle.com/pseudotensor\" target=\"_blank\">@pseudotensor</a>!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2085453": "Thanks to organizers and making a fun physics-based competition.\n\nI'm very impressed by the creativity (and simplicity) of the solutions (so far) posted from @junkoda and @shunrcn, I'm sure there are more to come.\n\nAlso thanks to the kaggle community for being such great competitors.\n\n# Solution:\n\n## Train\n- pyfstat to generate 20k signals with test-data-based gaps\n- Same portion of real vs. fake data as in test set (fake being roughly constant standard deviation)\n- All real type data generated by sampling from test data (e.g.start times, durations, standard deviations vs. time)\n- Gaussian random complex noise added to complex signal, as well as direct real noise to absolute magnitude.\n- Ignore h0, just add signal to noise in normalized way based upon total power in signal vs. noise along signal.  Signal power went from just barely human visible at 360x360 to 10x lower.\n- About 200k training samples, Adam, Cosine scheduler, batch=32*4\n\n## Train and Test\n- Fill-out time data to 5760, binning data to be uniformly spaced\n- Cleaned data by removing all line or other spot noise from test data (quite effectively and fast)\n- Fill-in gaps of all data with surrounding level of noise\n- Compute absolute value as well as H1 times complex conjugate of L1 (equiv of cross-correction in FFT space)\n- Reduced by np.mean (like avgpool) down to 360x360 to help reduce noise\n- Normalize\n\n## Model\n- tf_efficientnet_b5_ns, 20-30 epochs, 5 folds\n- No Aug\n- Mixup (was required to avoid overfitting, is that what others saw?)\n- TTA of various seeds for filling-in the noise between gaps (worried model would overfit on clusters of noise)\n- Blended best single model, 15 5-fold TTAs of that best, large kernel\n- Nicely, CV score matched public very well.\n\n# Things I wrongly disregarded as not helpful\n- Normalizing per time point.  I disregarded it too quickly, thinking the signal would vary in strength and that would make things harder.  I was simply wrong to think that.  Thanks @ren4yu \n\n# Things that failed or didn't help:\n- Treating as regression on power fraction (didn't help)\n- replknet large kernel model (didn't help)\n- Larger image sizes, 360x720 and 360x1440 (didn't help)\n- Denoising VAE with noise+signal given, signal as goal (failed totally to learn, even at small resolutions)\n- Stable diffusion (VAE they trained fails totally on our noise level)\n- resnet18 (failed to work at all for me)\n- mixnet_m (does ok)\n- mixnet_l (does ok)\n- convmixer_768_32 (does ok, but slow)\n- vit_base_patch16_224 and vit_relpos_small_patch16_224 (though ViT would pick up on sequence of patches, but failed totally on me, probably bug in usage)\n- large_kernel (https://www.kaggle.com/code/laeyoung/g2net-large-kernel-inference/data) but dropping avgpool since not needed.  Tried various kernel sizes, nothing really helped. Ok model but weaker than efficientnet\n- Training on fake and real data separately (only did bit worse)\n- Adding noise as Aug (didn't help)\n- BiLSTM + Conv approach like done for BH-BH case and other competitions (isft -> sft is not perfect, and even matching pyfstat sft, the human-visible signals get easily lost by the isft process, so time series approach would likely fail, so gave up)\n- Downloading real data to find segments that match in real test data, as @vslaykovsky said.  I did that, but quickly realized the data did not match for gaps etc.  So disregarded it.\n- Use pyfstat signal detector itself, both MCMC and grid search.  I had to hack the code to convert hd5f -> sft.  Does very poorly even on easy signals when range of search is even a tiny bigger than actual signal.  So failed to work.  \n- Look at test set overlap.  Form bounding boxes for every test data set, and find matches.  There are about 1000 matches, with varying overlaps.  I used those to remove the noise.  Clean-up standard deviation and mean at interfaces left over.  But didn't help, probably because subtractions made data less clean.\n\n# What I don't understand:\n1) I always got boost from blending with good amount of top kernel with 0.761 that was partially based upon large kernel.  I couldn't get independently high enough.  My blend used half my stuff and half of that kernel in the end that did best on private.  The large kernel model, part of that 0.761 score, seems to be doing better on fake data for weaker signals.  But I never could reproduce that with own large kernel or replknet models.\n2) Why did VAE fail so bad?  I never got out of AUC~0.5 domain.",
    "2085479": "Congrats Jonathan McKinney！Very excellent solution! ",
    "2086342": "Excellent solution! I can tell you put a lot of work into this competition. Congrats @pseudotensor!"
  }
}