{
  "id": 375910,
  "title": "1st place solution: Summing the power with GPU",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/375910",
  "author_name": "🐢 Jun Koda",
  "post_date": "2023-01-04T01:35:40.379000",
  "votes": 101,
  "comment_count": 38,
  "views": 0,
  "content": "<p>I thank Kaggle and the organizers for hosting this gravitational wave competition. I enjoyed the previous binary back hole merger competition, as well, and was impressed a lot by the gold-medal solutions. I thought large pretrained image models would be the strongest anyway, but they outperformed with 1-dimensional convolutional neural networks. I joined this competition so that I could build such deep neural network models detecting the wave, not the power, … but failed.</p>\n<p>The largest difference between the two competitions is that the Earth rotates during 120 days and it imprints complicated frequency pattern into the frequency. Adding wave is very delicate, requires very accurate phase patterns, and I was not able to add up the complex Fourier modes effectively within reasonable computational resources. After more than one month without any progress, I thought I should get a silver medal even with an unsatisfactory approach, give up adding the wave and add the power.</p>\n<h1>Solution</h1>\n<ul>\n<li>Sum power (absolute-value squared) along various signal patterns</li>\n<li>No machine learning</li>\n<li>No use of external data or leakage</li>\n</ul>\n<h2>Power summation</h2>\n<ol>\n<li>Extract signal frequency and amplitude [total power P(t)] from the simulations</li>\n<li>Subtract Doppler shift frequency from the data frequency for 4000 signal patterns </li>\n<li>Weight the data proportional to the signal amplitude pattern; this is the optimal linear weight w(t)</li>\n<li>Sum the weighted power along lines: 360 frequencies (intercept) × 241 slops in [-120, 120] (frequency bin / 120 days)</li>\n<li>Take the maximum</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F1bb4eb2f23c068c7c2998fc8a6e20ef8%2Fpower_sum.png?generation=1672802594937302&amp;alt=media\" alt=\"\"></p>\n<p>The values are highly skewed from the typical range [-1, 1] because these are maxima of 4000 templates × 360 frequencies × 241 slops.</p>\n<p>This took 5 days using GPU RTX 3090. Full range of slope is [-360, 360] but [-120, 120] already took long enough.</p>\n<h2>Real noise normalization</h2>\n<ol>\n<li>Normalize by the noise rms at each time h -&gt; h / sigma(t)</li>\n<li>Remove single-frequency noise by masking anomalously large frequency bin</li>\n<li>Normalize the frequency dependence by the remaining rms sigma(f)</li>\n</ol>\n<p>I know the noise is not always written as a product of time dependence and frequency dependence, but I did not have time for better treatment. Some false positives are remaining in my prediction.</p>\n<h2>Follow up summation with sinc kernel</h2>\n<p>The signal spread among frequency bins with sinc function (assuming the window function for the short-time Fourier transform is almost top hat). I use a sinc kernel with width 8 and stride 1/8 frequency bin to collect the signal. This is the optimal linear weighting in the frequency direction. I recompute the power sum with this kernel around the largest-power line in the first step for a subsample of 400 templates. This gives a surprisingly large boost to the public score 0.825 -&gt; 0.848</p>\n<p>Finally, I apply a sigmoid to the standardized power sum and submit, which is same as just submitting the power sum. I thought the prediction value must also depend on the noise level; if the signal is undetected, there should be larger possibility to be positive for larger noise because more data are undetected. I modeled this effect but could not improve the score. </p>\n<p>PS:<br>\nMore about the weight<br>\n<a href=\"https://www.kaggle.com/code/junkoda/optimal-weights-for-signal-to-noise-ratio\" target=\"_blank\">https://www.kaggle.com/code/junkoda/optimal-weights-for-signal-to-noise-ratio</a></p>\n<p>Code<br>\n<a href=\"https://github.com/junkoda/kaggle_g2net2_solution\" target=\"_blank\">https://github.com/junkoda/kaggle_g2net2_solution</a></p>",
  "messages": [
    {
      "id": 2085205,
      "postDate": "2023-01-04T01:35:40.380Z",
      "content": "<p>I thank Kaggle and the organizers for hosting this gravitational wave competition. I enjoyed the previous binary back hole merger competition, as well, and was impressed a lot by the gold-medal solutions. I thought large pretrained image models would be the strongest anyway, but they outperformed with 1-dimensional convolutional neural networks. I joined this competition so that I could build such deep neural network models detecting the wave, not the power, … but failed.</p>\n<p>The largest difference between the two competitions is that the Earth rotates during 120 days and it imprints complicated frequency pattern into the frequency. Adding wave is very delicate, requires very accurate phase patterns, and I was not able to add up the complex Fourier modes effectively within reasonable computational resources. After more than one month without any progress, I thought I should get a silver medal even with an unsatisfactory approach, give up adding the wave and add the power.</p>\n<h1>Solution</h1>\n<ul>\n<li>Sum power (absolute-value squared) along various signal patterns</li>\n<li>No machine learning</li>\n<li>No use of external data or leakage</li>\n</ul>\n<h2>Power summation</h2>\n<ol>\n<li>Extract signal frequency and amplitude [total power P(t)] from the simulations</li>\n<li>Subtract Doppler shift frequency from the data frequency for 4000 signal patterns </li>\n<li>Weight the data proportional to the signal amplitude pattern; this is the optimal linear weight w(t)</li>\n<li>Sum the weighted power along lines: 360 frequencies (intercept) × 241 slops in [-120, 120] (frequency bin / 120 days)</li>\n<li>Take the maximum</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F1bb4eb2f23c068c7c2998fc8a6e20ef8%2Fpower_sum.png?generation=1672802594937302&amp;alt=media\" alt=\"\"></p>\n<p>The values are highly skewed from the typical range [-1, 1] because these are maxima of 4000 templates × 360 frequencies × 241 slops.</p>\n<p>This took 5 days using GPU RTX 3090. Full range of slope is [-360, 360] but [-120, 120] already took long enough.</p>\n<h2>Real noise normalization</h2>\n<ol>\n<li>Normalize by the noise rms at each time h -&gt; h / sigma(t)</li>\n<li>Remove single-frequency noise by masking anomalously large frequency bin</li>\n<li>Normalize the frequency dependence by the remaining rms sigma(f)</li>\n</ol>\n<p>I know the noise is not always written as a product of time dependence and frequency dependence, but I did not have time for better treatment. Some false positives are remaining in my prediction.</p>\n<h2>Follow up summation with sinc kernel</h2>\n<p>The signal spread among frequency bins with sinc function (assuming the window function for the short-time Fourier transform is almost top hat). I use a sinc kernel with width 8 and stride 1/8 frequency bin to collect the signal. This is the optimal linear weighting in the frequency direction. I recompute the power sum with this kernel around the largest-power line in the first step for a subsample of 400 templates. This gives a surprisingly large boost to the public score 0.825 -&gt; 0.848</p>\n<p>Finally, I apply a sigmoid to the standardized power sum and submit, which is same as just submitting the power sum. I thought the prediction value must also depend on the noise level; if the signal is undetected, there should be larger possibility to be positive for larger noise because more data are undetected. I modeled this effect but could not improve the score. </p>\n<p>PS:<br>\nMore about the weight<br>\n<a href=\"https://www.kaggle.com/code/junkoda/optimal-weights-for-signal-to-noise-ratio\" target=\"_blank\">https://www.kaggle.com/code/junkoda/optimal-weights-for-signal-to-noise-ratio</a></p>\n<p>Code<br>\n<a href=\"https://github.com/junkoda/kaggle_g2net2_solution\" target=\"_blank\">https://github.com/junkoda/kaggle_g2net2_solution</a></p>",
      "rawMarkdown": "I thank Kaggle and the organizers for hosting this gravitational wave competition. I enjoyed the previous binary back hole merger competition, as well, and was impressed a lot by the gold-medal solutions. I thought large pretrained image models would be the strongest anyway, but they outperformed with 1-dimensional convolutional neural networks. I joined this competition so that I could build such deep neural network models detecting the wave, not the power, ... but failed.\n\nThe largest difference between the two competitions is that the Earth rotates during 120 days and it imprints complicated frequency pattern into the frequency. Adding wave is very delicate, requires very accurate phase patterns, and I was not able to add up the complex Fourier modes effectively within reasonable computational resources. After more than one month without any progress, I thought I should get a silver medal even with an unsatisfactory approach, give up adding the wave and add the power.\n\n# Solution\n\n- Sum power (absolute-value squared) along various signal patterns\n- No machine learning\n- No use of external data or leakage\n\n## Power summation\n\n1. Extract signal frequency and amplitude [total power P(t)] from the simulations\n2. Subtract Doppler shift frequency from the data frequency for 4000 signal patterns \n3. Weight the data proportional to the signal amplitude pattern; this is the optimal linear weight w(t)\n4. Sum the weighted power along lines: 360 frequencies (intercept) × 241 slops in [-120, 120] (frequency bin / 120 days)\n5. Take the maximum\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F1bb4eb2f23c068c7c2998fc8a6e20ef8%2Fpower_sum.png?generation=1672802594937302&alt=media)\n\nThe values are highly skewed from the typical range [-1, 1] because these are maxima of 4000 templates × 360 frequencies × 241 slops.\n\nThis took 5 days using GPU RTX 3090. Full range of slope is [-360, 360] but [-120, 120] already took long enough.\n\n##  Real noise normalization\n\n1. Normalize by the noise rms at each time h -> h / sigma(t)\n2. Remove single-frequency noise by masking anomalously large frequency bin\n3. Normalize the frequency dependence by the remaining rms sigma(f)\n\nI know the noise is not always written as a product of time dependence and frequency dependence, but I did not have time for better treatment. Some false positives are remaining in my prediction.\n\n## Follow up summation with sinc kernel\n\nThe signal spread among frequency bins with sinc function (assuming the window function for the short-time Fourier transform is almost top hat). I use a sinc kernel with width 8 and stride 1/8 frequency bin to collect the signal. This is the optimal linear weighting in the frequency direction. I recompute the power sum with this kernel around the largest-power line in the first step for a subsample of 400 templates. This gives a surprisingly large boost to the public score 0.825 -> 0.848\n\nFinally, I apply a sigmoid to the standardized power sum and submit, which is same as just submitting the power sum. I thought the prediction value must also depend on the noise level; if the signal is undetected, there should be larger possibility to be positive for larger noise because more data are undetected. I modeled this effect but could not improve the score. \n\nPS:\nMore about the weight\nhttps://www.kaggle.com/code/junkoda/optimal-weights-for-signal-to-noise-ratio\n\nCode\nhttps://github.com/junkoda/kaggle_g2net2_solution\n",
      "votes": 101
    },
    {
      "id": 2104800,
      "postDate": "2023-01-18T03:49:56.957Z",
      "content": "<p>Thanks for posting the detailed solution! It was amazing to note the procedure used in the weights set to the signal-to-noise ratio for improving the model. My solution involved the use of t-test and proximity analysis of signal-to-noise ratios. The notebook is very informative and helped me think on a different line.</p>",
      "rawMarkdown": "Thanks for posting the detailed solution! It was amazing to note the procedure used in the weights set to the signal-to-noise ratio for improving the model. My solution involved the use of t-test and proximity analysis of signal-to-noise ratios. The notebook is very informative and helped me think on a different line.\n",
      "votes": 1
    },
    {
      "id": 2086711,
      "postDate": "2023-01-05T00:45:43.887Z",
      "content": "<p>the use of the sinc kernel is superb</p>",
      "rawMarkdown": "the use of the sinc kernel is superb",
      "votes": 1
    },
    {
      "id": 2086075,
      "postDate": "2023-01-04T14:50:46.520Z",
      "content": "<p>Congratulations, a great achievement! Your write-up is a good example of being willing to pivot away from initial intuitions about a problem to find a better outcome with a different approach. A good lesson for us all.</p>",
      "rawMarkdown": "Congratulations, a great achievement! Your write-up is a good example of being willing to pivot away from initial intuitions about a problem to find a better outcome with a different approach. A good lesson for us all.",
      "votes": 1
    },
    {
      "id": 2085822,
      "postDate": "2023-01-04T12:25:20.540Z",
      "content": "<p>Congratulations &amp; Thanks for the insightful solution. </p>",
      "rawMarkdown": "Congratulations & Thanks for the insightful solution. ",
      "votes": 1
    },
    {
      "id": 2085726,
      "postDate": "2023-01-04T10:55:23.007Z",
      "content": "<p><a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> <br>\nCongratulations for the championship!<br>\nWhat do you mean when you say “slope”?</p>",
      "rawMarkdown": "@junkoda \nCongratulations for the championship!\nWhat do you mean when you say “slope”?",
      "votes": 1,
      "replies": [
        {
          "id": 2086614,
          "postDate": "2023-01-04T21:33:53.860Z",
          "content": "<p>Thanks for the question. I thought it might not clear, too. I mean slope and intercept for a linear equation y = a x + b.<br>\nf(t) = f0 + fdot t<br>\nf0 is the 360 frequencies and fdot is the 241 patterns of <code>slope</code></p>",
          "rawMarkdown": "Thanks for the question. I thought it might not clear, too. I mean slope and intercept for a linear equation y = a x + b.\nf(t) = f0 + fdot t\nf0 is the 360 frequencies and fdot is the 241 patterns of `slope`\n",
          "votes": 2,
          "replies": [
            {
              "id": 2115145,
              "postDate": "2023-01-25T14:20:05.333Z",
              "content": "<p>Thank you for your answer, and sorry for my late reaction.<br>\nThe question was cleared up!</p>",
              "rawMarkdown": "Thank you for your answer, and sorry for my late reaction.\nThe question was cleared up!"
            }
          ]
        }
      ]
    },
    {
      "id": 2085664,
      "postDate": "2023-01-04T09:58:49.363Z",
      "content": "<p>Congratulations on winning the first prize!!! 👍 Very impressive non-machine learning solution based on profound understanding of the competition. I used similar data normalization to reduce the impact of real noise in my CNN model, which is simple but effective.</p>",
      "rawMarkdown": "Congratulations on winning the first prize!!! 👍 Very impressive non-machine learning solution based on profound understanding of the competition. I used similar data normalization to reduce the impact of real noise in my CNN model, which is simple but effective.",
      "votes": 1
    },
    {
      "id": 2085610,
      "postDate": "2023-01-04T09:24:46.047Z",
      "content": "<p>Very cool.  I wonder if GW scientists can use this as a bootstrap step, where this is done first and then they can go back and do parameter estimation.  Doing it all in one go as a template is too costly or poor (pyfstat itself requires very narrow search around signal).</p>",
      "rawMarkdown": "Very cool.  I wonder if GW scientists can use this as a bootstrap step, where this is done first and then they can go back and do parameter estimation.  Doing it all in one go as a template is too costly or poor (pyfstat itself requires very narrow search around signal).",
      "votes": 1,
      "replies": [
        {
          "id": 2086620,
          "postDate": "2023-01-04T21:40:50.640Z",
          "content": "<p>I guess scientists usually use iterative approach, Markov chain Monte Carlo, like the 6th place simulated annealing, but I don't know.</p>",
          "rawMarkdown": "I guess scientists usually use iterative approach, Markov chain Monte Carlo, like the 6th place simulated annealing, but I don't know.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2085587,
      "postDate": "2023-01-04T09:04:17.277Z",
      "content": "<p>Thank you again! I've learned a lot starting from your notebook. And even if it is not impressive, to achieve a 0.73 score was an impossible task at first. At the beginning even 0.6 score was an achievement. Still with your efficient-net approach , 10-20k generated samples a keras single model  and two hours of TPU training the result is amazing for me. </p>",
      "rawMarkdown": "Thank you again! I've learned a lot starting from your notebook. And even if it is not impressive, to achieve a 0.73 score was an impossible task at first. At the beginning even 0.6 score was an achievement. Still with your efficient-net approach , 10-20k generated samples a keras single model  and two hours of TPU training the result is amazing for me. ",
      "votes": 1
    },
    {
      "id": 2085461,
      "postDate": "2023-01-04T07:09:04.867Z",
      "content": "<p>Congratulations, <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>! When I watched your score and the number os submissions, I was willing to bet that you are using a ‘physics-based’ method. We also used a very similar approach, but did not account for the amplitude modulation.</p>",
      "rawMarkdown": "Congratulations, @junkoda! When I watched your score and the number os submissions, I was willing to bet that you are using a ‘physics-based’ method. We also used a very similar approach, but did not account for the amplitude modulation.",
      "votes": 1,
      "replies": [
        {
          "id": 2085550,
          "postDate": "2023-01-04T08:22:07.683Z",
          "content": "<p>Thanks! I'm looking forward to reading you solution, too. Mine is brute force, but do you use something like Markov chain Monte Carlo? Amazing first medal!</p>",
          "rawMarkdown": "Thanks! I'm looking forward to reading you solution, too. Mine is brute force, but do you use something like Markov chain Monte Carlo? Amazing first medal!",
          "votes": 2,
          "replies": [
            {
              "id": 2085584,
              "postDate": "2023-01-04T09:02:20.037Z",
              "content": "<p>Thank you very much! We will post the solution soon.<br>\nYes, we used differential evolution with objective = - (max power from the two detectors). We also tried nested sampling, but it did not perform well and took longer than a grid scan.</p>",
              "rawMarkdown": "Thank you very much! We will post the solution soon.\nYes, we used differential evolution with objective = - (max power from the two detectors). We also tried nested sampling, but it did not perform well and took longer than a grid scan.",
              "votes": 1
            },
            {
              "id": 2085601,
              "postDate": "2023-01-04T09:16:11.767Z",
              "content": "<p>Actually, when we started scanning over the frequency derivative (in addition to the source position) and the computational cost started to explode, we considered brute-forcing the problem by implementing the summation in CUDA(.jl). But before that we gave the algorithms from <a href=\"https://github.com/robertfeldt/BlackBoxOptim.jl\" target=\"_blank\">BlackBoxOptim.jl</a> a try, and some flavours of differential evolution did so well that we gave up about CUDA. It would be interesting to see if these optimisers would work with your summation method too. In our case we didn’t implement the sinc kernel (and are now starting to regret it ;) ) so the peaks were slightly spread out, which may have helped the optimisers a bit at the cost of reducing the signal sensitivity.</p>",
              "rawMarkdown": "Actually, when we started scanning over the frequency derivative (in addition to the source position) and the computational cost started to explode, we considered brute-forcing the problem by implementing the summation in CUDA(.jl). But before that we gave the algorithms from [BlackBoxOptim.jl](https://github.com/robertfeldt/BlackBoxOptim.jl) a try, and some flavours of differential evolution did so well that we gave up about CUDA. It would be interesting to see if these optimisers would work with your summation method too. In our case we didn’t implement the sinc kernel (and are now starting to regret it ;) ) so the peaks were slightly spread out, which may have helped the optimisers a bit at the cost of reducing the signal sensitivity.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2085405,
      "postDate": "2023-01-04T06:35:08.043Z",
      "content": "<p>Seriously - this is gold! Thank you for sharing and for the clear explanation! It shows how being clever is the most important thing. Going back to the basics time and time again and just trying to stop yourself from overengineering. Thank you! 🙏</p>",
      "rawMarkdown": "Seriously - this is gold! Thank you for sharing and for the clear explanation! It shows how being clever is the most important thing. Going back to the basics time and time again and just trying to stop yourself from overengineering. Thank you! 🙏",
      "votes": 1
    },
    {
      "id": 2085373,
      "postDate": "2023-01-04T05:44:39.350Z",
      "content": "<p>Amazing! Congratulations <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> !</p>",
      "rawMarkdown": "Amazing! Congratulations @junkoda !",
      "votes": 1
    },
    {
      "id": 2085283,
      "postDate": "2023-01-04T03:30:04.280Z",
      "content": "<p>would you like post your failed deep neural network models attempts ? maybe we can learn from it. </p>",
      "rawMarkdown": "would you like post your failed deep neural network models attempts ? maybe we can learn from it. ",
      "votes": 1,
      "replies": [
        {
          "id": 2085302,
          "postDate": "2023-01-04T03:56:31.453Z",
          "content": "<p>I don't think I have anything worthwhile… For example, I made tiny ResNet-like models reading last year solutions, but those are not better than simple 2D image models even for simple linear frequency growth. There is a wonderful 1D wave detection model by the previous 2nd place winner:</p>\n<p><a href=\"https://github.com/analokmaus/kaggle-g2net-public\" target=\"_blank\">https://github.com/analokmaus/kaggle-g2net-public</a></p>",
          "rawMarkdown": "I don't think I have anything worthwhile... For example, I made tiny ResNet-like models reading last year solutions, but those are not better than simple 2D image models even for simple linear frequency growth. There is a wonderful 1D wave detection model by the previous 2nd place winner:\n\nhttps://github.com/analokmaus/kaggle-g2net-public",
          "votes": 1,
          "replies": [
            {
              "id": 2085323,
              "postDate": "2023-01-04T04:23:36.887Z",
              "content": "<p>thank you. i used your [Basic spectrogram image classification] as base, then added two more channels on top of the training dataset plus data augmentation using albumentations, be able to reach 0.8x (but the final score is 0.6x from 0.5x ) during the training-evaluation (may have data-leakage ) processes, no external data, no extra noise and pure signal generation used. </p>",
              "rawMarkdown": "thank you. i used your [Basic spectrogram image classification] as base, then added two more channels on top of the training dataset plus data augmentation using albumentations, be able to reach 0.8x (but the final score is 0.6x from 0.5x ) during the training-evaluation (may have data-leakage ) processes, no external data, no extra noise and pure signal generation used. "
            }
          ]
        }
      ]
    },
    {
      "id": 2085250,
      "postDate": "2023-01-04T02:51:19.400Z",
      "content": "<p>Congrats @🐢 Jun Koda！Very wonderful solution!</p>",
      "rawMarkdown": "Congrats @🐢 Jun Koda！Very wonderful solution!",
      "votes": 1,
      "replies": [
        {
          "id": 2085254,
          "postDate": "2023-01-04T02:55:55.157Z",
          "content": "<p>Thanks. It must be a disappointment for people who want to learn machine learning. I hope there is something useful for signal processing.</p>",
          "rawMarkdown": "Thanks. It must be a disappointment for people who want to learn machine learning. I hope there is something useful for signal processing.",
          "votes": 2,
          "replies": [
            {
              "id": 2085485,
              "postDate": "2023-01-04T07:35:33.973Z",
              "content": "<p>Haha, I think the best solution is the one that can solve practical problems, no matter what method is used.</p>",
              "rawMarkdown": "Haha, I think the best solution is the one that can solve practical problems, no matter what method is used.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2086818,
      "postDate": "2023-01-05T04:16:22.120Z",
      "content": "<p>You are amazing. Congratulations, and thank you for sharing the solution (I like the link from the leaderboard) </p>",
      "rawMarkdown": "You are amazing. Congratulations, and thank you for sharing the solution (I like the link from the leaderboard) ",
      "votes": 2
    },
    {
      "id": 2086319,
      "postDate": "2023-01-04T17:51:11.967Z",
      "content": "<p>This is an amazing solution! Thanks you for sharing your approach! It is a good reminder that machine learning should not always be the go to answer. Congrats on winning <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>!</p>",
      "rawMarkdown": "This is an amazing solution! Thanks you for sharing your approach! It is a good reminder that machine learning should not always be the go to answer. Congrats on winning @junkoda!",
      "votes": 2
    },
    {
      "id": 2086181,
      "postDate": "2023-01-04T16:15:13.637Z",
      "content": "<p>Congratulations on your first place!! 🎉<br>\nI was amazed by the performance without any leakage: We took a similar approach (GPU-accelerated random search of signals) but scored 0.835 in private LB without leakage.</p>\n<p>These might be a little obvious since I am not familiar with this field, but can I ask some questions?</p>\n<ol>\n<li>What do you mean by signal amplitude (and how did you simulate it)?</li>\n<li>In what way weighting proportional to signal amplitude is optimal?</li>\n</ol>",
      "rawMarkdown": "Congratulations on your first place!! 🎉\nI was amazed by the performance without any leakage: We took a similar approach (GPU-accelerated random search of signals) but scored 0.835 in private LB without leakage.\n\nThese might be a little obvious since I am not familiar with this field, but can I ask some questions?\n1. What do you mean by signal amplitude (and how did you simulate it)?\n2. In what way weighting proportional to signal amplitude is optimal?",
      "votes": 2,
      "replies": [
        {
          "id": 2087116,
          "postDate": "2023-01-05T10:45:18.843Z",
          "content": "<ol>\n<li><p>The \"simulation\" is the standard PyFstat data generation. I created without noise, then the amplitude squared is the total power: sum(|h|**2, axis=frequency)</p></li>\n<li><p>What I wrote was obviously too short and I was writing the detail with LaTex today, but it was not an appropriate day after days of excitement and caffeine overdose. I'll add it here and let you know after a few days.</p></li>\n</ol>\n<p>Your performance was amazing, too! Looks like working 24 hrs a day, updating the score at 3am, 4 am. I don't think I'll ever see a team getting the first in the last 3 minutes with such a large change in the score again.</p>",
          "rawMarkdown": "1. The \"simulation\" is the standard PyFstat data generation. I created without noise, then the amplitude squared is the total power: sum(|h|**2, axis=frequency)\n\n2. What I wrote was obviously too short and I was writing the detail with LaTex today, but it was not an appropriate day after days of excitement and caffeine overdose. I'll add it here and let you know after a few days.\n\nYour performance was amazing, too! Looks like working 24 hrs a day, updating the score at 3am, 4 am. I don't think I'll ever see a team getting the first in the last 3 minutes with such a large change in the score again.\n",
          "votes": 1,
          "replies": [
            {
              "id": 2092408,
              "postDate": "2023-01-09T10:27:21.107Z",
              "content": "<p>Sorry for the late reply. I didn't realize the amplitudes differ by timestamps.<br>\nWe worked hard until the end (we might have seemed quick because we were a team of three!), but I have to admit that your solution was more sophisticated. Anyway, thanks for the reply, and good game!</p>",
              "rawMarkdown": "Sorry for the late reply. I didn't realize the amplitudes differ by timestamps.\nWe worked hard until the end (we might have seemed quick because we were a team of three!), but I have to admit that your solution was more sophisticated. Anyway, thanks for the reply, and good game!"
            },
            {
              "id": 2104028,
              "postDate": "2023-01-17T14:33:00.763Z",
              "content": "<p>I wrote what the weight is more in detail in the notebook<br>\n<a href=\"https://www.kaggle.com/code/junkoda/optimal-weights-for-signal-to-noise-ratio\" target=\"_blank\">https://www.kaggle.com/code/junkoda/optimal-weights-for-signal-to-noise-ratio</a><br>\nbecause I got similar questions, too. I wanted to brush up a little more, but looks like you already noticed.</p>",
              "rawMarkdown": "I wrote what the weight is more in detail in the notebook\nhttps://www.kaggle.com/code/junkoda/optimal-weights-for-signal-to-noise-ratio\nbecause I got similar questions, too. I wanted to brush up a little more, but looks like you already noticed."
            },
            {
              "id": 2104037,
              "postDate": "2023-01-17T14:42:25.373Z",
              "content": "<p>Thanks so much!!! Let me read it in detail when I have time.<br>\nBy the way, it seems the link to the source code is not public…?</p>",
              "rawMarkdown": "Thanks so much!!! Let me read it in detail when I have time.\nBy the way, it seems the link to the source code is not public...?",
              "votes": 1
            },
            {
              "id": 2104615,
              "postDate": "2023-01-17T23:30:52.520Z",
              "content": "<p>Thanks for the feedback. I fixed the github repository to public. I hope it is accessible now (I'll continue to add more explanations).<br>\nI am impressed by the shortness of your code. I'll learn from you, too.</p>",
              "rawMarkdown": "Thanks for the feedback. I fixed the github repository to public. I hope it is accessible now (I'll continue to add more explanations).\nI am impressed by the shortness of your code. I'll learn from you, too.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2085299,
      "postDate": "2023-01-04T03:51:40.043Z",
      "content": "<p>Congratulations Jun. I'm very happy that you be the winner because you worked hard on it from the beginning and generously shared your Basic Spectrogram Classificator Notebook with us. I learnt a lot with you and the others. And all this knowledge did that I could work in a solution based in Isaac Horowitz's QFT (Quantitative Feedback Theory) applied in Industry but with detectors in this case. I though that the uncertainty could be quantified applying this theory but I fell in Bias, so I was wrong. However, this experience was a good try, and I could meet great people. I hope see you all here in future. Thanks for everything 👏👏👏</p>",
      "rawMarkdown": "Congratulations Jun. I'm very happy that you be the winner because you worked hard on it from the beginning and generously shared your Basic Spectrogram Classificator Notebook with us. I learnt a lot with you and the others. And all this knowledge did that I could work in a solution based in Isaac Horowitz's QFT (Quantitative Feedback Theory) applied in Industry but with detectors in this case. I though that the uncertainty could be quantified applying this theory but I fell in Bias, so I was wrong. However, this experience was a good try, and I could meet great people. I hope see you all here in future. Thanks for everything 👏👏👏",
      "votes": 2
    },
    {
      "id": 2085295,
      "postDate": "2023-01-04T03:49:04.627Z",
      "content": "<p>Hearty congratulations to <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> for the result!!<br>\nThis was indeed an involved challenge. All the best and regards!!</p>",
      "rawMarkdown": "Hearty congratulations to @junkoda for the result!!\nThis was indeed an involved challenge. All the best and regards!!",
      "votes": 2
    },
    {
      "id": 2087300,
      "postDate": "2023-01-05T13:59:19.480Z",
      "content": "<p>Congratulations! Your journey through the competition is amazing. Sinc Kernel approach with power sum makes detection easy.</p>",
      "rawMarkdown": "Congratulations! Your journey through the competition is amazing. Sinc Kernel approach with power sum makes detection easy."
    },
    {
      "id": 2086487,
      "postDate": "2023-01-04T20:01:50.863Z",
      "content": "<p>congratulation sir. Really happy for you</p>",
      "rawMarkdown": "congratulation sir. Really happy for you"
    },
    {
      "id": 2086408,
      "postDate": "2023-01-04T18:47:46.993Z",
      "content": "<blockquote>\n  <p>Some false positives are remaining in my prediction.</p>\n</blockquote>\n<p>Well, some test data are using real noise background from the detectors. Who knows, probably there is a signal lurking there after all :) </p>\n<p>Congrats on a physical solution!</p>",
      "rawMarkdown": ">Some false positives are remaining in my prediction.\n\nWell, some test data are using real noise background from the detectors. Who knows, probably there is a signal lurking there after all :) \n\nCongrats on a physical solution!"
    },
    {
      "id": 2085419,
      "postDate": "2023-01-04T06:48:38.597Z",
      "content": "<p><a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> What other solutions did you use? I am just curious - what are your takeaways from this competition?</p>",
      "rawMarkdown": "@junkoda What other solutions did you use? I am just curious - what are your takeaways from this competition?",
      "replies": [
        {
          "id": 2085555,
          "postDate": "2023-01-04T08:27:19.040Z",
          "content": "<p>For an unsuccessful wave model, I tried to add 48 complex data each day using a template, create 3D block of frequency x frequency time derivative x time, and do 3D CNN, but I need orders of magnitude more templates to add complex wave.</p>\n<p>Take away…, I am much happier with a practical solution than an incomplete ideal solution that didn't meet the deadline; the latter is my situation with the GPS Google Smartphone competition last year.</p>",
          "rawMarkdown": "For an unsuccessful wave model, I tried to add 48 complex data each day using a template, create 3D block of frequency x frequency time derivative x time, and do 3D CNN, but I need orders of magnitude more templates to add complex wave.\n\nTake away..., I am much happier with a practical solution than an incomplete ideal solution that didn't meet the deadline; the latter is my situation with the GPS Google Smartphone competition last year.",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2104800,
      "author_name": "Suma Mallapragada",
      "author_url": "",
      "post_date": "2023-01-18T03:49:56.957000",
      "content": "<p>Thanks for posting the detailed solution! It was amazing to note the procedure used in the weights set to the signal-to-noise ratio for improving the model. My solution involved the use of t-test and proximity analysis of signal-to-noise ratios. The notebook is very informative and helped me think on a different line.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2086711,
      "author_name": "Red Bear",
      "author_url": "",
      "post_date": "2023-01-05T00:45:43.887000",
      "content": "<p>the use of the sinc kernel is superb</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2086075,
      "author_name": "MichaelP",
      "author_url": "",
      "post_date": "2023-01-04T14:50:46.520000",
      "content": "<p>Congratulations, a great achievement! Your write-up is a good example of being willing to pivot away from initial intuitions about a problem to find a better outcome with a different approach. A good lesson for us all.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2085822,
      "author_name": "Gaju Ahmed",
      "author_url": "",
      "post_date": "2023-01-04T12:25:20.540000",
      "content": "<p>Congratulations &amp; Thanks for the insightful solution. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2085726,
      "author_name": "Y-Haneji",
      "author_url": "",
      "post_date": "2023-01-04T10:55:23.007000",
      "content": "<p><a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> <br>\nCongratulations for the championship!<br>\nWhat do you mean when you say “slope”?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2086614,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-01-04T21:33:53.860000",
          "content": "<p>Thanks for the question. I thought it might not clear, too. I mean slope and intercept for a linear equation y = a x + b.<br>\nf(t) = f0 + fdot t<br>\nf0 is the 360 frequencies and fdot is the 241 patterns of <code>slope</code></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2115145,
              "author_name": "Y-Haneji",
              "author_url": "",
              "post_date": "2023-01-25T14:20:05.333000",
              "content": "<p>Thank you for your answer, and sorry for my late reaction.<br>\nThe question was cleared up!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2085664,
      "author_name": "Jin Niu",
      "author_url": "",
      "post_date": "2023-01-04T09:58:49.363000",
      "content": "<p>Congratulations on winning the first prize!!! 👍 Very impressive non-machine learning solution based on profound understanding of the competition. I used similar data normalization to reduce the impact of real noise in my CNN model, which is simple but effective.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2085610,
      "author_name": "Jonathan McKinney",
      "author_url": "",
      "post_date": "2023-01-04T09:24:46.047000",
      "content": "<p>Very cool.  I wonder if GW scientists can use this as a bootstrap step, where this is done first and then they can go back and do parameter estimation.  Doing it all in one go as a template is too costly or poor (pyfstat itself requires very narrow search around signal).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2086620,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-01-04T21:40:50.640000",
          "content": "<p>I guess scientists usually use iterative approach, Markov chain Monte Carlo, like the 6th place simulated annealing, but I don't know.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2085587,
      "author_name": "George Chirita",
      "author_url": "",
      "post_date": "2023-01-04T09:04:17.277000",
      "content": "<p>Thank you again! I've learned a lot starting from your notebook. And even if it is not impressive, to achieve a 0.73 score was an impossible task at first. At the beginning even 0.6 score was an achievement. Still with your efficient-net approach , 10-20k generated samples a keras single model  and two hours of TPU training the result is amazing for me. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2085461,
      "author_name": "Inar Timiryasov",
      "author_url": "",
      "post_date": "2023-01-04T07:09:04.867000",
      "content": "<p>Congratulations, <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>! When I watched your score and the number os submissions, I was willing to bet that you are using a ‘physics-based’ method. We also used a very similar approach, but did not account for the amplitude modulation.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2085550,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-01-04T08:22:07.683000",
          "content": "<p>Thanks! I'm looking forward to reading you solution, too. Mine is brute force, but do you use something like Markov chain Monte Carlo? Amazing first medal!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2085584,
              "author_name": "Inar Timiryasov",
              "author_url": "",
              "post_date": "2023-01-04T09:02:20.037000",
              "content": "<p>Thank you very much! We will post the solution soon.<br>\nYes, we used differential evolution with objective = - (max power from the two detectors). We also tried nested sampling, but it did not perform well and took longer than a grid scan.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2085601,
              "author_name": "Jean-Loup Tastet",
              "author_url": "",
              "post_date": "2023-01-04T09:16:11.767000",
              "content": "<p>Actually, when we started scanning over the frequency derivative (in addition to the source position) and the computational cost started to explode, we considered brute-forcing the problem by implementing the summation in CUDA(.jl). But before that we gave the algorithms from <a href=\"https://github.com/robertfeldt/BlackBoxOptim.jl\" target=\"_blank\">BlackBoxOptim.jl</a> a try, and some flavours of differential evolution did so well that we gave up about CUDA. It would be interesting to see if these optimisers would work with your summation method too. In our case we didn’t implement the sinc kernel (and are now starting to regret it ;) ) so the peaks were slightly spread out, which may have helped the optimisers a bit at the cost of reducing the signal sensitivity.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2085405,
      "author_name": "PiotrKlinke",
      "author_url": "",
      "post_date": "2023-01-04T06:35:08.043000",
      "content": "<p>Seriously - this is gold! Thank you for sharing and for the clear explanation! It shows how being clever is the most important thing. Going back to the basics time and time again and just trying to stop yourself from overengineering. Thank you! 🙏</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2085373,
      "author_name": "BwandoWando",
      "author_url": "",
      "post_date": "2023-01-04T05:44:39.350000",
      "content": "<p>Amazing! Congratulations <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2085283,
      "author_name": "Kefan Xu",
      "author_url": "",
      "post_date": "2023-01-04T03:30:04.280000",
      "content": "<p>would you like post your failed deep neural network models attempts ? maybe we can learn from it. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2085302,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-01-04T03:56:31.453000",
          "content": "<p>I don't think I have anything worthwhile… For example, I made tiny ResNet-like models reading last year solutions, but those are not better than simple 2D image models even for simple linear frequency growth. There is a wonderful 1D wave detection model by the previous 2nd place winner:</p>\n<p><a href=\"https://github.com/analokmaus/kaggle-g2net-public\" target=\"_blank\">https://github.com/analokmaus/kaggle-g2net-public</a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2085323,
              "author_name": "Kefan Xu",
              "author_url": "",
              "post_date": "2023-01-04T04:23:36.887000",
              "content": "<p>thank you. i used your [Basic spectrogram image classification] as base, then added two more channels on top of the training dataset plus data augmentation using albumentations, be able to reach 0.8x (but the final score is 0.6x from 0.5x ) during the training-evaluation (may have data-leakage ) processes, no external data, no extra noise and pure signal generation used. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2085250,
      "author_name": "BarryZhou",
      "author_url": "",
      "post_date": "2023-01-04T02:51:19.400000",
      "content": "<p>Congrats @🐢 Jun Koda！Very wonderful solution!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2085254,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-01-04T02:55:55.157000",
          "content": "<p>Thanks. It must be a disappointment for people who want to learn machine learning. I hope there is something useful for signal processing.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2085485,
              "author_name": "BarryZhou",
              "author_url": "",
              "post_date": "2023-01-04T07:35:33.973000",
              "content": "<p>Haha, I think the best solution is the one that can solve practical problems, no matter what method is used.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2086818,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2023-01-05T04:16:22.120000",
      "content": "<p>You are amazing. Congratulations, and thank you for sharing the solution (I like the link from the leaderboard) </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2086319,
      "author_name": "Ravi Shah",
      "author_url": "",
      "post_date": "2023-01-04T17:51:11.967000",
      "content": "<p>This is an amazing solution! Thanks you for sharing your approach! It is a good reminder that machine learning should not always be the go to answer. Congrats on winning <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2086181,
      "author_name": "knshnb",
      "author_url": "",
      "post_date": "2023-01-04T16:15:13.637000",
      "content": "<p>Congratulations on your first place!! 🎉<br>\nI was amazed by the performance without any leakage: We took a similar approach (GPU-accelerated random search of signals) but scored 0.835 in private LB without leakage.</p>\n<p>These might be a little obvious since I am not familiar with this field, but can I ask some questions?</p>\n<ol>\n<li>What do you mean by signal amplitude (and how did you simulate it)?</li>\n<li>In what way weighting proportional to signal amplitude is optimal?</li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 2087116,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-01-05T10:45:18.843000",
          "content": "<ol>\n<li><p>The \"simulation\" is the standard PyFstat data generation. I created without noise, then the amplitude squared is the total power: sum(|h|**2, axis=frequency)</p></li>\n<li><p>What I wrote was obviously too short and I was writing the detail with LaTex today, but it was not an appropriate day after days of excitement and caffeine overdose. I'll add it here and let you know after a few days.</p></li>\n</ol>\n<p>Your performance was amazing, too! Looks like working 24 hrs a day, updating the score at 3am, 4 am. I don't think I'll ever see a team getting the first in the last 3 minutes with such a large change in the score again.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2092408,
              "author_name": "knshnb",
              "author_url": "",
              "post_date": "2023-01-09T10:27:21.107000",
              "content": "<p>Sorry for the late reply. I didn't realize the amplitudes differ by timestamps.<br>\nWe worked hard until the end (we might have seemed quick because we were a team of three!), but I have to admit that your solution was more sophisticated. Anyway, thanks for the reply, and good game!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2104028,
              "author_name": "🐢 Jun Koda",
              "author_url": "",
              "post_date": "2023-01-17T14:33:00.763000",
              "content": "<p>I wrote what the weight is more in detail in the notebook<br>\n<a href=\"https://www.kaggle.com/code/junkoda/optimal-weights-for-signal-to-noise-ratio\" target=\"_blank\">https://www.kaggle.com/code/junkoda/optimal-weights-for-signal-to-noise-ratio</a><br>\nbecause I got similar questions, too. I wanted to brush up a little more, but looks like you already noticed.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2104037,
              "author_name": "knshnb",
              "author_url": "",
              "post_date": "2023-01-17T14:42:25.373000",
              "content": "<p>Thanks so much!!! Let me read it in detail when I have time.<br>\nBy the way, it seems the link to the source code is not public…?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2104615,
              "author_name": "🐢 Jun Koda",
              "author_url": "",
              "post_date": "2023-01-17T23:30:52.520000",
              "content": "<p>Thanks for the feedback. I fixed the github repository to public. I hope it is accessible now (I'll continue to add more explanations).<br>\nI am impressed by the shortness of your code. I'll learn from you, too.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2085299,
      "author_name": "Zollkron",
      "author_url": "",
      "post_date": "2023-01-04T03:51:40.043000",
      "content": "<p>Congratulations Jun. I'm very happy that you be the winner because you worked hard on it from the beginning and generously shared your Basic Spectrogram Classificator Notebook with us. I learnt a lot with you and the others. And all this knowledge did that I could work in a solution based in Isaac Horowitz's QFT (Quantitative Feedback Theory) applied in Industry but with detectors in this case. I though that the uncertainty could be quantified applying this theory but I fell in Bias, so I was wrong. However, this experience was a good try, and I could meet great people. I hope see you all here in future. Thanks for everything 👏👏👏</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2085295,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2023-01-04T03:49:04.627000",
      "content": "<p>Hearty congratulations to <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> for the result!!<br>\nThis was indeed an involved challenge. All the best and regards!!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2087300,
      "author_name": "Suma Mallapragada",
      "author_url": "",
      "post_date": "2023-01-05T13:59:19.480000",
      "content": "<p>Congratulations! Your journey through the competition is amazing. Sinc Kernel approach with power sum makes detection easy.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2086487,
      "author_name": "Muhammad Waseem",
      "author_url": "",
      "post_date": "2023-01-04T20:01:50.863000",
      "content": "<p>congratulation sir. Really happy for you</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2086408,
      "author_name": "Alex Z",
      "author_url": "",
      "post_date": "2023-01-04T18:47:46.993000",
      "content": "<blockquote>\n  <p>Some false positives are remaining in my prediction.</p>\n</blockquote>\n<p>Well, some test data are using real noise background from the detectors. Who knows, probably there is a signal lurking there after all :) </p>\n<p>Congrats on a physical solution!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2085419,
      "author_name": "PiotrKlinke",
      "author_url": "",
      "post_date": "2023-01-04T06:48:38.597000",
      "content": "<p><a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> What other solutions did you use? I am just curious - what are your takeaways from this competition?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2085555,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2023-01-04T08:27:19.040000",
          "content": "<p>For an unsuccessful wave model, I tried to add 48 complex data each day using a template, create 3D block of frequency x frequency time derivative x time, and do 3D CNN, but I need orders of magnitude more templates to add complex wave.</p>\n<p>Take away…, I am much happier with a practical solution than an incomplete ideal solution that didn't meet the deadline; the latter is my situation with the GPS Google Smartphone competition last year.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2085205": "I thank Kaggle and the organizers for hosting this gravitational wave competition. I enjoyed the previous binary back hole merger competition, as well, and was impressed a lot by the gold-medal solutions. I thought large pretrained image models would be the strongest anyway, but they outperformed with 1-dimensional convolutional neural networks. I joined this competition so that I could build such deep neural network models detecting the wave, not the power, ... but failed.\n\nThe largest difference between the two competitions is that the Earth rotates during 120 days and it imprints complicated frequency pattern into the frequency. Adding wave is very delicate, requires very accurate phase patterns, and I was not able to add up the complex Fourier modes effectively within reasonable computational resources. After more than one month without any progress, I thought I should get a silver medal even with an unsatisfactory approach, give up adding the wave and add the power.\n\n# Solution\n\n- Sum power (absolute-value squared) along various signal patterns\n- No machine learning\n- No use of external data or leakage\n\n## Power summation\n\n1. Extract signal frequency and amplitude [total power P(t)] from the simulations\n2. Subtract Doppler shift frequency from the data frequency for 4000 signal patterns \n3. Weight the data proportional to the signal amplitude pattern; this is the optimal linear weight w(t)\n4. Sum the weighted power along lines: 360 frequencies (intercept) × 241 slops in [-120, 120] (frequency bin / 120 days)\n5. Take the maximum\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F1bb4eb2f23c068c7c2998fc8a6e20ef8%2Fpower_sum.png?generation=1672802594937302&alt=media)\n\nThe values are highly skewed from the typical range [-1, 1] because these are maxima of 4000 templates × 360 frequencies × 241 slops.\n\nThis took 5 days using GPU RTX 3090. Full range of slope is [-360, 360] but [-120, 120] already took long enough.\n\n##  Real noise normalization\n\n1. Normalize by the noise rms at each time h -> h / sigma(t)\n2. Remove single-frequency noise by masking anomalously large frequency bin\n3. Normalize the frequency dependence by the remaining rms sigma(f)\n\nI know the noise is not always written as a product of time dependence and frequency dependence, but I did not have time for better treatment. Some false positives are remaining in my prediction.\n\n## Follow up summation with sinc kernel\n\nThe signal spread among frequency bins with sinc function (assuming the window function for the short-time Fourier transform is almost top hat). I use a sinc kernel with width 8 and stride 1/8 frequency bin to collect the signal. This is the optimal linear weighting in the frequency direction. I recompute the power sum with this kernel around the largest-power line in the first step for a subsample of 400 templates. This gives a surprisingly large boost to the public score 0.825 -> 0.848\n\nFinally, I apply a sigmoid to the standardized power sum and submit, which is same as just submitting the power sum. I thought the prediction value must also depend on the noise level; if the signal is undetected, there should be larger possibility to be positive for larger noise because more data are undetected. I modeled this effect but could not improve the score. \n\nPS:\nMore about the weight\nhttps://www.kaggle.com/code/junkoda/optimal-weights-for-signal-to-noise-ratio\n\nCode\nhttps://github.com/junkoda/kaggle_g2net2_solution\n",
    "2104800": "Thanks for posting the detailed solution! It was amazing to note the procedure used in the weights set to the signal-to-noise ratio for improving the model. My solution involved the use of t-test and proximity analysis of signal-to-noise ratios. The notebook is very informative and helped me think on a different line.\n",
    "2086711": "the use of the sinc kernel is superb",
    "2086075": "Congratulations, a great achievement! Your write-up is a good example of being willing to pivot away from initial intuitions about a problem to find a better outcome with a different approach. A good lesson for us all.",
    "2085822": "Congratulations & Thanks for the insightful solution. ",
    "2085726": "@junkoda \nCongratulations for the championship!\nWhat do you mean when you say “slope”?",
    "2085664": "Congratulations on winning the first prize!!! 👍 Very impressive non-machine learning solution based on profound understanding of the competition. I used similar data normalization to reduce the impact of real noise in my CNN model, which is simple but effective.",
    "2085610": "Very cool.  I wonder if GW scientists can use this as a bootstrap step, where this is done first and then they can go back and do parameter estimation.  Doing it all in one go as a template is too costly or poor (pyfstat itself requires very narrow search around signal).",
    "2085587": "Thank you again! I've learned a lot starting from your notebook. And even if it is not impressive, to achieve a 0.73 score was an impossible task at first. At the beginning even 0.6 score was an achievement. Still with your efficient-net approach , 10-20k generated samples a keras single model  and two hours of TPU training the result is amazing for me. ",
    "2085461": "Congratulations, @junkoda! When I watched your score and the number os submissions, I was willing to bet that you are using a ‘physics-based’ method. We also used a very similar approach, but did not account for the amplitude modulation.",
    "2085405": "Seriously - this is gold! Thank you for sharing and for the clear explanation! It shows how being clever is the most important thing. Going back to the basics time and time again and just trying to stop yourself from overengineering. Thank you! 🙏",
    "2085373": "Amazing! Congratulations @junkoda !",
    "2085283": "would you like post your failed deep neural network models attempts ? maybe we can learn from it. ",
    "2085250": "Congrats @🐢 Jun Koda！Very wonderful solution!",
    "2086818": "You are amazing. Congratulations, and thank you for sharing the solution (I like the link from the leaderboard) ",
    "2086319": "This is an amazing solution! Thanks you for sharing your approach! It is a good reminder that machine learning should not always be the go to answer. Congrats on winning @junkoda!",
    "2086181": "Congratulations on your first place!! 🎉\nI was amazed by the performance without any leakage: We took a similar approach (GPU-accelerated random search of signals) but scored 0.835 in private LB without leakage.\n\nThese might be a little obvious since I am not familiar with this field, but can I ask some questions?\n1. What do you mean by signal amplitude (and how did you simulate it)?\n2. In what way weighting proportional to signal amplitude is optimal?",
    "2085299": "Congratulations Jun. I'm very happy that you be the winner because you worked hard on it from the beginning and generously shared your Basic Spectrogram Classificator Notebook with us. I learnt a lot with you and the others. And all this knowledge did that I could work in a solution based in Isaac Horowitz's QFT (Quantitative Feedback Theory) applied in Industry but with detectors in this case. I though that the uncertainty could be quantified applying this theory but I fell in Bias, so I was wrong. However, this experience was a good try, and I could meet great people. I hope see you all here in future. Thanks for everything 👏👏👏",
    "2085295": "Hearty congratulations to @junkoda for the result!!\nThis was indeed an involved challenge. All the best and regards!!",
    "2087300": "Congratulations! Your journey through the competition is amazing. Sinc Kernel approach with power sum makes detection easy.",
    "2086487": "congratulation sir. Really happy for you",
    "2086408": ">Some false positives are remaining in my prediction.\n\nWell, some test data are using real noise background from the detectors. Who knows, probably there is a signal lurking there after all :) \n\nCongrats on a physical solution!",
    "2085419": "@junkoda What other solutions did you use? I am just curious - what are your takeaways from this competition?"
  }
}