{
  "id": 376253,
  "title": "8th place solution: matched filter and CNN",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/376253",
  "author_name": "s_shohei",
  "post_date": "2023-01-05T13:52:32.137000",
  "votes": 13,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Congratulations to the winners, and I would like to express my gratitude to the organizers and my teammate.</p>\n<h1>Summary</h1>\n<ul>\n<li>Matched filter using PyFstat</li>\n<li>CNN with pseudo-label and synthesized data</li>\n<li>Ensemble MF and CNN results.</li>\n</ul>\n<h1>Matched filter</h1>\n<p>Matched filter (MF) is a highly competitive method for searching continuous waves. PyFstat provides a MF module, which we used for this competition. However, PyFstat (and its underlying LALpulsar) requires a specific file structure (.sft file) that is not compatible with hdf5 or ndarray. We converted the provided data to the sft file.  <br>\nWhen provided with four values: f0, f1, alpha, and delta, matched filter calculates a \"twoF\" value, which is the statistical measure of \"likelihood\" of the target wave's existence, </p>\n<ul>\n<li>We used PyFstat's <a href=\"https://pyfstat.readthedocs.io/en/latest/pyfstat.html#pyfstat.core.SemiCoherentSearch\" target=\"_blank\">SemiCoherentSearch</a> with \"nsegs=1000\".  <br>\nHere, we used relatively large \"nsegs\" value to smooth the function shape of \"twoF\" w.r.t. parameters f0/f1/alpha/delta. This trick helps find the existence of target waves in much less grid search trials.</li>\n<li>We conducted a grid search for alpha and delta.</li>\n<li>We used Optuna to search for f0 and f1 in order to maximize the towF value.</li>\n<li>We performed around 600 Optuna explorations per sample.</li>\n<li>You can find more information about our implementation here:   <br>\n<a href=\"https://www.kaggle.com/code/iiyamaiiyama/g2net-pyfstat-matched-filter\" target=\"_blank\">https://www.kaggle.com/code/iiyamaiiyama/g2net-pyfstat-matched-filter</a></li>\n</ul>\n<h2>CV result</h2>\n<p>The matched filter histograms for the training data (600 samples) are shown. The right image is an enlarged version of the left image. As you can see, if MF result(twoF) is greater than 4600, the precision is 100%. Therefore we set 4600 as the threshold value.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2304617%2Fe87236cd94cbd61d554ccf9fca31a494%2Fsol1.png?generation=1672925392581375&amp;alt=media\" alt=\"\"></p>\n<p>The AUC of the 600 samples is 0.8406.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2304617%2Ff4b47e343fa5b288fc78ae0eaa4d89a2%2Fsol2.png?generation=1672925412196942&amp;alt=media\" alt=\"\"></p>\n<h2>LB result</h2>\n<p>If we submitted only the MF results, the private/public leaderboard score was 0.764/0.760, which is not a very competitive result. We believe it may be due to the non-stationary noise. Therefore, we decided to ensemble MF and CNN results to improve our score.</p>\n<h1>CNN</h1>\n<h2>dataset</h2>\n<p>The training data only contains 600 samples. We had to deal with it.</p>\n<h3>Stationary noise</h3>\n<ul>\n<li>Pseudo-label using matched filter<br>\nWe selected samples with MF results greater than 4600 from the test set. Around 1700 samples.</li>\n<li>Synthesized data<br>\nWe synthesized new data from train1 and train2 by using the following formula:  <br>\n<code>new_data = train1 * alpha + train2 * beta</code>  <br>\n(alpha and beta satisfy the condition <code>alpha^2 + beta^2 == 1</code>)</li>\n</ul>\n<h3>Non-stationary noise</h3>\n<p>We used the following notebook as a reference for creating our own new data.   <br>\n<a href=\"https://www.kaggle.com/code/vslaykovsky/g2net-pytorch-generated-realistic-noise\" target=\"_blank\">https://www.kaggle.com/code/vslaykovsky/g2net-pytorch-generated-realistic-noise</a>  <br>\nAfter creating the new data, we augmented the data by synthesizing the two data as well as the stationary noise data.</p>\n<h2>Model</h2>\n<p>We had two separate models: one for stationary noise and one for non-stationary noise. The models themselves were ordinary CNNs. We used data augmentation techniques such as h/v flip, mask and roll.</p>\n<h1>Ensemble</h1>\n<p>Finally, we ensemble MF and CNN results. For samples with MF result greater than 4600, we assigned a score of 0.9-1.0. For all other samples, we used the CNN scores of 0.0-0.9. By Prioritizing samples with 100% precision and placing less confident CNN results afterwards, we can maximize the AUC score.</p>",
  "messages": [
    {
      "id": 2087289,
      "postDate": "2023-01-05T13:52:32.137Z",
      "content": "<p>Congratulations to the winners, and I would like to express my gratitude to the organizers and my teammate.</p>\n<h1>Summary</h1>\n<ul>\n<li>Matched filter using PyFstat</li>\n<li>CNN with pseudo-label and synthesized data</li>\n<li>Ensemble MF and CNN results.</li>\n</ul>\n<h1>Matched filter</h1>\n<p>Matched filter (MF) is a highly competitive method for searching continuous waves. PyFstat provides a MF module, which we used for this competition. However, PyFstat (and its underlying LALpulsar) requires a specific file structure (.sft file) that is not compatible with hdf5 or ndarray. We converted the provided data to the sft file.  <br>\nWhen provided with four values: f0, f1, alpha, and delta, matched filter calculates a \"twoF\" value, which is the statistical measure of \"likelihood\" of the target wave's existence, </p>\n<ul>\n<li>We used PyFstat's <a href=\"https://pyfstat.readthedocs.io/en/latest/pyfstat.html#pyfstat.core.SemiCoherentSearch\" target=\"_blank\">SemiCoherentSearch</a> with \"nsegs=1000\".  <br>\nHere, we used relatively large \"nsegs\" value to smooth the function shape of \"twoF\" w.r.t. parameters f0/f1/alpha/delta. This trick helps find the existence of target waves in much less grid search trials.</li>\n<li>We conducted a grid search for alpha and delta.</li>\n<li>We used Optuna to search for f0 and f1 in order to maximize the towF value.</li>\n<li>We performed around 600 Optuna explorations per sample.</li>\n<li>You can find more information about our implementation here:   <br>\n<a href=\"https://www.kaggle.com/code/iiyamaiiyama/g2net-pyfstat-matched-filter\" target=\"_blank\">https://www.kaggle.com/code/iiyamaiiyama/g2net-pyfstat-matched-filter</a></li>\n</ul>\n<h2>CV result</h2>\n<p>The matched filter histograms for the training data (600 samples) are shown. The right image is an enlarged version of the left image. As you can see, if MF result(twoF) is greater than 4600, the precision is 100%. Therefore we set 4600 as the threshold value.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2304617%2Fe87236cd94cbd61d554ccf9fca31a494%2Fsol1.png?generation=1672925392581375&amp;alt=media\" alt=\"\"></p>\n<p>The AUC of the 600 samples is 0.8406.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2304617%2Ff4b47e343fa5b288fc78ae0eaa4d89a2%2Fsol2.png?generation=1672925412196942&amp;alt=media\" alt=\"\"></p>\n<h2>LB result</h2>\n<p>If we submitted only the MF results, the private/public leaderboard score was 0.764/0.760, which is not a very competitive result. We believe it may be due to the non-stationary noise. Therefore, we decided to ensemble MF and CNN results to improve our score.</p>\n<h1>CNN</h1>\n<h2>dataset</h2>\n<p>The training data only contains 600 samples. We had to deal with it.</p>\n<h3>Stationary noise</h3>\n<ul>\n<li>Pseudo-label using matched filter<br>\nWe selected samples with MF results greater than 4600 from the test set. Around 1700 samples.</li>\n<li>Synthesized data<br>\nWe synthesized new data from train1 and train2 by using the following formula:  <br>\n<code>new_data = train1 * alpha + train2 * beta</code>  <br>\n(alpha and beta satisfy the condition <code>alpha^2 + beta^2 == 1</code>)</li>\n</ul>\n<h3>Non-stationary noise</h3>\n<p>We used the following notebook as a reference for creating our own new data.   <br>\n<a href=\"https://www.kaggle.com/code/vslaykovsky/g2net-pytorch-generated-realistic-noise\" target=\"_blank\">https://www.kaggle.com/code/vslaykovsky/g2net-pytorch-generated-realistic-noise</a>  <br>\nAfter creating the new data, we augmented the data by synthesizing the two data as well as the stationary noise data.</p>\n<h2>Model</h2>\n<p>We had two separate models: one for stationary noise and one for non-stationary noise. The models themselves were ordinary CNNs. We used data augmentation techniques such as h/v flip, mask and roll.</p>\n<h1>Ensemble</h1>\n<p>Finally, we ensemble MF and CNN results. For samples with MF result greater than 4600, we assigned a score of 0.9-1.0. For all other samples, we used the CNN scores of 0.0-0.9. By Prioritizing samples with 100% precision and placing less confident CNN results afterwards, we can maximize the AUC score.</p>",
      "rawMarkdown": "Congratulations to the winners, and I would like to express my gratitude to the organizers and my teammate.\n\n# Summary\n* Matched filter using PyFstat\n* CNN with pseudo-label and synthesized data\n* Ensemble MF and CNN results.\n\n# Matched filter\nMatched filter (MF) is a highly competitive method for searching continuous waves. PyFstat provides a MF module, which we used for this competition. However, PyFstat (and its underlying LALpulsar) requires a specific file structure (.sft file) that is not compatible with hdf5 or ndarray. We converted the provided data to the sft file.  \nWhen provided with four values: f0, f1, alpha, and delta, matched filter calculates a \"twoF\" value, which is the statistical measure of \"likelihood\" of the target wave's existence, \n\n* We used PyFstat's [SemiCoherentSearch](https://pyfstat.readthedocs.io/en/latest/pyfstat.html#pyfstat.core.SemiCoherentSearch) with \"nsegs=1000\".  \n  Here, we used relatively large \"nsegs\" value to smooth the function shape of \"twoF\" w.r.t. parameters f0/f1/alpha/delta. This trick helps find the existence of target waves in much less grid search trials.\n* We conducted a grid search for alpha and delta.\n* We used Optuna to search for f0 and f1 in order to maximize the towF value.\n* We performed around 600 Optuna explorations per sample.\n* You can find more information about our implementation here:   \n  https://www.kaggle.com/code/iiyamaiiyama/g2net-pyfstat-matched-filter\n\n## CV result\nThe matched filter histograms for the training data (600 samples) are shown. The right image is an enlarged version of the left image. As you can see, if MF result(twoF) is greater than 4600, the precision is 100%. Therefore we set 4600 as the threshold value.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2304617%2Fe87236cd94cbd61d554ccf9fca31a494%2Fsol1.png?generation=1672925392581375&alt=media )\n\n\nThe AUC of the 600 samples is 0.8406.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2304617%2Ff4b47e343fa5b288fc78ae0eaa4d89a2%2Fsol2.png?generation=1672925412196942&alt=media =332x221)\n\n## LB result\nIf we submitted only the MF results, the private/public leaderboard score was 0.764/0.760, which is not a very competitive result. We believe it may be due to the non-stationary noise. Therefore, we decided to ensemble MF and CNN results to improve our score.\n\n# CNN\n## dataset\nThe training data only contains 600 samples. We had to deal with it.\n\n### Stationary noise\n* Pseudo-label using matched filter\nWe selected samples with MF results greater than 4600 from the test set. Around 1700 samples.\n* Synthesized data\nWe synthesized new data from train1 and train2 by using the following formula:  \n  `new_data = train1 * alpha + train2 * beta`  \n  (alpha and beta satisfy the condition `alpha^2 + beta^2 == 1`)\n\n### Non-stationary noise\nWe used the following notebook as a reference for creating our own new data.   \nhttps://www.kaggle.com/code/vslaykovsky/g2net-pytorch-generated-realistic-noise  \nAfter creating the new data, we augmented the data by synthesizing the two data as well as the stationary noise data.\n\n## Model\nWe had two separate models: one for stationary noise and one for non-stationary noise. The models themselves were ordinary CNNs. We used data augmentation techniques such as h/v flip, mask and roll.\n\n# Ensemble\nFinally, we ensemble MF and CNN results. For samples with MF result greater than 4600, we assigned a score of 0.9-1.0. For all other samples, we used the CNN scores of 0.0-0.9. By Prioritizing samples with 100% precision and placing less confident CNN results afterwards, we can maximize the AUC score.",
      "votes": 13
    },
    {
      "id": 2088381,
      "postDate": "2023-01-06T09:23:57.490Z",
      "content": "<p>Congratulations with the gold medal, very nice approach!  I can see that the kaggle notebook takes 3 hours to run. How long it takes you to analyse the test set?</p>",
      "rawMarkdown": "Congratulations with the gold medal, very nice approach!  I can see that the kaggle notebook takes 3 hours to run. How long it takes you to analyse the test set?",
      "votes": 1,
      "replies": [
        {
          "id": 2088533,
          "postDate": "2023-01-06T12:03:55.990Z",
          "content": "<p>Thank you! Congratulations to you too.</p>\n<p>To analyze the test set, we spent a maximum of 10 hours per sample. If the result was greater than 4600, we stopped the search for that sample.<br>\nIt took around two weeks to run the entire test set in real time using 96 vCPU machine.</p>",
          "rawMarkdown": "Thank you! Congratulations to you too.\n\nTo analyze the test set, we spent a maximum of 10 hours per sample. If the result was greater than 4600, we stopped the search for that sample.\nIt took around two weeks to run the entire test set in real time using 96 vCPU machine.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2087357,
      "postDate": "2023-01-05T14:48:06.047Z",
      "content": "<p>Congratulations on your gold medal!! Awesomes_shohei! Thank you for sharing. It's very helpful for us to learn.</p>",
      "rawMarkdown": "Congratulations on your gold medal!! Awesomes_shohei! Thank you for sharing. It's very helpful for us to learn.",
      "votes": 1
    },
    {
      "id": 2087598,
      "postDate": "2023-01-05T17:44:12.113Z",
      "content": "<p>Very cool.  I tried to use pyfstat and I also hacked their code to generate sft's from the given hdf5 files.  However, I was using their MCMC and grid search codes and it never worked except for their simple test case when the search parameters were very close to the signal.  I didn't see this other semi-coherent search at the time, but this was the last day of the competition so I didn't have much time to go through more. </p>\n<p>That is, after making sfts, then I was using the 2nd half of codes like:<br>\n<a href=\"https://pyfstat.readthedocs.io/en/latest/mcmc_examples/index.html\" target=\"_blank\">https://pyfstat.readthedocs.io/en/latest/mcmc_examples/index.html</a><br>\n<a href=\"https://pyfstat.readthedocs.io/en/latest/mcmc_vs_grid_simple_example/index.html\" target=\"_blank\">https://pyfstat.readthedocs.io/en/latest/mcmc_vs_grid_simple_example/index.html</a><br>\netc.</p>\n<p>But even working around their band issues (I had to expand the data to band=1 to avoid boundary effects in the search (as they mention in the examples), even if I narrowed the search right around the frequency of the signal, it couldn't find it.</p>\n<p>Are you saying all I had to do was use the \"semi-coherent\" version of their code and it would have worked?  darn :)</p>\n<p>FYI here's example codes:<br>\n<a href=\"https://www.kaggle.com/pseudotensor/hdf5-to-sft\" target=\"_blank\">https://www.kaggle.com/pseudotensor/hdf5-to-sft</a><br>\n<a href=\"https://www.kaggle.com/pseudotensor/mcmcsearch-pyfstat\" target=\"_blank\">https://www.kaggle.com/pseudotensor/mcmcsearch-pyfstat</a></p>\n<p>If you can see or explain where I went wrong or if it was just about using semi-coherent, I'd be thankful :)</p>",
      "rawMarkdown": "Very cool.  I tried to use pyfstat and I also hacked their code to generate sft's from the given hdf5 files.  However, I was using their MCMC and grid search codes and it never worked except for their simple test case when the search parameters were very close to the signal.  I didn't see this other semi-coherent search at the time, but this was the last day of the competition so I didn't have much time to go through more. \n\nThat is, after making sfts, then I was using the 2nd half of codes like:\nhttps://pyfstat.readthedocs.io/en/latest/mcmc_examples/index.html\nhttps://pyfstat.readthedocs.io/en/latest/mcmc_vs_grid_simple_example/index.html\netc.\n\nBut even working around their band issues (I had to expand the data to band=1 to avoid boundary effects in the search (as they mention in the examples), even if I narrowed the search right around the frequency of the signal, it couldn't find it.\n\nAre you saying all I had to do was use the \"semi-coherent\" version of their code and it would have worked?  darn :)\n\nFYI here's example codes:\nhttps://www.kaggle.com/pseudotensor/hdf5-to-sft\nhttps://www.kaggle.com/pseudotensor/mcmcsearch-pyfstat\n\nIf you can see or explain where I went wrong or if it was just about using semi-coherent, I'd be thankful :)",
      "votes": 2,
      "replies": [
        {
          "id": 2087977,
          "postDate": "2023-01-06T00:39:18.363Z",
          "content": "<p>Thank you.</p>\n<p>Initially, we tried a fully-coherent search. However, as you pointed out, fully-coherent search requires near-perfect (e.g. 99.99999%) parameter estimation to achieve a \"spiked\" result. Therefore, we decided to use a semi-coherent search with a large nsegs.</p>\n<p>You can find a comparison of semi and fully coherent search results on page 4 of the following slide by Rodrigo Tenorio, host of this competition:<br>\n<a href=\"https://www.uv.es/igwm2021/slides/Rodrigo_Tenorio.pdf\" target=\"_blank\">https://www.uv.es/igwm2021/slides/Rodrigo_Tenorio.pdf</a></p>\n<p>We also tried using the MCMCSemiCoherentSearch method, but it was not as competitive as a normal semi-coherent search.</p>",
          "rawMarkdown": "Thank you.\n\nInitially, we tried a fully-coherent search. However, as you pointed out, fully-coherent search requires near-perfect (e.g. 99.99999%) parameter estimation to achieve a \"spiked\" result. Therefore, we decided to use a semi-coherent search with a large nsegs.\n\nYou can find a comparison of semi and fully coherent search results on page 4 of the following slide by Rodrigo Tenorio, host of this competition:\nhttps://www.uv.es/igwm2021/slides/Rodrigo_Tenorio.pdf\n\nWe also tried using the MCMCSemiCoherentSearch method, but it was not as competitive as a normal semi-coherent search.\n",
          "votes": 2,
          "replies": [
            {
              "id": 2087986,
              "postDate": "2023-01-06T00:50:32.967Z",
              "content": "<p>Thanks you so much for explaining :)</p>",
              "rawMarkdown": "Thanks you so much for explaining :)",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2088381,
      "author_name": "Inar Timiryasov",
      "author_url": "",
      "post_date": "2023-01-06T09:23:57.490000",
      "content": "<p>Congratulations with the gold medal, very nice approach!  I can see that the kaggle notebook takes 3 hours to run. How long it takes you to analyse the test set?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2088533,
          "author_name": "s_shohei",
          "author_url": "",
          "post_date": "2023-01-06T12:03:55.990000",
          "content": "<p>Thank you! Congratulations to you too.</p>\n<p>To analyze the test set, we spent a maximum of 10 hours per sample. If the result was greater than 4600, we stopped the search for that sample.<br>\nIt took around two weeks to run the entire test set in real time using 96 vCPU machine.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2087357,
      "author_name": "BarryZhou",
      "author_url": "",
      "post_date": "2023-01-05T14:48:06.047000",
      "content": "<p>Congratulations on your gold medal!! Awesomes_shohei! Thank you for sharing. It's very helpful for us to learn.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2087598,
      "author_name": "Jonathan McKinney",
      "author_url": "",
      "post_date": "2023-01-05T17:44:12.113000",
      "content": "<p>Very cool.  I tried to use pyfstat and I also hacked their code to generate sft's from the given hdf5 files.  However, I was using their MCMC and grid search codes and it never worked except for their simple test case when the search parameters were very close to the signal.  I didn't see this other semi-coherent search at the time, but this was the last day of the competition so I didn't have much time to go through more. </p>\n<p>That is, after making sfts, then I was using the 2nd half of codes like:<br>\n<a href=\"https://pyfstat.readthedocs.io/en/latest/mcmc_examples/index.html\" target=\"_blank\">https://pyfstat.readthedocs.io/en/latest/mcmc_examples/index.html</a><br>\n<a href=\"https://pyfstat.readthedocs.io/en/latest/mcmc_vs_grid_simple_example/index.html\" target=\"_blank\">https://pyfstat.readthedocs.io/en/latest/mcmc_vs_grid_simple_example/index.html</a><br>\netc.</p>\n<p>But even working around their band issues (I had to expand the data to band=1 to avoid boundary effects in the search (as they mention in the examples), even if I narrowed the search right around the frequency of the signal, it couldn't find it.</p>\n<p>Are you saying all I had to do was use the \"semi-coherent\" version of their code and it would have worked?  darn :)</p>\n<p>FYI here's example codes:<br>\n<a href=\"https://www.kaggle.com/pseudotensor/hdf5-to-sft\" target=\"_blank\">https://www.kaggle.com/pseudotensor/hdf5-to-sft</a><br>\n<a href=\"https://www.kaggle.com/pseudotensor/mcmcsearch-pyfstat\" target=\"_blank\">https://www.kaggle.com/pseudotensor/mcmcsearch-pyfstat</a></p>\n<p>If you can see or explain where I went wrong or if it was just about using semi-coherent, I'd be thankful :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2087977,
          "author_name": "s_shohei",
          "author_url": "",
          "post_date": "2023-01-06T00:39:18.363000",
          "content": "<p>Thank you.</p>\n<p>Initially, we tried a fully-coherent search. However, as you pointed out, fully-coherent search requires near-perfect (e.g. 99.99999%) parameter estimation to achieve a \"spiked\" result. Therefore, we decided to use a semi-coherent search with a large nsegs.</p>\n<p>You can find a comparison of semi and fully coherent search results on page 4 of the following slide by Rodrigo Tenorio, host of this competition:<br>\n<a href=\"https://www.uv.es/igwm2021/slides/Rodrigo_Tenorio.pdf\" target=\"_blank\">https://www.uv.es/igwm2021/slides/Rodrigo_Tenorio.pdf</a></p>\n<p>We also tried using the MCMCSemiCoherentSearch method, but it was not as competitive as a normal semi-coherent search.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2087986,
              "author_name": "Jonathan McKinney",
              "author_url": "",
              "post_date": "2023-01-06T00:50:32.967000",
              "content": "<p>Thanks you so much for explaining :)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2087289": "Congratulations to the winners, and I would like to express my gratitude to the organizers and my teammate.\n\n# Summary\n* Matched filter using PyFstat\n* CNN with pseudo-label and synthesized data\n* Ensemble MF and CNN results.\n\n# Matched filter\nMatched filter (MF) is a highly competitive method for searching continuous waves. PyFstat provides a MF module, which we used for this competition. However, PyFstat (and its underlying LALpulsar) requires a specific file structure (.sft file) that is not compatible with hdf5 or ndarray. We converted the provided data to the sft file.  \nWhen provided with four values: f0, f1, alpha, and delta, matched filter calculates a \"twoF\" value, which is the statistical measure of \"likelihood\" of the target wave's existence, \n\n* We used PyFstat's [SemiCoherentSearch](https://pyfstat.readthedocs.io/en/latest/pyfstat.html#pyfstat.core.SemiCoherentSearch) with \"nsegs=1000\".  \n  Here, we used relatively large \"nsegs\" value to smooth the function shape of \"twoF\" w.r.t. parameters f0/f1/alpha/delta. This trick helps find the existence of target waves in much less grid search trials.\n* We conducted a grid search for alpha and delta.\n* We used Optuna to search for f0 and f1 in order to maximize the towF value.\n* We performed around 600 Optuna explorations per sample.\n* You can find more information about our implementation here:   \n  https://www.kaggle.com/code/iiyamaiiyama/g2net-pyfstat-matched-filter\n\n## CV result\nThe matched filter histograms for the training data (600 samples) are shown. The right image is an enlarged version of the left image. As you can see, if MF result(twoF) is greater than 4600, the precision is 100%. Therefore we set 4600 as the threshold value.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2304617%2Fe87236cd94cbd61d554ccf9fca31a494%2Fsol1.png?generation=1672925392581375&alt=media )\n\n\nThe AUC of the 600 samples is 0.8406.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2304617%2Ff4b47e343fa5b288fc78ae0eaa4d89a2%2Fsol2.png?generation=1672925412196942&alt=media =332x221)\n\n## LB result\nIf we submitted only the MF results, the private/public leaderboard score was 0.764/0.760, which is not a very competitive result. We believe it may be due to the non-stationary noise. Therefore, we decided to ensemble MF and CNN results to improve our score.\n\n# CNN\n## dataset\nThe training data only contains 600 samples. We had to deal with it.\n\n### Stationary noise\n* Pseudo-label using matched filter\nWe selected samples with MF results greater than 4600 from the test set. Around 1700 samples.\n* Synthesized data\nWe synthesized new data from train1 and train2 by using the following formula:  \n  `new_data = train1 * alpha + train2 * beta`  \n  (alpha and beta satisfy the condition `alpha^2 + beta^2 == 1`)\n\n### Non-stationary noise\nWe used the following notebook as a reference for creating our own new data.   \nhttps://www.kaggle.com/code/vslaykovsky/g2net-pytorch-generated-realistic-noise  \nAfter creating the new data, we augmented the data by synthesizing the two data as well as the stationary noise data.\n\n## Model\nWe had two separate models: one for stationary noise and one for non-stationary noise. The models themselves were ordinary CNNs. We used data augmentation techniques such as h/v flip, mask and roll.\n\n# Ensemble\nFinally, we ensemble MF and CNN results. For samples with MF result greater than 4600, we assigned a score of 0.9-1.0. For all other samples, we used the CNN scores of 0.0-0.9. By Prioritizing samples with 100% precision and placing less confident CNN results afterwards, we can maximize the AUC score.",
    "2088381": "Congratulations with the gold medal, very nice approach!  I can see that the kaggle notebook takes 3 hours to run. How long it takes you to analyse the test set?",
    "2087357": "Congratulations on your gold medal!! Awesomes_shohei! Thank you for sharing. It's very helpful for us to learn.",
    "2087598": "Very cool.  I tried to use pyfstat and I also hacked their code to generate sft's from the given hdf5 files.  However, I was using their MCMC and grid search codes and it never worked except for their simple test case when the search parameters were very close to the signal.  I didn't see this other semi-coherent search at the time, but this was the last day of the competition so I didn't have much time to go through more. \n\nThat is, after making sfts, then I was using the 2nd half of codes like:\nhttps://pyfstat.readthedocs.io/en/latest/mcmc_examples/index.html\nhttps://pyfstat.readthedocs.io/en/latest/mcmc_vs_grid_simple_example/index.html\netc.\n\nBut even working around their band issues (I had to expand the data to band=1 to avoid boundary effects in the search (as they mention in the examples), even if I narrowed the search right around the frequency of the signal, it couldn't find it.\n\nAre you saying all I had to do was use the \"semi-coherent\" version of their code and it would have worked?  darn :)\n\nFYI here's example codes:\nhttps://www.kaggle.com/pseudotensor/hdf5-to-sft\nhttps://www.kaggle.com/pseudotensor/mcmcsearch-pyfstat\n\nIf you can see or explain where I went wrong or if it was just about using semi-coherent, I'd be thankful :)"
  }
}