{
  "id": 375923,
  "title": "6th place solution (Simulated Annealing Approach)",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/375923",
  "author_name": "Shun_PI",
  "post_date": "2023-01-04T03:05:59.808000",
  "votes": 41,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, thanks to the host and congratulations to the winners.<br>\nI'm so glad to win my first solo gold medal in this competition!<br>\nMy solution is based on <a href=\"https://en.wikipedia.org/wiki/Simulated_annealing\" target=\"_blank\">Simulated annealing</a> (SA), a technique often used in optimization problems. It is NOT based on machine learning/deep learning/data generation.</p>\n<h3>Preprocess</h3>\n<p>The following process was performed for each of H1 and L1 and added together to produce a 360*256 image.</p>\n<ul>\n<li><p>Calculate the square of the amplitude</p></li>\n<li><p>Removal of horizontal line noise (only real noise)</p>\n<ul>\n<li>Calculate the sum for each row and create a list of 360 elements. Then take a difference of this list. If there is a series of large plus value and large minus value with short intervals, it is assumed that there is horizontal line noise. Then the values of these rows is replaced by the average of the each column. This operation is repeated until it can no longer be done.</li>\n<li>This removes almost all of the real noise and improves the score by about 0.005.<ul>\n<li>Some example images before and after horizontal line noise removal are as follows.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fdd6cf77558775aabeaf6f8692627cae9%2F1-2.png?generation=1672801107095404&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fc5a3f0d84fd30bdfc21ad2f112c7adcc%2F3-4.png?generation=1672801151754877&amp;alt=media\" alt=\"\"></li></ul></li></ul></li>\n<li><p>Removal of vertical stripe noise</p>\n<ul>\n<li>Subtract mean for each column, divide by std^2, and correct so that the sum of std of all columns is a constant.</li>\n<li>By dividing by the square of std instead of std, the columns with larger original variance will have smaller final variance. The idea is to reduce the contribution of columns with large original variance because they cannot be relied upon. The score increases slightly by this method.</li></ul></li>\n<li><p>Filling the time gap</p>\n<ul>\n<li>Calculate the difference of the timestamps as d and repeat inserting zero-filled columns round(d / 1800 - 1) times.</li></ul></li>\n<li><p>Average in the time direction and reduce columns size to 256</p></li>\n</ul>\n<h3>Wave Detection with SA</h3>\n<p>I will explain the essential part of my solution: Simulated annealing (SA).<br>\nFirst, assume that the signal is perfectly sinusoidal and follows the following equation</p>\n<p>f(x) = Asin(vx+θ) + h</p>\n<p>where A is the amplitude, v is the frequency, θ is the phase, and h is the position in the frequency direction.<br>\nThe search range for each parameter is set as follows based on the experimental results.</p>\n<ul>\n<li>A: [-100, 100].</li>\n<li>v: fixed at 0.31 (fixing the frequency slightly improved the score)</li>\n<li>θ: any real number</li>\n<li>h: any real number</li>\n</ul>\n<p>Then we want to search the parameters that gives the clearest waveform.<br>\nTo do this, I used SA, as following procedures.</p>\n<ul>\n<li>The following operation is repeated tens of thousands of times with decreasing the temperature.<ul>\n<li>Changes the above parameters slightly, and calculates the sum of the squares of the amplitudes along the waveform as the score (penalty if it refers outside the image). If the score is greater than the last score, accepts the change. Even if the score is smaller than the last score,  accepts the change probabilistically, depending on the score difference and the annealing temperature. Otherwise rejects the change. </li></ul></li>\n<li>In addition, to prevent getting a score of local optimum, the above whole process is repeated about several hundred times, and the kth percentile value (k chosen from around 90~100) is used as the final score.</li>\n<li>It takes several hours to a day (depending on the number of iterations) on my CPU (Core i9-9900K) to calculate the scores of all the test data.<ul>\n<li>Postprocessing and submitting this can achieve LB score above 0.79 ,with a few hours calculation.</li></ul></li>\n</ul>\n<p>Some example images of detected waveforms for some difficult test cases are as follows.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2F664a106bc0a7a3a394bc78c55a1d80f1%2F5.png?generation=1672801516897133&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Feaa5412252d039795d351203575fd2d1%2F6.png?generation=1672801527565415&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2F344d212d77257fd883c702ea2bfd7780%2F7.png?generation=1672801537224085&amp;alt=media\" alt=\"\"></p>\n<h3>Postprocess</h3>\n<p>I performed a number of SAs with various numbers of iterations, parameter search ranges, wave thickness, etc., and calculated the ensemble score by weighted averaging.<br>\nI gave up on calculating the CV score, and adjusted the weights so that Public LB would be higher.<br>\nIn addition, the real noise data tended to score slightly higher than the generated noise data (due to its tendency to overfit against noise). Therefore, I first converted the final score to a rank, and then divided the rank value by about 1.1 for real noise data only, improving the score by several points.</p>",
  "messages": [
    {
      "id": 2085264,
      "postDate": "2023-01-04T03:05:59.807Z",
      "content": "<p>First of all, thanks to the host and congratulations to the winners.<br>\nI'm so glad to win my first solo gold medal in this competition!<br>\nMy solution is based on <a href=\"https://en.wikipedia.org/wiki/Simulated_annealing\" target=\"_blank\">Simulated annealing</a> (SA), a technique often used in optimization problems. It is NOT based on machine learning/deep learning/data generation.</p>\n<h3>Preprocess</h3>\n<p>The following process was performed for each of H1 and L1 and added together to produce a 360*256 image.</p>\n<ul>\n<li><p>Calculate the square of the amplitude</p></li>\n<li><p>Removal of horizontal line noise (only real noise)</p>\n<ul>\n<li>Calculate the sum for each row and create a list of 360 elements. Then take a difference of this list. If there is a series of large plus value and large minus value with short intervals, it is assumed that there is horizontal line noise. Then the values of these rows is replaced by the average of the each column. This operation is repeated until it can no longer be done.</li>\n<li>This removes almost all of the real noise and improves the score by about 0.005.<ul>\n<li>Some example images before and after horizontal line noise removal are as follows.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fdd6cf77558775aabeaf6f8692627cae9%2F1-2.png?generation=1672801107095404&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fc5a3f0d84fd30bdfc21ad2f112c7adcc%2F3-4.png?generation=1672801151754877&amp;alt=media\" alt=\"\"></li></ul></li></ul></li>\n<li><p>Removal of vertical stripe noise</p>\n<ul>\n<li>Subtract mean for each column, divide by std^2, and correct so that the sum of std of all columns is a constant.</li>\n<li>By dividing by the square of std instead of std, the columns with larger original variance will have smaller final variance. The idea is to reduce the contribution of columns with large original variance because they cannot be relied upon. The score increases slightly by this method.</li></ul></li>\n<li><p>Filling the time gap</p>\n<ul>\n<li>Calculate the difference of the timestamps as d and repeat inserting zero-filled columns round(d / 1800 - 1) times.</li></ul></li>\n<li><p>Average in the time direction and reduce columns size to 256</p></li>\n</ul>\n<h3>Wave Detection with SA</h3>\n<p>I will explain the essential part of my solution: Simulated annealing (SA).<br>\nFirst, assume that the signal is perfectly sinusoidal and follows the following equation</p>\n<p>f(x) = Asin(vx+θ) + h</p>\n<p>where A is the amplitude, v is the frequency, θ is the phase, and h is the position in the frequency direction.<br>\nThe search range for each parameter is set as follows based on the experimental results.</p>\n<ul>\n<li>A: [-100, 100].</li>\n<li>v: fixed at 0.31 (fixing the frequency slightly improved the score)</li>\n<li>θ: any real number</li>\n<li>h: any real number</li>\n</ul>\n<p>Then we want to search the parameters that gives the clearest waveform.<br>\nTo do this, I used SA, as following procedures.</p>\n<ul>\n<li>The following operation is repeated tens of thousands of times with decreasing the temperature.<ul>\n<li>Changes the above parameters slightly, and calculates the sum of the squares of the amplitudes along the waveform as the score (penalty if it refers outside the image). If the score is greater than the last score, accepts the change. Even if the score is smaller than the last score,  accepts the change probabilistically, depending on the score difference and the annealing temperature. Otherwise rejects the change. </li></ul></li>\n<li>In addition, to prevent getting a score of local optimum, the above whole process is repeated about several hundred times, and the kth percentile value (k chosen from around 90~100) is used as the final score.</li>\n<li>It takes several hours to a day (depending on the number of iterations) on my CPU (Core i9-9900K) to calculate the scores of all the test data.<ul>\n<li>Postprocessing and submitting this can achieve LB score above 0.79 ,with a few hours calculation.</li></ul></li>\n</ul>\n<p>Some example images of detected waveforms for some difficult test cases are as follows.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2F664a106bc0a7a3a394bc78c55a1d80f1%2F5.png?generation=1672801516897133&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Feaa5412252d039795d351203575fd2d1%2F6.png?generation=1672801527565415&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2F344d212d77257fd883c702ea2bfd7780%2F7.png?generation=1672801537224085&amp;alt=media\" alt=\"\"></p>\n<h3>Postprocess</h3>\n<p>I performed a number of SAs with various numbers of iterations, parameter search ranges, wave thickness, etc., and calculated the ensemble score by weighted averaging.<br>\nI gave up on calculating the CV score, and adjusted the weights so that Public LB would be higher.<br>\nIn addition, the real noise data tended to score slightly higher than the generated noise data (due to its tendency to overfit against noise). Therefore, I first converted the final score to a rank, and then divided the rank value by about 1.1 for real noise data only, improving the score by several points.</p>",
      "rawMarkdown": "First of all, thanks to the host and congratulations to the winners.\nI'm so glad to win my first solo gold medal in this competition!\nMy solution is based on [Simulated annealing](https://en.wikipedia.org/wiki/Simulated_annealing) (SA), a technique often used in optimization problems. It is NOT based on machine learning/deep learning/data generation.\n\n### Preprocess\n\nThe following process was performed for each of H1 and L1 and added together to produce a 360*256 image.\n- Calculate the square of the amplitude\n- Removal of horizontal line noise (only real noise)\n\t- Calculate the sum for each row and create a list of 360 elements. Then take a difference of this list. If there is a series of large plus value and large minus value with short intervals, it is assumed that there is horizontal line noise. Then the values of these rows is replaced by the average of the each column. This operation is repeated until it can no longer be done.\n\t- This removes almost all of the real noise and improves the score by about 0.005.\n        - Some example images before and after horizontal line noise removal are as follows.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fdd6cf77558775aabeaf6f8692627cae9%2F1-2.png?generation=1672801107095404&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fc5a3f0d84fd30bdfc21ad2f112c7adcc%2F3-4.png?generation=1672801151754877&alt=media)\n\n- Removal of vertical stripe noise\n\t- Subtract mean for each column, divide by std^2, and correct so that the sum of std of all columns is a constant.\n\t- By dividing by the square of std instead of std, the columns with larger original variance will have smaller final variance. The idea is to reduce the contribution of columns with large original variance because they cannot be relied upon. The score increases slightly by this method.\n- Filling the time gap\n\t- Calculate the difference of the timestamps as d and repeat inserting zero-filled columns round(d / 1800 - 1) times.\n- Average in the time direction and reduce columns size to 256\n\n### Wave Detection with SA\n\nI will explain the essential part of my solution: Simulated annealing (SA).\nFirst, assume that the signal is perfectly sinusoidal and follows the following equation\n\nf(x) = Asin(vx+θ) + h\n\nwhere A is the amplitude, v is the frequency, θ is the phase, and h is the position in the frequency direction.\nThe search range for each parameter is set as follows based on the experimental results.\n- A: [-100, 100].\n- v: fixed at 0.31 (fixing the frequency slightly improved the score)\n- θ: any real number\n- h: any real number\n\nThen we want to search the parameters that gives the clearest waveform.\nTo do this, I used SA, as following procedures.\n\n- The following operation is repeated tens of thousands of times with decreasing the temperature.\n  - Changes the above parameters slightly, and calculates the sum of the squares of the amplitudes along the waveform as the score (penalty if it refers outside the image). If the score is greater than the last score, accepts the change. Even if the score is smaller than the last score,  accepts the change probabilistically, depending on the score difference and the annealing temperature. Otherwise rejects the change. \n- In addition, to prevent getting a score of local optimum, the above whole process is repeated about several hundred times, and the kth percentile value (k chosen from around 90~100) is used as the final score.\n- It takes several hours to a day (depending on the number of iterations) on my CPU (Core i9-9900K) to calculate the scores of all the test data.\n    - Postprocessing and submitting this can achieve LB score above 0.79 ,with a few hours calculation.\n\nSome example images of detected waveforms for some difficult test cases are as follows.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2F664a106bc0a7a3a394bc78c55a1d80f1%2F5.png?generation=1672801516897133&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Feaa5412252d039795d351203575fd2d1%2F6.png?generation=1672801527565415&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2F344d212d77257fd883c702ea2bfd7780%2F7.png?generation=1672801537224085&alt=media)\n\n### Postprocess\n\nI performed a number of SAs with various numbers of iterations, parameter search ranges, wave thickness, etc., and calculated the ensemble score by weighted averaging.\nI gave up on calculating the CV score, and adjusted the weights so that Public LB would be higher.\nIn addition, the real noise data tended to score slightly higher than the generated noise data (due to its tendency to overfit against noise). Therefore, I first converted the final score to a rank, and then divided the rank value by about 1.1 for real noise data only, improving the score by several points.\n",
      "votes": 41
    },
    {
      "id": 2087065,
      "postDate": "2023-01-05T09:40:20.560Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/shunrcn\" target=\"_blank\">@shunrcn</a> </p>",
      "rawMarkdown": "Congrats @shunrcn ",
      "votes": 1
    },
    {
      "id": 2086339,
      "postDate": "2023-01-04T18:02:20.353Z",
      "content": "<p>Thanks for sharing this fascinating solution! Very unique solution. Congrats <a href=\"https://www.kaggle.com/shunrcn\" target=\"_blank\">@shunrcn</a></p>",
      "rawMarkdown": "Thanks for sharing this fascinating solution! Very unique solution. Congrats @shunrcn",
      "votes": 1
    },
    {
      "id": 2085481,
      "postDate": "2023-01-04T07:32:28.233Z",
      "content": "<p>Congrats Shun_PI！Very excellent solution! </p>",
      "rawMarkdown": "Congrats Shun_PI！Very excellent solution! ",
      "votes": 1
    },
    {
      "id": 2085490,
      "postDate": "2023-01-04T07:38:14.140Z",
      "content": "<p>I'm so impressed by the range of cool solutions in this competition. The diversity seems unparalleled. </p>",
      "rawMarkdown": "I'm so impressed by the range of cool solutions in this competition. The diversity seems unparalleled. ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2087065,
      "author_name": "Megh Dedhia",
      "author_url": "",
      "post_date": "2023-01-05T09:40:20.560000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/shunrcn\" target=\"_blank\">@shunrcn</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2086339,
      "author_name": "Ravi Shah",
      "author_url": "",
      "post_date": "2023-01-04T18:02:20.353000",
      "content": "<p>Thanks for sharing this fascinating solution! Very unique solution. Congrats <a href=\"https://www.kaggle.com/shunrcn\" target=\"_blank\">@shunrcn</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2085481,
      "author_name": "BarryZhou",
      "author_url": "",
      "post_date": "2023-01-04T07:32:28.233000",
      "content": "<p>Congrats Shun_PI！Very excellent solution! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2085490,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2023-01-04T07:38:14.140000",
      "content": "<p>I'm so impressed by the range of cool solutions in this competition. The diversity seems unparalleled. </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2085264": "First of all, thanks to the host and congratulations to the winners.\nI'm so glad to win my first solo gold medal in this competition!\nMy solution is based on [Simulated annealing](https://en.wikipedia.org/wiki/Simulated_annealing) (SA), a technique often used in optimization problems. It is NOT based on machine learning/deep learning/data generation.\n\n### Preprocess\n\nThe following process was performed for each of H1 and L1 and added together to produce a 360*256 image.\n- Calculate the square of the amplitude\n- Removal of horizontal line noise (only real noise)\n\t- Calculate the sum for each row and create a list of 360 elements. Then take a difference of this list. If there is a series of large plus value and large minus value with short intervals, it is assumed that there is horizontal line noise. Then the values of these rows is replaced by the average of the each column. This operation is repeated until it can no longer be done.\n\t- This removes almost all of the real noise and improves the score by about 0.005.\n        - Some example images before and after horizontal line noise removal are as follows.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fdd6cf77558775aabeaf6f8692627cae9%2F1-2.png?generation=1672801107095404&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fc5a3f0d84fd30bdfc21ad2f112c7adcc%2F3-4.png?generation=1672801151754877&alt=media)\n\n- Removal of vertical stripe noise\n\t- Subtract mean for each column, divide by std^2, and correct so that the sum of std of all columns is a constant.\n\t- By dividing by the square of std instead of std, the columns with larger original variance will have smaller final variance. The idea is to reduce the contribution of columns with large original variance because they cannot be relied upon. The score increases slightly by this method.\n- Filling the time gap\n\t- Calculate the difference of the timestamps as d and repeat inserting zero-filled columns round(d / 1800 - 1) times.\n- Average in the time direction and reduce columns size to 256\n\n### Wave Detection with SA\n\nI will explain the essential part of my solution: Simulated annealing (SA).\nFirst, assume that the signal is perfectly sinusoidal and follows the following equation\n\nf(x) = Asin(vx+θ) + h\n\nwhere A is the amplitude, v is the frequency, θ is the phase, and h is the position in the frequency direction.\nThe search range for each parameter is set as follows based on the experimental results.\n- A: [-100, 100].\n- v: fixed at 0.31 (fixing the frequency slightly improved the score)\n- θ: any real number\n- h: any real number\n\nThen we want to search the parameters that gives the clearest waveform.\nTo do this, I used SA, as following procedures.\n\n- The following operation is repeated tens of thousands of times with decreasing the temperature.\n  - Changes the above parameters slightly, and calculates the sum of the squares of the amplitudes along the waveform as the score (penalty if it refers outside the image). If the score is greater than the last score, accepts the change. Even if the score is smaller than the last score,  accepts the change probabilistically, depending on the score difference and the annealing temperature. Otherwise rejects the change. \n- In addition, to prevent getting a score of local optimum, the above whole process is repeated about several hundred times, and the kth percentile value (k chosen from around 90~100) is used as the final score.\n- It takes several hours to a day (depending on the number of iterations) on my CPU (Core i9-9900K) to calculate the scores of all the test data.\n    - Postprocessing and submitting this can achieve LB score above 0.79 ,with a few hours calculation.\n\nSome example images of detected waveforms for some difficult test cases are as follows.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2F664a106bc0a7a3a394bc78c55a1d80f1%2F5.png?generation=1672801516897133&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Feaa5412252d039795d351203575fd2d1%2F6.png?generation=1672801527565415&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2F344d212d77257fd883c702ea2bfd7780%2F7.png?generation=1672801537224085&alt=media)\n\n### Postprocess\n\nI performed a number of SAs with various numbers of iterations, parameter search ranges, wave thickness, etc., and calculated the ensemble score by weighted averaging.\nI gave up on calculating the CV score, and adjusted the weights so that Public LB would be higher.\nIn addition, the real noise data tended to score slightly higher than the generated noise data (due to its tendency to overfit against noise). Therefore, I first converted the final score to a rank, and then divided the rank value by about 1.1 for real noise data only, improving the score by several points.\n",
    "2087065": "Congrats @shunrcn ",
    "2086339": "Thanks for sharing this fascinating solution! Very unique solution. Congrats @shunrcn",
    "2085481": "Congrats Shun_PI！Very excellent solution! ",
    "2085490": "I'm so impressed by the range of cool solutions in this competition. The diversity seems unparalleled. "
  }
}