{
  "id": 376504,
  "title": "2nd Place Solution: GPU-Accelerated Random Search",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/376504",
  "author_name": "knshnb",
  "post_date": "2023-01-06T17:05:18.511000",
  "votes": 21,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thanks to the host and Kaggle staff for holding the competition and congratulations to the winners! I also appreciate my teammates ( <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> and <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a>) a lot.</p>\n<p>Eight hours before the contest ended, we realized that real data in close frequency range might have exactly the same noise, and by actually identifying some of them, we boosted the performance to 0.855/0.849 (1st in public LB!). In this post, we focus on our main solution without this leak magic, which could still win 2nd place (0.835/0.826).</p>\n<h2>Basic Algorithm</h2>\n<p>After struggling with training neural network models, in the last two weeks we found that a very simple solution could work: a random search of signals. Using the velocity for each timestamp computed by PyFstat, the shapes of waves are determined by four parameters: f0, f1, alpha, and delta. We searched the combination of these parameters that maximizes the mean powers (=square of absolute values) of the corresponding part in each data. The essential part of our algorithm is simple as follows (NumPy-like pseudocode):</p>\n<pre><code>stft_sq: (, n_timestamp)\nfrequency_Hz: ()\nvelocity: (, n_timestamp)\n\n _  (n_random_search):\n    f0, f1, alpha, delta = random_sample_params()\n    signal = calc_signal_shape(f0, f1, alpha, delta, velocity)  \n    frequency_idx = np.((signal - frequency_Hz[]) / (frequency_Hz[] - frequency_Hz[]))  \n    signal_part = stft_sq[frequency_idx, np.arange(n_timestamp)]  \n    score = np.sqrt(signal_part.mean())\n</code></pre>\n<p>As a prediction, we simply outputted the mean score of the two detectors (L1 and H1) for each data.</p>\n<p>We implemented batch-wise execution of this algorithm using <a href=\"https://github.com/cupy/cupy\" target=\"_blank\">Cupy</a> to accelerate on GPU. It enabled the search of 3276800 points in around 20 seconds per data on NVIDIA V100. Therefore, it took around 3 GPU hours and 2 GPU days for the execution of all train data and test data, respectively.</p>\n<h2>Details</h2>\n<p>We tried several frequency widths around signals (we take only the nearest one in the above pseudocode) and weighting methods. The best one was the nearest two and linear interpolation by the differences from signal frequencies.</p>\n<p>When calculating signal shapes, we fixed <code>tref</code> for the whole data to a certain value because it can be covered by changing f0 and f1.</p>\n<p>Choosing the right parameter distribution was important. From the analysis of train data, we realized that data with higher scores than a certain threshold (around 1.541) were almost surely positive. The distribution of found parameters in test data with high scores are as follows  (the right figure is f0 scaled by each data's frequency range):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2Fa7a7e2c88250b1bd0577259da7a2c808%2Fparameter-distribution.png?generation=1673022472671396&amp;alt=media\"><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F8d0831bfed267e4e608344eca4fd6998%2Ff0_ratio.png?generation=1673022226325898&amp;alt=media\"><br>\nFrom this, we decided to sample each parameter from the following distributions, which enhanced the performance a lot:</p>\n<pre><code>alpha: Uniform([0, 2 \\pi])\ndelta: arcsine(Uniform([-1, 1]))\nf0: Beta(2, 2) between frequency range 20% extended to both sides\nf1: 1/3 are from -2 * 10^(Uniform([-11, -9])), 2/3 are from 2 * 10^(Uniform([-11, -8]))\n</code></pre>\n<p>Around 1/5 of the test data included real noise with nonstationarity and peak in certain frequencies, etc. For these data, we performed time-wise normalization after a simple rule-based frequency mask like below:</p>\n<pre><code> ():\n    freq_std = np.std(stft_sq, axis=)\n    error_freq = (freq_std &gt; np.median(freq_std) * median_coeff) &amp; (freq_std &gt; np.percentile(freq_std, percent))\n    stft_sq[error_freq] = stft_sq[~error_freq].mean(axis=)\n     stft_sq\n</code></pre>\n<p>It seems most of the strong noises are removed by this preprocessing (an example of <code>id:56b090eaf</code> is below), but we did not have enough time to put much effort into this, so there might be some room for improvement.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2Fe6e2af680a8f328ac8581036a8483af3%2Fdelete_freq.png?generation=1673021886640173&amp;alt=media\"></p>\n<h2>Validation</h2>\n<p>Since the number of training data was very limited, we calculated the average AUC of 5 different random seeds. It correlated with the public LB to some extent. The best mean train AUC was 0.902940, which seems to overfit a little, still. (Did not have enough time after identifying parameter distribution…)</p>\n<p>For real data, we generated a validation set by adding signals to low-score samples, but it did not correlate well with LB, so we focused more on LB scores.</p>\n<h2>Things that might be improved</h2>\n<ul>\n<li>More sophisticated preprocessing of real data</li>\n<li>2-stage search</li>\n<li>Rule-based or machine learning postprocessing of the search results</li>\n<li>Consider amplitude differences by timestamps</li>\n</ul>",
  "messages": [
    {
      "id": 2088816,
      "postDate": "2023-01-06T17:05:18.510Z",
      "content": "<p>Thanks to the host and Kaggle staff for holding the competition and congratulations to the winners! I also appreciate my teammates ( <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> and <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a>) a lot.</p>\n<p>Eight hours before the contest ended, we realized that real data in close frequency range might have exactly the same noise, and by actually identifying some of them, we boosted the performance to 0.855/0.849 (1st in public LB!). In this post, we focus on our main solution without this leak magic, which could still win 2nd place (0.835/0.826).</p>\n<h2>Basic Algorithm</h2>\n<p>After struggling with training neural network models, in the last two weeks we found that a very simple solution could work: a random search of signals. Using the velocity for each timestamp computed by PyFstat, the shapes of waves are determined by four parameters: f0, f1, alpha, and delta. We searched the combination of these parameters that maximizes the mean powers (=square of absolute values) of the corresponding part in each data. The essential part of our algorithm is simple as follows (NumPy-like pseudocode):</p>\n<pre><code>stft_sq: (, n_timestamp)\nfrequency_Hz: ()\nvelocity: (, n_timestamp)\n\n _  (n_random_search):\n    f0, f1, alpha, delta = random_sample_params()\n    signal = calc_signal_shape(f0, f1, alpha, delta, velocity)  \n    frequency_idx = np.((signal - frequency_Hz[]) / (frequency_Hz[] - frequency_Hz[]))  \n    signal_part = stft_sq[frequency_idx, np.arange(n_timestamp)]  \n    score = np.sqrt(signal_part.mean())\n</code></pre>\n<p>As a prediction, we simply outputted the mean score of the two detectors (L1 and H1) for each data.</p>\n<p>We implemented batch-wise execution of this algorithm using <a href=\"https://github.com/cupy/cupy\" target=\"_blank\">Cupy</a> to accelerate on GPU. It enabled the search of 3276800 points in around 20 seconds per data on NVIDIA V100. Therefore, it took around 3 GPU hours and 2 GPU days for the execution of all train data and test data, respectively.</p>\n<h2>Details</h2>\n<p>We tried several frequency widths around signals (we take only the nearest one in the above pseudocode) and weighting methods. The best one was the nearest two and linear interpolation by the differences from signal frequencies.</p>\n<p>When calculating signal shapes, we fixed <code>tref</code> for the whole data to a certain value because it can be covered by changing f0 and f1.</p>\n<p>Choosing the right parameter distribution was important. From the analysis of train data, we realized that data with higher scores than a certain threshold (around 1.541) were almost surely positive. The distribution of found parameters in test data with high scores are as follows  (the right figure is f0 scaled by each data's frequency range):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2Fa7a7e2c88250b1bd0577259da7a2c808%2Fparameter-distribution.png?generation=1673022472671396&amp;alt=media\"><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F8d0831bfed267e4e608344eca4fd6998%2Ff0_ratio.png?generation=1673022226325898&amp;alt=media\"><br>\nFrom this, we decided to sample each parameter from the following distributions, which enhanced the performance a lot:</p>\n<pre><code>alpha: Uniform([0, 2 \\pi])\ndelta: arcsine(Uniform([-1, 1]))\nf0: Beta(2, 2) between frequency range 20% extended to both sides\nf1: 1/3 are from -2 * 10^(Uniform([-11, -9])), 2/3 are from 2 * 10^(Uniform([-11, -8]))\n</code></pre>\n<p>Around 1/5 of the test data included real noise with nonstationarity and peak in certain frequencies, etc. For these data, we performed time-wise normalization after a simple rule-based frequency mask like below:</p>\n<pre><code> ():\n    freq_std = np.std(stft_sq, axis=)\n    error_freq = (freq_std &gt; np.median(freq_std) * median_coeff) &amp; (freq_std &gt; np.percentile(freq_std, percent))\n    stft_sq[error_freq] = stft_sq[~error_freq].mean(axis=)\n     stft_sq\n</code></pre>\n<p>It seems most of the strong noises are removed by this preprocessing (an example of <code>id:56b090eaf</code> is below), but we did not have enough time to put much effort into this, so there might be some room for improvement.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2Fe6e2af680a8f328ac8581036a8483af3%2Fdelete_freq.png?generation=1673021886640173&amp;alt=media\"></p>\n<h2>Validation</h2>\n<p>Since the number of training data was very limited, we calculated the average AUC of 5 different random seeds. It correlated with the public LB to some extent. The best mean train AUC was 0.902940, which seems to overfit a little, still. (Did not have enough time after identifying parameter distribution…)</p>\n<p>For real data, we generated a validation set by adding signals to low-score samples, but it did not correlate well with LB, so we focused more on LB scores.</p>\n<h2>Things that might be improved</h2>\n<ul>\n<li>More sophisticated preprocessing of real data</li>\n<li>2-stage search</li>\n<li>Rule-based or machine learning postprocessing of the search results</li>\n<li>Consider amplitude differences by timestamps</li>\n</ul>",
      "rawMarkdown": "Thanks to the host and Kaggle staff for holding the competition and congratulations to the winners! I also appreciate my teammates ( @charmq and @yoichi7yamakawa) a lot.\n\nEight hours before the contest ended, we realized that real data in close frequency range might have exactly the same noise, and by actually identifying some of them, we boosted the performance to 0.855/0.849 (1st in public LB!). In this post, we focus on our main solution without this leak magic, which could still win 2nd place (0.835/0.826).\n\n## Basic Algorithm\nAfter struggling with training neural network models, in the last two weeks we found that a very simple solution could work: a random search of signals. Using the velocity for each timestamp computed by PyFstat, the shapes of waves are determined by four parameters: f0, f1, alpha, and delta. We searched the combination of these parameters that maximizes the mean powers (=square of absolute values) of the corresponding part in each data. The essential part of our algorithm is simple as follows (NumPy-like pseudocode):\n\n```python\nstft_sq: (360, n_timestamp)\nfrequency_Hz: (360)\nvelocity: (3, n_timestamp)\n\nfor _ in range(n_random_search):\n    f0, f1, alpha, delta = random_sample_params()\n    signal = calc_signal_shape(f0, f1, alpha, delta, velocity)  # (n_timestamp): frequency for each timestamp\n    frequency_idx = np.round((signal - frequency_Hz[0]) / (frequency_Hz[1] - frequency_Hz[0]))  # (n_timestamp)\n    signal_part = stft_sq[frequency_idx, np.arange(n_timestamp)]  # (n_timestamp): powers of corresponding part in data\n    score = np.sqrt(signal_part.mean())\n```\n\nAs a prediction, we simply outputted the mean score of the two detectors (L1 and H1) for each data.\n\nWe implemented batch-wise execution of this algorithm using [Cupy](https://github.com/cupy/cupy) to accelerate on GPU. It enabled the search of 3276800 points in around 20 seconds per data on NVIDIA V100. Therefore, it took around 3 GPU hours and 2 GPU days for the execution of all train data and test data, respectively.\n\n## Details\nWe tried several frequency widths around signals (we take only the nearest one in the above pseudocode) and weighting methods. The best one was the nearest two and linear interpolation by the differences from signal frequencies.\n\nWhen calculating signal shapes, we fixed `tref` for the whole data to a certain value because it can be covered by changing f0 and f1.\n\nChoosing the right parameter distribution was important. From the analysis of train data, we realized that data with higher scores than a certain threshold (around 1.541) were almost surely positive. The distribution of found parameters in test data with high scores are as follows  (the right figure is f0 scaled by each data's frequency range):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2Fa7a7e2c88250b1bd0577259da7a2c808%2Fparameter-distribution.png?generation=1673022472671396&alt=media\" width=\"300\"><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F8d0831bfed267e4e608344eca4fd6998%2Ff0_ratio.png?generation=1673022226325898&alt=media\" width=\"300\">\nFrom this, we decided to sample each parameter from the following distributions, which enhanced the performance a lot:\n```\nalpha: Uniform([0, 2 \\pi])\ndelta: arcsine(Uniform([-1, 1]))\nf0: Beta(2, 2) between frequency range 20% extended to both sides\nf1: 1/3 are from -2 * 10^(Uniform([-11, -9])), 2/3 are from 2 * 10^(Uniform([-11, -8]))\n```\n\nAround 1/5 of the test data included real noise with nonstationarity and peak in certain frequencies, etc. For these data, we performed time-wise normalization after a simple rule-based frequency mask like below:\n```python\ndef remove_freq_peak(stft_sq, median_coeff: float = 1.1, percent: int = 75):\n    freq_std = np.std(stft_sq, axis=1)\n    error_freq = (freq_std > np.median(freq_std) * median_coeff) & (freq_std > np.percentile(freq_std, percent))\n    stft_sq[error_freq] = stft_sq[~error_freq].mean(axis=0)\n    return stft_sq\n```\nIt seems most of the strong noises are removed by this preprocessing (an example of `id:56b090eaf` is below), but we did not have enough time to put much effort into this, so there might be some room for improvement.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2Fe6e2af680a8f328ac8581036a8483af3%2Fdelete_freq.png?generation=1673021886640173&alt=media\" width=\"400\">\n\n## Validation\nSince the number of training data was very limited, we calculated the average AUC of 5 different random seeds. It correlated with the public LB to some extent. The best mean train AUC was 0.902940, which seems to overfit a little, still. (Did not have enough time after identifying parameter distribution...)\n\nFor real data, we generated a validation set by adding signals to low-score samples, but it did not correlate well with LB, so we focused more on LB scores.\n\n## Things that might be improved\n- More sophisticated preprocessing of real data\n- 2-stage search\n- Rule-based or machine learning postprocessing of the search results\n- Consider amplitude differences by timestamps",
      "votes": 21
    },
    {
      "id": 2091389,
      "postDate": "2023-01-08T11:13:47.187Z",
      "content": "<p>Congratulations! Great Work!<br>\nI am unsure about the velocity with shape (3, n_timestamp). Can you clarify it, please?</p>",
      "rawMarkdown": "Congratulations! Great Work!\nI am unsure about the velocity with shape (3, n_timestamp). Can you clarify it, please?",
      "votes": 1,
      "replies": [
        {
          "id": 2092447,
          "postDate": "2023-01-09T10:49:10.957Z",
          "content": "<p>Thanks! That's three-dimensional velocity vectors for different timestamps, which you can get by using PyFstat (appear as <code>velocities</code> in the <a href=\"https://github.com/PyFstat/PyFstat/blob/master/examples/tutorials/1_generating_signals.ipynb\" target=\"_blank\">tutorial</a>).</p>",
          "rawMarkdown": "Thanks! That's three-dimensional velocity vectors for different timestamps, which you can get by using PyFstat (appear as `velocities` in the [tutorial](https://github.com/PyFstat/PyFstat/blob/master/examples/tutorials/1_generating_signals.ipynb))."
        }
      ]
    },
    {
      "id": 2090220,
      "postDate": "2023-01-07T05:42:53.840Z",
      "content": "<p>Congratulations! The solution through signal exploration, parameter sampling, denoising with outlier frequency detection made me think on a different line. My solution involved t-test to detect deviations and the guarantee in presence of signals through signal to noise ratio.</p>",
      "rawMarkdown": "Congratulations! The solution through signal exploration, parameter sampling, denoising with outlier frequency detection made me think on a different line. My solution involved t-test to detect deviations and the guarantee in presence of signals through signal to noise ratio.",
      "votes": 1
    },
    {
      "id": 2088818,
      "postDate": "2023-01-06T17:08:49.843Z",
      "content": "<p>Congratulations on getting 2nd place!  🎉</p>",
      "rawMarkdown": "Congratulations on getting 2nd place!  🎉",
      "votes": 1
    },
    {
      "id": 2103981,
      "postDate": "2023-01-17T14:02:17.690Z",
      "content": "<p>I published a source code: <a href=\"https://github.com/knshnb/kaggle-g2net2-2nd-place\" target=\"_blank\">https://github.com/knshnb/kaggle-g2net2-2nd-place</a></p>",
      "rawMarkdown": "I published a source code: https://github.com/knshnb/kaggle-g2net2-2nd-place",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2091389,
      "author_name": "SemenB",
      "author_url": "",
      "post_date": "2023-01-08T11:13:47.187000",
      "content": "<p>Congratulations! Great Work!<br>\nI am unsure about the velocity with shape (3, n_timestamp). Can you clarify it, please?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2092447,
          "author_name": "knshnb",
          "author_url": "",
          "post_date": "2023-01-09T10:49:10.957000",
          "content": "<p>Thanks! That's three-dimensional velocity vectors for different timestamps, which you can get by using PyFstat (appear as <code>velocities</code> in the <a href=\"https://github.com/PyFstat/PyFstat/blob/master/examples/tutorials/1_generating_signals.ipynb\" target=\"_blank\">tutorial</a>).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2090220,
      "author_name": "Suma Mallapragada",
      "author_url": "",
      "post_date": "2023-01-07T05:42:53.840000",
      "content": "<p>Congratulations! The solution through signal exploration, parameter sampling, denoising with outlier frequency detection made me think on a different line. My solution involved t-test to detect deviations and the guarantee in presence of signals through signal to noise ratio.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2088818,
      "author_name": "BwandoWando",
      "author_url": "",
      "post_date": "2023-01-06T17:08:49.843000",
      "content": "<p>Congratulations on getting 2nd place!  🎉</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2103981,
      "author_name": "knshnb",
      "author_url": "",
      "post_date": "2023-01-17T14:02:17.690000",
      "content": "<p>I published a source code: <a href=\"https://github.com/knshnb/kaggle-g2net2-2nd-place\" target=\"_blank\">https://github.com/knshnb/kaggle-g2net2-2nd-place</a></p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2088816": "Thanks to the host and Kaggle staff for holding the competition and congratulations to the winners! I also appreciate my teammates ( @charmq and @yoichi7yamakawa) a lot.\n\nEight hours before the contest ended, we realized that real data in close frequency range might have exactly the same noise, and by actually identifying some of them, we boosted the performance to 0.855/0.849 (1st in public LB!). In this post, we focus on our main solution without this leak magic, which could still win 2nd place (0.835/0.826).\n\n## Basic Algorithm\nAfter struggling with training neural network models, in the last two weeks we found that a very simple solution could work: a random search of signals. Using the velocity for each timestamp computed by PyFstat, the shapes of waves are determined by four parameters: f0, f1, alpha, and delta. We searched the combination of these parameters that maximizes the mean powers (=square of absolute values) of the corresponding part in each data. The essential part of our algorithm is simple as follows (NumPy-like pseudocode):\n\n```python\nstft_sq: (360, n_timestamp)\nfrequency_Hz: (360)\nvelocity: (3, n_timestamp)\n\nfor _ in range(n_random_search):\n    f0, f1, alpha, delta = random_sample_params()\n    signal = calc_signal_shape(f0, f1, alpha, delta, velocity)  # (n_timestamp): frequency for each timestamp\n    frequency_idx = np.round((signal - frequency_Hz[0]) / (frequency_Hz[1] - frequency_Hz[0]))  # (n_timestamp)\n    signal_part = stft_sq[frequency_idx, np.arange(n_timestamp)]  # (n_timestamp): powers of corresponding part in data\n    score = np.sqrt(signal_part.mean())\n```\n\nAs a prediction, we simply outputted the mean score of the two detectors (L1 and H1) for each data.\n\nWe implemented batch-wise execution of this algorithm using [Cupy](https://github.com/cupy/cupy) to accelerate on GPU. It enabled the search of 3276800 points in around 20 seconds per data on NVIDIA V100. Therefore, it took around 3 GPU hours and 2 GPU days for the execution of all train data and test data, respectively.\n\n## Details\nWe tried several frequency widths around signals (we take only the nearest one in the above pseudocode) and weighting methods. The best one was the nearest two and linear interpolation by the differences from signal frequencies.\n\nWhen calculating signal shapes, we fixed `tref` for the whole data to a certain value because it can be covered by changing f0 and f1.\n\nChoosing the right parameter distribution was important. From the analysis of train data, we realized that data with higher scores than a certain threshold (around 1.541) were almost surely positive. The distribution of found parameters in test data with high scores are as follows  (the right figure is f0 scaled by each data's frequency range):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2Fa7a7e2c88250b1bd0577259da7a2c808%2Fparameter-distribution.png?generation=1673022472671396&alt=media\" width=\"300\"><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2F8d0831bfed267e4e608344eca4fd6998%2Ff0_ratio.png?generation=1673022226325898&alt=media\" width=\"300\">\nFrom this, we decided to sample each parameter from the following distributions, which enhanced the performance a lot:\n```\nalpha: Uniform([0, 2 \\pi])\ndelta: arcsine(Uniform([-1, 1]))\nf0: Beta(2, 2) between frequency range 20% extended to both sides\nf1: 1/3 are from -2 * 10^(Uniform([-11, -9])), 2/3 are from 2 * 10^(Uniform([-11, -8]))\n```\n\nAround 1/5 of the test data included real noise with nonstationarity and peak in certain frequencies, etc. For these data, we performed time-wise normalization after a simple rule-based frequency mask like below:\n```python\ndef remove_freq_peak(stft_sq, median_coeff: float = 1.1, percent: int = 75):\n    freq_std = np.std(stft_sq, axis=1)\n    error_freq = (freq_std > np.median(freq_std) * median_coeff) & (freq_std > np.percentile(freq_std, percent))\n    stft_sq[error_freq] = stft_sq[~error_freq].mean(axis=0)\n    return stft_sq\n```\nIt seems most of the strong noises are removed by this preprocessing (an example of `id:56b090eaf` is below), but we did not have enough time to put much effort into this, so there might be some room for improvement.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9088007%2Fe6e2af680a8f328ac8581036a8483af3%2Fdelete_freq.png?generation=1673021886640173&alt=media\" width=\"400\">\n\n## Validation\nSince the number of training data was very limited, we calculated the average AUC of 5 different random seeds. It correlated with the public LB to some extent. The best mean train AUC was 0.902940, which seems to overfit a little, still. (Did not have enough time after identifying parameter distribution...)\n\nFor real data, we generated a validation set by adding signals to low-score samples, but it did not correlate well with LB, so we focused more on LB scores.\n\n## Things that might be improved\n- More sophisticated preprocessing of real data\n- 2-stage search\n- Rule-based or machine learning postprocessing of the search results\n- Consider amplitude differences by timestamps",
    "2091389": "Congratulations! Great Work!\nI am unsure about the velocity with shape (3, n_timestamp). Can you clarify it, please?",
    "2090220": "Congratulations! The solution through signal exploration, parameter sampling, denoising with outlier frequency detection made me think on a different line. My solution involved t-test to detect deviations and the guarantee in presence of signals through signal to noise ratio.",
    "2088818": "Congratulations on getting 2nd place!  🎉",
    "2103981": "I published a source code: https://github.com/knshnb/kaggle-g2net2-2nd-place"
  }
}