{
  "id": 376724,
  "title": "13th Place Solution: plain machine learning.",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/376724",
  "author_name": "assign",
  "post_date": "2023-01-08T06:40:43.654000",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>summary</h1>\n<ul>\n<li>Dataset</li>\n<li>Normalization</li>\n<li>PyFstat data generation / parameter range</li>\n<li>Large Kernel and FFT</li>\n<li>Train more with EMA</li>\n<li>Etc</li>\n</ul>\n<h1>Dataset</h1>\n<ul>\n<li>test/: used for data augmentation during training (finally not applied)</li>\n<li>train/ : validation data</li>\n<li>Generate from PyFstat: data for training</li>\n</ul>\n<h1>Normalization</h1>\n<pre><code>from scipy.stats import norm\ndef Fnormalize(X):\n    X /= X.sum(-2, keepdims=True)\n    return X\ndef Pnormalize(X):\n    n = np.prod(X.shape[-2:])\n    POS = min(int(n * 0.999), n - 10)\n    EXP = norm.ppf((POS + 1 - np.pi / 8) / (n - np.pi / 4 + 1))\n    scale = np.partition(X.flatten(), POS, -1)[POS]\n    X /= scale / EXP.astype(scale.dtype) ** 2\n    return X\ndef normalize(X):\n    X = (X[..., None].view(X.real.dtype) ** 2).sum(-1)\n    X = Fnormalize(X)\n    X = Pnormalize(X)\n    return X\n</code></pre>\n<p>This data normalization has a simple theoretical background.<br>\n<code>Fnormalize</code> : Changes the complex input to the sum of the squares of the real and imaginary parts of the elements, transforming the input into a <code>chi2</code> distribution.<br>\n<code>Pnormalize</code>: Scales so that the square of the <code>POS</code>th largest number of the normal distribution is equal to the <code>POS</code>th largest number of the data.<br>\nIt would be a good idea to interpret <code>Pnormalize</code> for the <code>chi2</code> distribution, but since this function only scales the input, I think you'll get similar performance.<br>\nIn the same model, the performance is improved by changing the normalization technique of the inference step.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Pnormalize only</td>\n<td>0.766</td>\n<td>0.739</td>\n</tr>\n<tr>\n<td>Fnormalize + Pnormalize</td>\n<td>0.770</td>\n<td>0.751</td>\n</tr>\n</tbody>\n</table>\n<h1>PyFstat data generation / parameter range</h1>\n<p>Before implement the dataset sampling code, I read the PyFstat library code and tracked the parameter range. However, <code>psi</code> and <code>phi</code> were not found.<br>\nSurprisingly, only 5 days before the end of the competition, I found out about the proper range <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/373973\" target=\"_blank\">here</a>, and the model performance improved.</p>\n<p>Setting 1(Gen1):</p>\n<pre><code>tstart: [630720013, 1861492413)\nF0: [45, 600)\nF1: 10 ^ truncnorm.isf(rng.uniform(), a=-100, b=3, loc=-15, scale=2)\nAlpha: [0, 2pi)\nDelta: [-pi/2, pi/2)\ncosi: [-1, 1)\npsi: [-1, 1)\nphi: [-1, 1)\nTsft: 1800\nSFTWindowType: \"tukey\"\nSFTWindowBeta: 0.0001\n</code></pre>\n<p>Setting 2(Gen2):</p>\n<pre><code>tstart: [630720013, 1861492413)\nF0: [45, 600)\nF1: 10 ^ truncnorm.isf(rng.uniform(), a=-100, b=3, loc=-15, scale=2)\nAlpha: [0, 2pi)\nDelta: [-pi/2, pi/2)\ncosi: [-1, 1)\npsi: [-pi/4, pi/4)\nphi: [0, 2pi)\nTsft: 1800\nSFTWindowType: \"tukey\"\nSFTWindowBeta: 0.0001\n</code></pre>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Validation ROCAUC</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>exp/v0.2-1229-1 Best valid (Gen1)</td>\n<td>0.8569</td>\n<td>0.761</td>\n<td>0.747</td>\n</tr>\n<tr>\n<td>exp/v0.2-1229-1 Last epoch (Gen1)</td>\n<td>0.7278</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>exp/v0.2-1230-2 Best valid (Gen2)</td>\n<td>0.8662812499999999</td>\n<td>0.759</td>\n<td>0.747</td>\n</tr>\n<tr>\n<td>exp/v0.2-1230-2 Last epoch (Gen2)</td>\n<td>0.8624</td>\n<td>0.770</td>\n<td>0.756</td>\n</tr>\n</tbody>\n</table>\n<h1>Large Kernel and FFT</h1>\n<p>I implemented <a href=\"https://github.com/klae01/fft-conv-pytorch\" target=\"_blank\">fft convolution</a> to use large kernels in training. You can refer to the notebook that was released one week before the end of the competition. <a href=\"https://www.kaggle.com/code/assign/g2net-large-kernel-inference-fft-conv2d\" target=\"_blank\">link</a></p>\n<h2>Weight Decay</h2>\n<p>It has been observed that large kernels are prone to bias and are weak. I applied a strong weight decay of 0.001 to that layer.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Validation ROCAUC</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>exp/v0.2-1229-1 weight decay: 1e-3 Best valid (Gen1)</td>\n<td>0.8569</td>\n<td>0.761</td>\n<td>0.747</td>\n</tr>\n<tr>\n<td>exp/v0.2-1230-1 weight decay: 1e-4 Best valid (Gen1)</td>\n<td>0.8538749999999999</td>\n<td>0.756</td>\n<td>0.743</td>\n</tr>\n</tbody>\n</table>\n<h2>Dense stem</h2>\n<p>Some performance improvement is achieved by changing the depth-wise convolution applied per L1/H1 detector channel of the model to a normal convolution.<br>\nCompared to the best validation score, the gap is large, but the last epoch still performs better.</p>\n<table>\n<thead>\n<tr>\n<th>exp/v0.2-1231-1</th>\n<th>Validation ROCAUC</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Best valid (Gen2)</td>\n<td>0.8635312500000001</td>\n<td>0.762</td>\n<td>0.748</td>\n</tr>\n<tr>\n<td>Last epoch (Gen2)</td>\n<td>0.8532</td>\n<td>0.771</td>\n<td>0.760</td>\n</tr>\n</tbody>\n</table>\n<h1>Train more with EMA</h1>\n<p>I trained with 2x epochs with the same settings, but the model diverged. However, as the model with the highest validation score, it showed an improved score than before.<br>\nTraining ended 3 hours before the end of the competition, and this is the last training.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Gen 1</td>\n<td>0.761</td>\n<td>0.747</td>\n</tr>\n<tr>\n<td>Gen 2</td>\n<td>0.770</td>\n<td>0.756</td>\n</tr>\n<tr>\n<td>Gen 2 + Dense stem</td>\n<td>0.771</td>\n<td>0.760</td>\n</tr>\n<tr>\n<td>Gen 2 + Dense stem + EMA</td>\n<td>0.773</td>\n<td>0.761</td>\n</tr>\n</tbody>\n</table>\n<h1>Etc</h1>\n<ul>\n<li>I didn't use the idea of putting horizontal/vertical lines in the PyFstat data.</li>\n<li>To avoid overfitting all hypotheses to the train data, the data belonging to <code>train/</code> were only analyzed at the statistical level. Might be a dumb choice :)</li>\n<li>I have applied the idea of creating gaps in continuous inputs and filling them with noise, but the validation loss at the beginning of training is greatly improved, but the training becomes very unstable, and the training/validation/submission scores are all lower. It's an idea that was ultimately abandoned.</li>\n<li>PyFstat is slow because they uses disk. This is the bottleneck. Use <code>tmpfs</code>.</li>\n<li>The model was trained with 256 batches of 64k samples reused 32 times each.</li>\n<li>The time spent implementing the tool was too long compared to the time spent improving performance. For example, an implementation that multiprocesses a pipeline that reuses and discards data and guarantees the same order for seeds / fft convolution optimization / etc.</li>\n<li>Surprisingly, SNR ranges from training [6, 10]. Better than [5,10], better performance than [6, 15]. In the training phase, a low SNR range causes the model to diverge, while a high SNR range makes the model less discerning when no signal is present.</li>\n</ul>\n<p>Note: This is a rough comparison of the functions that investigate the SNR mentioned above.</p>\n<table>\n<thead>\n<tr>\n<th>PyFstat parameter</th>\n<th></th>\n<th></th>\n<th>SNR estimation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>h0</td>\n<td>sqrtSX</td>\n<td>cosi</td>\n<td></td>\n</tr>\n<tr>\n<td>1</td>\n<td>[10,100]</td>\n<td>0</td>\n<td>[9.2,101]</td>\n</tr>\n<tr>\n<td>1</td>\n<td>[10,100]</td>\n<td>1 or -1</td>\n<td>[2.3,19]</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 2091230,
      "postDate": "2023-01-08T06:40:43.653Z",
      "content": "<h1>summary</h1>\n<ul>\n<li>Dataset</li>\n<li>Normalization</li>\n<li>PyFstat data generation / parameter range</li>\n<li>Large Kernel and FFT</li>\n<li>Train more with EMA</li>\n<li>Etc</li>\n</ul>\n<h1>Dataset</h1>\n<ul>\n<li>test/: used for data augmentation during training (finally not applied)</li>\n<li>train/ : validation data</li>\n<li>Generate from PyFstat: data for training</li>\n</ul>\n<h1>Normalization</h1>\n<pre><code>from scipy.stats import norm\ndef Fnormalize(X):\n    X /= X.sum(-2, keepdims=True)\n    return X\ndef Pnormalize(X):\n    n = np.prod(X.shape[-2:])\n    POS = min(int(n * 0.999), n - 10)\n    EXP = norm.ppf((POS + 1 - np.pi / 8) / (n - np.pi / 4 + 1))\n    scale = np.partition(X.flatten(), POS, -1)[POS]\n    X /= scale / EXP.astype(scale.dtype) ** 2\n    return X\ndef normalize(X):\n    X = (X[..., None].view(X.real.dtype) ** 2).sum(-1)\n    X = Fnormalize(X)\n    X = Pnormalize(X)\n    return X\n</code></pre>\n<p>This data normalization has a simple theoretical background.<br>\n<code>Fnormalize</code> : Changes the complex input to the sum of the squares of the real and imaginary parts of the elements, transforming the input into a <code>chi2</code> distribution.<br>\n<code>Pnormalize</code>: Scales so that the square of the <code>POS</code>th largest number of the normal distribution is equal to the <code>POS</code>th largest number of the data.<br>\nIt would be a good idea to interpret <code>Pnormalize</code> for the <code>chi2</code> distribution, but since this function only scales the input, I think you'll get similar performance.<br>\nIn the same model, the performance is improved by changing the normalization technique of the inference step.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Pnormalize only</td>\n<td>0.766</td>\n<td>0.739</td>\n</tr>\n<tr>\n<td>Fnormalize + Pnormalize</td>\n<td>0.770</td>\n<td>0.751</td>\n</tr>\n</tbody>\n</table>\n<h1>PyFstat data generation / parameter range</h1>\n<p>Before implement the dataset sampling code, I read the PyFstat library code and tracked the parameter range. However, <code>psi</code> and <code>phi</code> were not found.<br>\nSurprisingly, only 5 days before the end of the competition, I found out about the proper range <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/373973\" target=\"_blank\">here</a>, and the model performance improved.</p>\n<p>Setting 1(Gen1):</p>\n<pre><code>tstart: [630720013, 1861492413)\nF0: [45, 600)\nF1: 10 ^ truncnorm.isf(rng.uniform(), a=-100, b=3, loc=-15, scale=2)\nAlpha: [0, 2pi)\nDelta: [-pi/2, pi/2)\ncosi: [-1, 1)\npsi: [-1, 1)\nphi: [-1, 1)\nTsft: 1800\nSFTWindowType: \"tukey\"\nSFTWindowBeta: 0.0001\n</code></pre>\n<p>Setting 2(Gen2):</p>\n<pre><code>tstart: [630720013, 1861492413)\nF0: [45, 600)\nF1: 10 ^ truncnorm.isf(rng.uniform(), a=-100, b=3, loc=-15, scale=2)\nAlpha: [0, 2pi)\nDelta: [-pi/2, pi/2)\ncosi: [-1, 1)\npsi: [-pi/4, pi/4)\nphi: [0, 2pi)\nTsft: 1800\nSFTWindowType: \"tukey\"\nSFTWindowBeta: 0.0001\n</code></pre>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Validation ROCAUC</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>exp/v0.2-1229-1 Best valid (Gen1)</td>\n<td>0.8569</td>\n<td>0.761</td>\n<td>0.747</td>\n</tr>\n<tr>\n<td>exp/v0.2-1229-1 Last epoch (Gen1)</td>\n<td>0.7278</td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>exp/v0.2-1230-2 Best valid (Gen2)</td>\n<td>0.8662812499999999</td>\n<td>0.759</td>\n<td>0.747</td>\n</tr>\n<tr>\n<td>exp/v0.2-1230-2 Last epoch (Gen2)</td>\n<td>0.8624</td>\n<td>0.770</td>\n<td>0.756</td>\n</tr>\n</tbody>\n</table>\n<h1>Large Kernel and FFT</h1>\n<p>I implemented <a href=\"https://github.com/klae01/fft-conv-pytorch\" target=\"_blank\">fft convolution</a> to use large kernels in training. You can refer to the notebook that was released one week before the end of the competition. <a href=\"https://www.kaggle.com/code/assign/g2net-large-kernel-inference-fft-conv2d\" target=\"_blank\">link</a></p>\n<h2>Weight Decay</h2>\n<p>It has been observed that large kernels are prone to bias and are weak. I applied a strong weight decay of 0.001 to that layer.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Validation ROCAUC</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>exp/v0.2-1229-1 weight decay: 1e-3 Best valid (Gen1)</td>\n<td>0.8569</td>\n<td>0.761</td>\n<td>0.747</td>\n</tr>\n<tr>\n<td>exp/v0.2-1230-1 weight decay: 1e-4 Best valid (Gen1)</td>\n<td>0.8538749999999999</td>\n<td>0.756</td>\n<td>0.743</td>\n</tr>\n</tbody>\n</table>\n<h2>Dense stem</h2>\n<p>Some performance improvement is achieved by changing the depth-wise convolution applied per L1/H1 detector channel of the model to a normal convolution.<br>\nCompared to the best validation score, the gap is large, but the last epoch still performs better.</p>\n<table>\n<thead>\n<tr>\n<th>exp/v0.2-1231-1</th>\n<th>Validation ROCAUC</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Best valid (Gen2)</td>\n<td>0.8635312500000001</td>\n<td>0.762</td>\n<td>0.748</td>\n</tr>\n<tr>\n<td>Last epoch (Gen2)</td>\n<td>0.8532</td>\n<td>0.771</td>\n<td>0.760</td>\n</tr>\n</tbody>\n</table>\n<h1>Train more with EMA</h1>\n<p>I trained with 2x epochs with the same settings, but the model diverged. However, as the model with the highest validation score, it showed an improved score than before.<br>\nTraining ended 3 hours before the end of the competition, and this is the last training.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Gen 1</td>\n<td>0.761</td>\n<td>0.747</td>\n</tr>\n<tr>\n<td>Gen 2</td>\n<td>0.770</td>\n<td>0.756</td>\n</tr>\n<tr>\n<td>Gen 2 + Dense stem</td>\n<td>0.771</td>\n<td>0.760</td>\n</tr>\n<tr>\n<td>Gen 2 + Dense stem + EMA</td>\n<td>0.773</td>\n<td>0.761</td>\n</tr>\n</tbody>\n</table>\n<h1>Etc</h1>\n<ul>\n<li>I didn't use the idea of putting horizontal/vertical lines in the PyFstat data.</li>\n<li>To avoid overfitting all hypotheses to the train data, the data belonging to <code>train/</code> were only analyzed at the statistical level. Might be a dumb choice :)</li>\n<li>I have applied the idea of creating gaps in continuous inputs and filling them with noise, but the validation loss at the beginning of training is greatly improved, but the training becomes very unstable, and the training/validation/submission scores are all lower. It's an idea that was ultimately abandoned.</li>\n<li>PyFstat is slow because they uses disk. This is the bottleneck. Use <code>tmpfs</code>.</li>\n<li>The model was trained with 256 batches of 64k samples reused 32 times each.</li>\n<li>The time spent implementing the tool was too long compared to the time spent improving performance. For example, an implementation that multiprocesses a pipeline that reuses and discards data and guarantees the same order for seeds / fft convolution optimization / etc.</li>\n<li>Surprisingly, SNR ranges from training [6, 10]. Better than [5,10], better performance than [6, 15]. In the training phase, a low SNR range causes the model to diverge, while a high SNR range makes the model less discerning when no signal is present.</li>\n</ul>\n<p>Note: This is a rough comparison of the functions that investigate the SNR mentioned above.</p>\n<table>\n<thead>\n<tr>\n<th>PyFstat parameter</th>\n<th></th>\n<th></th>\n<th>SNR estimation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>h0</td>\n<td>sqrtSX</td>\n<td>cosi</td>\n<td></td>\n</tr>\n<tr>\n<td>1</td>\n<td>[10,100]</td>\n<td>0</td>\n<td>[9.2,101]</td>\n</tr>\n<tr>\n<td>1</td>\n<td>[10,100]</td>\n<td>1 or -1</td>\n<td>[2.3,19]</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "# summary\n- Dataset\n- Normalization\n- PyFstat data generation / parameter range\n- Large Kernel and FFT\n- Train more with EMA\n- Etc\n\n# Dataset\n- test/: used for data augmentation during training (finally not applied)\n- train/ : validation data\n- Generate from PyFstat: data for training\n\n# Normalization\n```\nfrom scipy.stats import norm\ndef Fnormalize(X):\n    X /= X.sum(-2, keepdims=True)\n    return X\ndef Pnormalize(X):\n    n = np.prod(X.shape[-2:])\n    POS = min(int(n * 0.999), n - 10)\n    EXP = norm.ppf((POS + 1 - np.pi / 8) / (n - np.pi / 4 + 1))\n    scale = np.partition(X.flatten(), POS, -1)[POS]\n    X /= scale / EXP.astype(scale.dtype) ** 2\n    return X\ndef normalize(X):\n    X = (X[..., None].view(X.real.dtype) ** 2).sum(-1)\n    X = Fnormalize(X)\n    X = Pnormalize(X)\n    return X\n```\nThis data normalization has a simple theoretical background.\n`Fnormalize` : Changes the complex input to the sum of the squares of the real and imaginary parts of the elements, transforming the input into a `chi2` distribution.\n`Pnormalize`: Scales so that the square of the `POS`th largest number of the normal distribution is equal to the `POS`th largest number of the data.\nIt would be a good idea to interpret `Pnormalize` for the `chi2` distribution, but since this function only scales the input, I think you'll get similar performance.\nIn the same model, the performance is improved by changing the normalization technique of the inference step.\n|  |Private  |Public|\n| --- | --- | --- |\n|  Pnormalize only|0.766  |0.739|\n|  Fnormalize + Pnormalize|0.770  |0.751|\n\n# PyFstat data generation / parameter range\n\nBefore implement the dataset sampling code, I read the PyFstat library code and tracked the parameter range. However, `psi` and `phi` were not found.\nSurprisingly, only 5 days before the end of the competition, I found out about the proper range [here](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/373973), and the model performance improved.\n\n Setting 1(Gen1):\n```\ntstart: [630720013, 1861492413)\nF0: [45, 600)\nF1: 10 ^ truncnorm.isf(rng.uniform(), a=-100, b=3, loc=-15, scale=2)\nAlpha: [0, 2pi)\nDelta: [-pi/2, pi/2)\ncosi: [-1, 1)\npsi: [-1, 1)\nphi: [-1, 1)\nTsft: 1800\nSFTWindowType: \"tukey\"\nSFTWindowBeta: 0.0001\n```\n Setting 2(Gen2):\n```\ntstart: [630720013, 1861492413)\nF0: [45, 600)\nF1: 10 ^ truncnorm.isf(rng.uniform(), a=-100, b=3, loc=-15, scale=2)\nAlpha: [0, 2pi)\nDelta: [-pi/2, pi/2)\ncosi: [-1, 1)\npsi: [-pi/4, pi/4)\nphi: [0, 2pi)\nTsft: 1800\nSFTWindowType: \"tukey\"\nSFTWindowBeta: 0.0001\n```\n|  |Validation ROCAUC|Private  |Public|\n| --- | --- | --- | --- |\n|  exp/v0.2-1229-1 Best valid (Gen1)|0.8569|0.761  |0.747|\n|  exp/v0.2-1229-1 Last epoch (Gen1)|0.7278|  | |\n|  exp/v0.2-1230-2 Best valid (Gen2)|0.8662812499999999|0.759  |0.747|\n|  exp/v0.2-1230-2 Last epoch (Gen2)|0.8624|0.770  |0.756|\n\n#  Large Kernel and FFT\nI implemented [fft convolution](https://github.com/klae01/fft-conv-pytorch) to use large kernels in training. You can refer to the notebook that was released one week before the end of the competition. [link](https://www.kaggle.com/code/assign/g2net-large-kernel-inference-fft-conv2d)\n\n## Weight Decay\nIt has been observed that large kernels are prone to bias and are weak. I applied a strong weight decay of 0.001 to that layer.\n\n| |Validation ROCAUC|Private|Public|\n| --- | --- | --- | --- |\n|exp/v0.2-1229-1 weight decay: 1e-3 Best valid (Gen1) |0.8569|0.761|0.747|\n|exp/v0.2-1230-1 weight decay: 1e-4 Best valid (Gen1) |0.8538749999999999|0.756|0.743|\n\n## Dense stem\nSome performance improvement is achieved by changing the depth-wise convolution applied per L1/H1 detector channel of the model to a normal convolution.\nCompared to the best validation score, the gap is large, but the last epoch still performs better.\n|exp/v0.2-1231-1|Validation ROCAUC|Private|Public|\n| --- | --- | --- | --- |\n|Best valid (Gen2)|0.8635312500000001|0.762|0.748|\n|Last epoch (Gen2)|0.8532|0.771|0.760|\n\n# Train more with EMA\nI trained with 2x epochs with the same settings, but the model diverged. However, as the model with the highest validation score, it showed an improved score than before.\nTraining ended 3 hours before the end of the competition, and this is the last training.\n\n| |Private|Public|\n| --- | --- | --- |\n|Gen 1|0.761|0.747|\n|Gen 2|0.770|0.756|\n|Gen 2 + Dense stem|0.771|0.760|\n|Gen 2 + Dense stem + EMA|0.773|0.761|\n\n# Etc\n- I didn't use the idea of putting horizontal/vertical lines in the PyFstat data.\n- To avoid overfitting all hypotheses to the train data, the data belonging to `train/` were only analyzed at the statistical level. Might be a dumb choice :)\n- I have applied the idea of creating gaps in continuous inputs and filling them with noise, but the validation loss at the beginning of training is greatly improved, but the training becomes very unstable, and the training/validation/submission scores are all lower. It's an idea that was ultimately abandoned.\n- PyFstat is slow because they uses disk. This is the bottleneck. Use `tmpfs`.\n- The model was trained with 256 batches of 64k samples reused 32 times each.\n- The time spent implementing the tool was too long compared to the time spent improving performance. For example, an implementation that multiprocesses a pipeline that reuses and discards data and guarantees the same order for seeds / fft convolution optimization / etc.\n- Surprisingly, SNR ranges from training [6, 10]. Better than [5,10], better performance than [6, 15]. In the training phase, a low SNR range causes the model to diverge, while a high SNR range makes the model less discerning when no signal is present.\n\nNote: This is a rough comparison of the functions that investigate the SNR mentioned above.\n\n| PyFstat parameter ||| SNR estimation |\n| --- | --- | --- | --- |\n|h0|sqrtSX|cosi||\n|1|[10,100]|0|[9.2,101]|\n|1|[10,100]|1 or -1|[2.3,19]|\n",
      "votes": 6
    },
    {
      "id": 2095602,
      "postDate": "2023-01-11T13:51:04.897Z",
      "content": "<p><a href=\"https://www.kaggle.com/assign\" target=\"_blank\">@assign</a>  Did you do any sort of time alignment between H1 and L1?</p>",
      "rawMarkdown": "@assign  Did you do any sort of time alignment between H1 and L1?"
    },
    {
      "id": 2091425,
      "postDate": "2023-01-08T11:51:19.603Z",
      "content": "<p>Awesome analysis! Thank you!</p>",
      "rawMarkdown": "Awesome analysis! Thank you!"
    }
  ],
  "comments": [
    {
      "id": 2095602,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2023-01-11T13:51:04.897000",
      "content": "<p><a href=\"https://www.kaggle.com/assign\" target=\"_blank\">@assign</a>  Did you do any sort of time alignment between H1 and L1?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2091425,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2023-01-08T11:51:19.603000",
      "content": "<p>Awesome analysis! Thank you!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2091230": "# summary\n- Dataset\n- Normalization\n- PyFstat data generation / parameter range\n- Large Kernel and FFT\n- Train more with EMA\n- Etc\n\n# Dataset\n- test/: used for data augmentation during training (finally not applied)\n- train/ : validation data\n- Generate from PyFstat: data for training\n\n# Normalization\n```\nfrom scipy.stats import norm\ndef Fnormalize(X):\n    X /= X.sum(-2, keepdims=True)\n    return X\ndef Pnormalize(X):\n    n = np.prod(X.shape[-2:])\n    POS = min(int(n * 0.999), n - 10)\n    EXP = norm.ppf((POS + 1 - np.pi / 8) / (n - np.pi / 4 + 1))\n    scale = np.partition(X.flatten(), POS, -1)[POS]\n    X /= scale / EXP.astype(scale.dtype) ** 2\n    return X\ndef normalize(X):\n    X = (X[..., None].view(X.real.dtype) ** 2).sum(-1)\n    X = Fnormalize(X)\n    X = Pnormalize(X)\n    return X\n```\nThis data normalization has a simple theoretical background.\n`Fnormalize` : Changes the complex input to the sum of the squares of the real and imaginary parts of the elements, transforming the input into a `chi2` distribution.\n`Pnormalize`: Scales so that the square of the `POS`th largest number of the normal distribution is equal to the `POS`th largest number of the data.\nIt would be a good idea to interpret `Pnormalize` for the `chi2` distribution, but since this function only scales the input, I think you'll get similar performance.\nIn the same model, the performance is improved by changing the normalization technique of the inference step.\n|  |Private  |Public|\n| --- | --- | --- |\n|  Pnormalize only|0.766  |0.739|\n|  Fnormalize + Pnormalize|0.770  |0.751|\n\n# PyFstat data generation / parameter range\n\nBefore implement the dataset sampling code, I read the PyFstat library code and tracked the parameter range. However, `psi` and `phi` were not found.\nSurprisingly, only 5 days before the end of the competition, I found out about the proper range [here](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/373973), and the model performance improved.\n\n Setting 1(Gen1):\n```\ntstart: [630720013, 1861492413)\nF0: [45, 600)\nF1: 10 ^ truncnorm.isf(rng.uniform(), a=-100, b=3, loc=-15, scale=2)\nAlpha: [0, 2pi)\nDelta: [-pi/2, pi/2)\ncosi: [-1, 1)\npsi: [-1, 1)\nphi: [-1, 1)\nTsft: 1800\nSFTWindowType: \"tukey\"\nSFTWindowBeta: 0.0001\n```\n Setting 2(Gen2):\n```\ntstart: [630720013, 1861492413)\nF0: [45, 600)\nF1: 10 ^ truncnorm.isf(rng.uniform(), a=-100, b=3, loc=-15, scale=2)\nAlpha: [0, 2pi)\nDelta: [-pi/2, pi/2)\ncosi: [-1, 1)\npsi: [-pi/4, pi/4)\nphi: [0, 2pi)\nTsft: 1800\nSFTWindowType: \"tukey\"\nSFTWindowBeta: 0.0001\n```\n|  |Validation ROCAUC|Private  |Public|\n| --- | --- | --- | --- |\n|  exp/v0.2-1229-1 Best valid (Gen1)|0.8569|0.761  |0.747|\n|  exp/v0.2-1229-1 Last epoch (Gen1)|0.7278|  | |\n|  exp/v0.2-1230-2 Best valid (Gen2)|0.8662812499999999|0.759  |0.747|\n|  exp/v0.2-1230-2 Last epoch (Gen2)|0.8624|0.770  |0.756|\n\n#  Large Kernel and FFT\nI implemented [fft convolution](https://github.com/klae01/fft-conv-pytorch) to use large kernels in training. You can refer to the notebook that was released one week before the end of the competition. [link](https://www.kaggle.com/code/assign/g2net-large-kernel-inference-fft-conv2d)\n\n## Weight Decay\nIt has been observed that large kernels are prone to bias and are weak. I applied a strong weight decay of 0.001 to that layer.\n\n| |Validation ROCAUC|Private|Public|\n| --- | --- | --- | --- |\n|exp/v0.2-1229-1 weight decay: 1e-3 Best valid (Gen1) |0.8569|0.761|0.747|\n|exp/v0.2-1230-1 weight decay: 1e-4 Best valid (Gen1) |0.8538749999999999|0.756|0.743|\n\n## Dense stem\nSome performance improvement is achieved by changing the depth-wise convolution applied per L1/H1 detector channel of the model to a normal convolution.\nCompared to the best validation score, the gap is large, but the last epoch still performs better.\n|exp/v0.2-1231-1|Validation ROCAUC|Private|Public|\n| --- | --- | --- | --- |\n|Best valid (Gen2)|0.8635312500000001|0.762|0.748|\n|Last epoch (Gen2)|0.8532|0.771|0.760|\n\n# Train more with EMA\nI trained with 2x epochs with the same settings, but the model diverged. However, as the model with the highest validation score, it showed an improved score than before.\nTraining ended 3 hours before the end of the competition, and this is the last training.\n\n| |Private|Public|\n| --- | --- | --- |\n|Gen 1|0.761|0.747|\n|Gen 2|0.770|0.756|\n|Gen 2 + Dense stem|0.771|0.760|\n|Gen 2 + Dense stem + EMA|0.773|0.761|\n\n# Etc\n- I didn't use the idea of putting horizontal/vertical lines in the PyFstat data.\n- To avoid overfitting all hypotheses to the train data, the data belonging to `train/` were only analyzed at the statistical level. Might be a dumb choice :)\n- I have applied the idea of creating gaps in continuous inputs and filling them with noise, but the validation loss at the beginning of training is greatly improved, but the training becomes very unstable, and the training/validation/submission scores are all lower. It's an idea that was ultimately abandoned.\n- PyFstat is slow because they uses disk. This is the bottleneck. Use `tmpfs`.\n- The model was trained with 256 batches of 64k samples reused 32 times each.\n- The time spent implementing the tool was too long compared to the time spent improving performance. For example, an implementation that multiprocesses a pipeline that reuses and discards data and guarantees the same order for seeds / fft convolution optimization / etc.\n- Surprisingly, SNR ranges from training [6, 10]. Better than [5,10], better performance than [6, 15]. In the training phase, a low SNR range causes the model to diverge, while a high SNR range makes the model less discerning when no signal is present.\n\nNote: This is a rough comparison of the functions that investigate the SNR mentioned above.\n\n| PyFstat parameter ||| SNR estimation |\n| --- | --- | --- | --- |\n|h0|sqrtSX|cosi||\n|1|[10,100]|0|[9.2,101]|\n|1|[10,100]|1 or -1|[2.3,19]|\n",
    "2095602": "@assign  Did you do any sort of time alignment between H1 and L1?",
    "2091425": "Awesome analysis! Thank you!"
  }
}