{
  "id": 376602,
  "title": "Some questions and confusions after competition",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/376602",
  "author_name": "Jiawei Zhang",
  "post_date": "2023-01-07T09:26:41.571000",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First I want to thank the kaggle team and the competition organizers for the event. I'm an undergraduate student and it's my first time participating in kaggle competition. I am more than glad to learn so much knowledge from the competition.  And I'm trying to understand and learn from the top solutions in the days after the competition (also thanks to their selfless sharing). And I have two small questions that I could not figure out.</p>\n<p>① Many teams proposed to us Laeyoung's data preprocessing method.<br>\n<code>def normalize(X):</code><br>\n<code>X = (X[..., None].view(X.real.dtype) ** 2).sum(-1)</code><br>\n<code>POS = int(X.size * 0.99903)</code><br>\n<code>EXP = norm.ppf((POS + 0.4) / (X.size + 0.215))</code><br>\n<code>scale = np.partition(X.flatten(), POS, -1)[POS]</code><br>\n<code>X /= scale / EXP.astype(scale.dtype) ** 2</code><br>\n<code>return X</code><br>\nI can not understand how the lines of code work even after the competition, I only know that the first line is adding the real and imaginary parts squared. The later lines may imply some knowledge of probability theory?</p>\n<p>② From the beginning of the competition, everyone was using the power spectrum. I tried to use the amplitude spectrum early on, but the results would be slightly worse. Is there a signal processing theory that suggests the power spectrum will perform better than the amplitude spectrum or is just empirical?</p>\n<p>Thanks in advance for answering a student's questions.</p>",
  "messages": [
    {
      "id": 2090393,
      "postDate": "2023-01-07T09:26:41.570Z",
      "content": "<p>First I want to thank the kaggle team and the competition organizers for the event. I'm an undergraduate student and it's my first time participating in kaggle competition. I am more than glad to learn so much knowledge from the competition.  And I'm trying to understand and learn from the top solutions in the days after the competition (also thanks to their selfless sharing). And I have two small questions that I could not figure out.</p>\n<p>① Many teams proposed to us Laeyoung's data preprocessing method.<br>\n<code>def normalize(X):</code><br>\n<code>X = (X[..., None].view(X.real.dtype) ** 2).sum(-1)</code><br>\n<code>POS = int(X.size * 0.99903)</code><br>\n<code>EXP = norm.ppf((POS + 0.4) / (X.size + 0.215))</code><br>\n<code>scale = np.partition(X.flatten(), POS, -1)[POS]</code><br>\n<code>X /= scale / EXP.astype(scale.dtype) ** 2</code><br>\n<code>return X</code><br>\nI can not understand how the lines of code work even after the competition, I only know that the first line is adding the real and imaginary parts squared. The later lines may imply some knowledge of probability theory?</p>\n<p>② From the beginning of the competition, everyone was using the power spectrum. I tried to use the amplitude spectrum early on, but the results would be slightly worse. Is there a signal processing theory that suggests the power spectrum will perform better than the amplitude spectrum or is just empirical?</p>\n<p>Thanks in advance for answering a student's questions.</p>",
      "rawMarkdown": "First I want to thank the kaggle team and the competition organizers for the event. I'm an undergraduate student and it's my first time participating in kaggle competition. I am more than glad to learn so much knowledge from the competition.  And I'm trying to understand and learn from the top solutions in the days after the competition (also thanks to their selfless sharing). And I have two small questions that I could not figure out.\n\n① Many teams proposed to us Laeyoung's data preprocessing method.\n`def normalize(X):`\n`   X = (X[..., None].view(X.real.dtype) ** 2).sum(-1)`\n`   POS = int(X.size * 0.99903)`\n`   EXP = norm.ppf((POS + 0.4) / (X.size + 0.215))`\n`    scale = np.partition(X.flatten(), POS, -1)[POS]`\n`    X /= scale / EXP.astype(scale.dtype) ** 2`\n`   return X`\nI can not understand how the lines of code work even after the competition, I only know that the first line is adding the real and imaginary parts squared. The later lines may imply some knowledge of probability theory?\n\n② From the beginning of the competition, everyone was using the power spectrum. I tried to use the amplitude spectrum early on, but the results would be slightly worse. Is there a signal processing theory that suggests the power spectrum will perform better than the amplitude spectrum or is just empirical?\n\nThanks in advance for answering a student's questions.",
      "votes": 4
    },
    {
      "id": 2090841,
      "postDate": "2023-01-07T17:26:59.760Z",
      "content": "<p>I looked at the code when it was posted and was like What the heck…. :)<br>\nscale looks like an element located at 99.9 percentile, which is like robust max (maximum excluding the outliers)<br>\nEXP. norm.ppf((POS + 0.4) / (X.size + 0.215)) - this is (POS + 0.4) / (X.size + 0.215) percentile of normal distribution. I have zero understanding of what 0.215 and 0.4 coefficients mean, but (POS + 0.4) / (X.size + 0.215) is some number a bit less than 1, so EXP is a percentile of normal distribution. EXP.astype(scale.dtype) ** 2 - will be a percentile of our cauchi distribution (squared normal)<br>\nX /= scale / EXP.astype(scale.dtype) ** 2 - this like is effectively tries to normalize the data and bring it to theoretical cauchi from two normal distributions with 0 mean and std of 1<br>\nWhat do you guys think?</p>\n<p>As for taking the square root or not, I don't think there's a theoretical explanation why one should work better than another for CNNs. Taking a square root as a pre-processing step in image processing kind of compresses the dynamic range so that larger values are dampened, while smaller values are much less affected. Say if the initial range was 1 to 100, then after the sqrt it will be 1 to 10. I would expect the latter to be better, but you just can't assume anything with neural nets and experimenting is the best way to test things.</p>",
      "rawMarkdown": "I looked at the code when it was posted and was like What the heck.... :)\nscale looks like an element located at 99.9 percentile, which is like robust max (maximum excluding the outliers)\nEXP. norm.ppf((POS + 0.4) / (X.size + 0.215)) - this is (POS + 0.4) / (X.size + 0.215) percentile of normal distribution. I have zero understanding of what 0.215 and 0.4 coefficients mean, but (POS + 0.4) / (X.size + 0.215) is some number a bit less than 1, so EXP is a percentile of normal distribution. EXP.astype(scale.dtype) ** 2 - will be a percentile of our cauchi distribution (squared normal)\nX /= scale / EXP.astype(scale.dtype) ** 2 - this like is effectively tries to normalize the data and bring it to theoretical cauchi from two normal distributions with 0 mean and std of 1\nWhat do you guys think?\n\nAs for taking the square root or not, I don't think there's a theoretical explanation why one should work better than another for CNNs. Taking a square root as a pre-processing step in image processing kind of compresses the dynamic range so that larger values are dampened, while smaller values are much less affected. Say if the initial range was 1 to 100, then after the sqrt it will be 1 to 10. I would expect the latter to be better, but you just can't assume anything with neural nets and experimenting is the best way to test things.",
      "votes": 2,
      "replies": [
        {
          "id": 2091034,
          "postDate": "2023-01-08T00:12:41.987Z",
          "content": "<p>Truly Thank you, your answer really inspires me a lot.</p>",
          "rawMarkdown": "Truly Thank you, your answer really inspires me a lot."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2090841,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2023-01-07T17:26:59.760000",
      "content": "<p>I looked at the code when it was posted and was like What the heck…. :)<br>\nscale looks like an element located at 99.9 percentile, which is like robust max (maximum excluding the outliers)<br>\nEXP. norm.ppf((POS + 0.4) / (X.size + 0.215)) - this is (POS + 0.4) / (X.size + 0.215) percentile of normal distribution. I have zero understanding of what 0.215 and 0.4 coefficients mean, but (POS + 0.4) / (X.size + 0.215) is some number a bit less than 1, so EXP is a percentile of normal distribution. EXP.astype(scale.dtype) ** 2 - will be a percentile of our cauchi distribution (squared normal)<br>\nX /= scale / EXP.astype(scale.dtype) ** 2 - this like is effectively tries to normalize the data and bring it to theoretical cauchi from two normal distributions with 0 mean and std of 1<br>\nWhat do you guys think?</p>\n<p>As for taking the square root or not, I don't think there's a theoretical explanation why one should work better than another for CNNs. Taking a square root as a pre-processing step in image processing kind of compresses the dynamic range so that larger values are dampened, while smaller values are much less affected. Say if the initial range was 1 to 100, then after the sqrt it will be 1 to 10. I would expect the latter to be better, but you just can't assume anything with neural nets and experimenting is the best way to test things.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2091034,
          "author_name": "Jiawei Zhang",
          "author_url": "",
          "post_date": "2023-01-08T00:12:41.987000",
          "content": "<p>Truly Thank you, your answer really inspires me a lot.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2090393": "First I want to thank the kaggle team and the competition organizers for the event. I'm an undergraduate student and it's my first time participating in kaggle competition. I am more than glad to learn so much knowledge from the competition.  And I'm trying to understand and learn from the top solutions in the days after the competition (also thanks to their selfless sharing). And I have two small questions that I could not figure out.\n\n① Many teams proposed to us Laeyoung's data preprocessing method.\n`def normalize(X):`\n`   X = (X[..., None].view(X.real.dtype) ** 2).sum(-1)`\n`   POS = int(X.size * 0.99903)`\n`   EXP = norm.ppf((POS + 0.4) / (X.size + 0.215))`\n`    scale = np.partition(X.flatten(), POS, -1)[POS]`\n`    X /= scale / EXP.astype(scale.dtype) ** 2`\n`   return X`\nI can not understand how the lines of code work even after the competition, I only know that the first line is adding the real and imaginary parts squared. The later lines may imply some knowledge of probability theory?\n\n② From the beginning of the competition, everyone was using the power spectrum. I tried to use the amplitude spectrum early on, but the results would be slightly worse. Is there a signal processing theory that suggests the power spectrum will perform better than the amplitude spectrum or is just empirical?\n\nThanks in advance for answering a student's questions.",
    "2090841": "I looked at the code when it was posted and was like What the heck.... :)\nscale looks like an element located at 99.9 percentile, which is like robust max (maximum excluding the outliers)\nEXP. norm.ppf((POS + 0.4) / (X.size + 0.215)) - this is (POS + 0.4) / (X.size + 0.215) percentile of normal distribution. I have zero understanding of what 0.215 and 0.4 coefficients mean, but (POS + 0.4) / (X.size + 0.215) is some number a bit less than 1, so EXP is a percentile of normal distribution. EXP.astype(scale.dtype) ** 2 - will be a percentile of our cauchi distribution (squared normal)\nX /= scale / EXP.astype(scale.dtype) ** 2 - this like is effectively tries to normalize the data and bring it to theoretical cauchi from two normal distributions with 0 mean and std of 1\nWhat do you guys think?\n\nAs for taking the square root or not, I don't think there's a theoretical explanation why one should work better than another for CNNs. Taking a square root as a pre-processing step in image processing kind of compresses the dynamic range so that larger values are dampened, while smaller values are much less affected. Say if the initial range was 1 to 100, then after the sqrt it will be 1 to 10. I would expect the latter to be better, but you just can't assume anything with neural nets and experimenting is the best way to test things."
  }
}