{
  "id": 377661,
  "title": "Can someone teach me about a famous preprocessing ?",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/377661",
  "author_name": "Shibata",
  "post_date": "2023-01-12T08:57:12.144000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>This exciting and unique competition came to an end and I have learned a lot by myself during the competition period and also from top solutions. Then, I would like to ask about one often used preprocessing. </p>\n<blockquote>\n  <p>def normalize(X): <br>\n      X = (X[…, None].view(X.real.dtype) ** 2).sum(-1)<br>\n      POS = int(X.size * 0.99903)<br>\n      EXP = norm.ppf((POS + 0.4) / (X.size + 0.215))<br>\n      scale = np.partition(X.flatten(), POS, -1)[POS]<br>\n      X /= scale / EXP.astype(scale.dtype) ** 2<br>\n      return X</p>\n</blockquote>\n<p>This preprocessing was shown in this great <a href=\"https://www.kaggle.com/code/laeyoung/g2net-large-kernel-inference\" target=\"_blank\">notebook</a> and this solution <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/376724\" target=\"_blank\">post</a>.<br>\nI could understand power spectrum follows chi-squared distribution, but haven’t figured out about this normalization code. Could someone give me a detailed explanation or any material to see?<br>\nThanks in advance!</p>",
  "messages": [
    {
      "id": 2096778,
      "postDate": "2023-01-12T08:57:12.143Z",
      "content": "<p>This exciting and unique competition came to an end and I have learned a lot by myself during the competition period and also from top solutions. Then, I would like to ask about one often used preprocessing. </p>\n<blockquote>\n  <p>def normalize(X): <br>\n      X = (X[…, None].view(X.real.dtype) ** 2).sum(-1)<br>\n      POS = int(X.size * 0.99903)<br>\n      EXP = norm.ppf((POS + 0.4) / (X.size + 0.215))<br>\n      scale = np.partition(X.flatten(), POS, -1)[POS]<br>\n      X /= scale / EXP.astype(scale.dtype) ** 2<br>\n      return X</p>\n</blockquote>\n<p>This preprocessing was shown in this great <a href=\"https://www.kaggle.com/code/laeyoung/g2net-large-kernel-inference\" target=\"_blank\">notebook</a> and this solution <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/376724\" target=\"_blank\">post</a>.<br>\nI could understand power spectrum follows chi-squared distribution, but haven’t figured out about this normalization code. Could someone give me a detailed explanation or any material to see?<br>\nThanks in advance!</p>",
      "rawMarkdown": "This exciting and unique competition came to an end and I have learned a lot by myself during the competition period and also from top solutions. Then, I would like to ask about one often used preprocessing. \n>def normalize(X): \n    X = (X[..., None].view(X.real.dtype) ** 2).sum(-1)\n    POS = int(X.size * 0.99903)\n    EXP = norm.ppf((POS + 0.4) / (X.size + 0.215))\n    scale = np.partition(X.flatten(), POS, -1)[POS]\n    X /= scale / EXP.astype(scale.dtype) ** 2\n    return X\n\nThis preprocessing was shown in this great [notebook](https://www.kaggle.com/code/laeyoung/g2net-large-kernel-inference) and this solution [post](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/376724).\nI could understand power spectrum follows chi-squared distribution, but haven’t figured out about this normalization code. Could someone give me a detailed explanation or any material to see?\nThanks in advance!",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2096778": "This exciting and unique competition came to an end and I have learned a lot by myself during the competition period and also from top solutions. Then, I would like to ask about one often used preprocessing. \n>def normalize(X): \n    X = (X[..., None].view(X.real.dtype) ** 2).sum(-1)\n    POS = int(X.size * 0.99903)\n    EXP = norm.ppf((POS + 0.4) / (X.size + 0.215))\n    scale = np.partition(X.flatten(), POS, -1)[POS]\n    X /= scale / EXP.astype(scale.dtype) ** 2\n    return X\n\nThis preprocessing was shown in this great [notebook](https://www.kaggle.com/code/laeyoung/g2net-large-kernel-inference) and this solution [post](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/376724).\nI could understand power spectrum follows chi-squared distribution, but haven’t figured out about this normalization code. Could someone give me a detailed explanation or any material to see?\nThanks in advance!"
  }
}