{
  "id": 364537,
  "title": "Input SNR question",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/364537",
  "author_name": "Husam Elfadil",
  "post_date": "2022-11-07T04:53:26.073000",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The data description mentions \"The typical amplitudes of the resulting signals are one or two orders of magnitude lower than the amplitude of the detector noise.\"</p>\n<p>Does this that we should assume the SNR is known and it can be one of these two values? Or is this just a general description of CW and we should assume unknown SNR?</p>",
  "messages": [
    {
      "id": 2020350,
      "postDate": "2022-11-07T11:38:34.110Z",
      "content": "<p>The distribution of amplitudes of the signals injected in the samples is not provided, but you already have a hint of the range (one or two magnitudes lower). I interpret it as a general description since those are indeed the ranges expected/searched for in real CW searches if I am not wrong. So I would assume unknown SNR / depth.</p>\n<p>Anyways if you generate data yourself and keep track of the depth (sqrtSX / h0) used in each positive sample, you will quickly learn that image classifiers have trouble with deeper signals (obviously), and then it's a matter of trying to go as deep as possible. Furthermore, since the train data contains 400 positive samples and none have real non-stationary noise or any other glitches, you can look at your CV AUC score with generated data versus AUC obtained predicting the kaggle train set, and infer a bit how deep are signals there. And by the way, test set seems to be even harder in terms of depth, even for the samples with generated noise.</p>",
      "rawMarkdown": "The distribution of amplitudes of the signals injected in the samples is not provided, but you already have a hint of the range (one or two magnitudes lower). I interpret it as a general description since those are indeed the ranges expected/searched for in real CW searches if I am not wrong. So I would assume unknown SNR / depth.\n\nAnyways if you generate data yourself and keep track of the depth (sqrtSX / h0) used in each positive sample, you will quickly learn that image classifiers have trouble with deeper signals (obviously), and then it's a matter of trying to go as deep as possible. Furthermore, since the train data contains 400 positive samples and none have real non-stationary noise or any other glitches, you can look at your CV AUC score with generated data versus AUC obtained predicting the kaggle train set, and infer a bit how deep are signals there. And by the way, test set seems to be even harder in terms of depth, even for the samples with generated noise.",
      "votes": 7,
      "replies": [
        {
          "id": 2023108,
          "postDate": "2022-11-09T14:36:36.757Z",
          "content": "<p>Better explained impossible. I'm agree in all, and thats the main point in this challenge. Awesome post! Congrats :)</p>",
          "rawMarkdown": "Better explained impossible. I'm agree in all, and thats the main point in this challenge. Awesome post! Congrats :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2019937,
      "postDate": "2022-11-07T04:53:26.073Z",
      "content": "<p>The data description mentions \"The typical amplitudes of the resulting signals are one or two orders of magnitude lower than the amplitude of the detector noise.\"</p>\n<p>Does this that we should assume the SNR is known and it can be one of these two values? Or is this just a general description of CW and we should assume unknown SNR?</p>",
      "rawMarkdown": "The data description mentions \"The typical amplitudes of the resulting signals are one or two orders of magnitude lower than the amplitude of the detector noise.\"\n\nDoes this that we should assume the SNR is known and it can be one of these two values? Or is this just a general description of CW and we should assume unknown SNR?",
      "votes": 4
    },
    {
      "id": 2020640,
      "postDate": "2022-11-07T17:04:29.653Z",
      "content": "<p>Ok thanks for the elaboration.</p>",
      "rawMarkdown": "Ok thanks for the elaboration."
    }
  ],
  "comments": [
    {
      "id": 2020350,
      "author_name": "Victor Gonzalez",
      "author_url": "",
      "post_date": "2022-11-07T11:38:34.110000",
      "content": "<p>The distribution of amplitudes of the signals injected in the samples is not provided, but you already have a hint of the range (one or two magnitudes lower). I interpret it as a general description since those are indeed the ranges expected/searched for in real CW searches if I am not wrong. So I would assume unknown SNR / depth.</p>\n<p>Anyways if you generate data yourself and keep track of the depth (sqrtSX / h0) used in each positive sample, you will quickly learn that image classifiers have trouble with deeper signals (obviously), and then it's a matter of trying to go as deep as possible. Furthermore, since the train data contains 400 positive samples and none have real non-stationary noise or any other glitches, you can look at your CV AUC score with generated data versus AUC obtained predicting the kaggle train set, and infer a bit how deep are signals there. And by the way, test set seems to be even harder in terms of depth, even for the samples with generated noise.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2023108,
          "author_name": "Zollkron",
          "author_url": "",
          "post_date": "2022-11-09T14:36:36.757000",
          "content": "<p>Better explained impossible. I'm agree in all, and thats the main point in this challenge. Awesome post! Congrats :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2020640,
      "author_name": "Husam Elfadil",
      "author_url": "",
      "post_date": "2022-11-07T17:04:29.653000",
      "content": "<p>Ok thanks for the elaboration.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2020350": "The distribution of amplitudes of the signals injected in the samples is not provided, but you already have a hint of the range (one or two magnitudes lower). I interpret it as a general description since those are indeed the ranges expected/searched for in real CW searches if I am not wrong. So I would assume unknown SNR / depth.\n\nAnyways if you generate data yourself and keep track of the depth (sqrtSX / h0) used in each positive sample, you will quickly learn that image classifiers have trouble with deeper signals (obviously), and then it's a matter of trying to go as deep as possible. Furthermore, since the train data contains 400 positive samples and none have real non-stationary noise or any other glitches, you can look at your CV AUC score with generated data versus AUC obtained predicting the kaggle train set, and infer a bit how deep are signals there. And by the way, test set seems to be even harder in terms of depth, even for the samples with generated noise.",
    "2019937": "The data description mentions \"The typical amplitudes of the resulting signals are one or two orders of magnitude lower than the amplitude of the detector noise.\"\n\nDoes this that we should assume the SNR is known and it can be one of these two values? Or is this just a general description of CW and we should assume unknown SNR?",
    "2020640": "Ok thanks for the elaboration."
  }
}