{
  "id": 363460,
  "title": "Weird property of std(spectrogram) in train data",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363460",
  "author_name": "Josef Slavicek",
  "post_date": "2022-11-01T17:45:55.097000",
  "votes": 11,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello <a href=\"https://www.kaggle.com/rodrigotenorio\" target=\"_blank\">@rodrigotenorio</a> , <a href=\"https://www.kaggle.com/michaeljwill\" target=\"_blank\">@michaeljwill</a> , I discovered very surprising property of train data. If I compute std of spectrogram in complex64, I'm getting almost allways the same value of 1.4973568266036865e-22. If I do the same computation in complex128, the values starts do differ - and they are partly useful as indicator of target label:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F59264%2Fe913756c71cfdb5b926f70543095733c%2F__results___6_1.png?generation=1667324396733833&amp;alt=media\" alt=\"\"></p>\n<p>The image is taken from <a href=\"https://www.kaggle.com/code/josefslavicek/g2net-weird-behavior-of-std-spectrogram\" target=\"_blank\">notebook</a> which demonstrate the behavior.</p>\n<p>My guess is, that most of training data was generated with single value of sqrtSX. Is this property of data known and considered as OK? For me, it feels almost like target label leak :-)</p>",
  "messages": [
    {
      "id": 2013187,
      "postDate": "2022-11-01T17:45:55.097Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/rodrigotenorio\" target=\"_blank\">@rodrigotenorio</a> , <a href=\"https://www.kaggle.com/michaeljwill\" target=\"_blank\">@michaeljwill</a> , I discovered very surprising property of train data. If I compute std of spectrogram in complex64, I'm getting almost allways the same value of 1.4973568266036865e-22. If I do the same computation in complex128, the values starts do differ - and they are partly useful as indicator of target label:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F59264%2Fe913756c71cfdb5b926f70543095733c%2F__results___6_1.png?generation=1667324396733833&amp;alt=media\" alt=\"\"></p>\n<p>The image is taken from <a href=\"https://www.kaggle.com/code/josefslavicek/g2net-weird-behavior-of-std-spectrogram\" target=\"_blank\">notebook</a> which demonstrate the behavior.</p>\n<p>My guess is, that most of training data was generated with single value of sqrtSX. Is this property of data known and considered as OK? For me, it feels almost like target label leak :-)</p>",
      "rawMarkdown": "Hello @rodrigotenorio , @michaeljwill , I discovered very surprising property of train data. If I compute std of spectrogram in complex64, I'm getting almost allways the same value of 1.4973568266036865e-22. If I do the same computation in complex128, the values starts do differ - and they are partly useful as indicator of target label:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F59264%2Fe913756c71cfdb5b926f70543095733c%2F__results___6_1.png?generation=1667324396733833&alt=media)\n\nThe image is taken from [notebook](https://www.kaggle.com/code/josefslavicek/g2net-weird-behavior-of-std-spectrogram) which demonstrate the behavior.\n\nMy guess is, that most of training data was generated with single value of sqrtSX. Is this property of data known and considered as OK? For me, it feels almost like target label leak :-)",
      "votes": 11
    },
    {
      "id": 2014503,
      "postDate": "2022-11-02T15:59:36.067Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/josefslavicek\" target=\"_blank\">@josefslavicek</a> , all, thank you very much for publicly starting this thread.</p>\n<p>We (@michaeljwill and myself) had a brief discussion about this, I hope the following clarifies what is going on.</p>\n<p>tl;dr This is not a data leak, we understand the reason this is happening and it is expected. Nothing to worry about.</p>\n<blockquote>\n  <p>and they are partly useful as indicator of target label</p>\n</blockquote>\n<p>Yes, this is expected: stdev is computed from the data, and very strong signals will deviate the value away from the \"noise-only\" value. This should already tell you how \"strong\" signals in the test set are, by the way :) .</p>\n<blockquote>\n  <p>Around 80% of test data seems to be affected by this</p>\n</blockquote>\n<p>This is also well understood. We wanted to give as much real data as possible in this challenge, but we didn't have a reliable way of generating more than what the LIGO detectors produced during the run, and we did not want to run into issues about having too much correlation among different samples. For this reason, we decided to generate synthetic data using Gaussian noise, for which a PSD (or ASD, or <code>sqrtSX</code>, they are all essentially equivalent in this context) needs to be specified. </p>\n<blockquote>\n  <p>It would  us sidechannel which is not present in real GW search data.</p>\n</blockquote>\n<p>This is not a problem in CW searches. For \"clean\" (relatively quiet) bands we are able to estimate the PSD of the noise with quite good accuracy, and that is sufficient for us to actually perform our analysis. The relevant bit is that, regardless of that PSD, the relevant quantity when it comes to detecting a signal is the SNR (which is roughly a ratio between signal amplitude  <code>h0</code> and noise amplitude <code>ASD</code>). Signals are usually so weak (read \"SNR is low\") that we cannot use PSD as a trigger for signals; in fact, high PSDs are usually related to noise artifacts.</p>\n<p>All in all, the test set portraits a data distribution which is a close as possible to a general CW search as we could generate. There is a significant fraction of simulated data, as we state in the data page, and that is due to the impossibility of properly and reliably generating \"real data\" on demand, but that does not drift this competition away from the sort of \"realistic\" setup we want you to focus on.</p>\n<p>We are happy to clarify any other concerns, so please keep raising any issues you may be concerned :)</p>\n<p>Cheers,</p>",
      "rawMarkdown": "Hi @josefslavicek , all, thank you very much for publicly starting this thread.\n\nWe (@michaeljwill and myself) had a brief discussion about this, I hope the following clarifies what is going on.\n\ntl;dr This is not a data leak, we understand the reason this is happening and it is expected. Nothing to worry about.\n\n> and they are partly useful as indicator of target label\n\nYes, this is expected: stdev is computed from the data, and very strong signals will deviate the value away from the \"noise-only\" value. This should already tell you how \"strong\" signals in the test set are, by the way :) .\n\n > Around 80% of test data seems to be affected by this\n\nThis is also well understood. We wanted to give as much real data as possible in this challenge, but we didn't have a reliable way of generating more than what the LIGO detectors produced during the run, and we did not want to run into issues about having too much correlation among different samples. For this reason, we decided to generate synthetic data using Gaussian noise, for which a PSD (or ASD, or `sqrtSX`, they are all essentially equivalent in this context) needs to be specified. \n\n> It would  us sidechannel which is not present in real GW search data.\n\nThis is not a problem in CW searches. For \"clean\" (relatively quiet) bands we are able to estimate the PSD of the noise with quite good accuracy, and that is sufficient for us to actually perform our analysis. The relevant bit is that, regardless of that PSD, the relevant quantity when it comes to detecting a signal is the SNR (which is roughly a ratio between signal amplitude  `h0` and noise amplitude `ASD`). Signals are usually so weak (read \"SNR is low\") that we cannot use PSD as a trigger for signals; in fact, high PSDs are usually related to noise artifacts.\n\nAll in all, the test set portraits a data distribution which is a close as possible to a general CW search as we could generate. There is a significant fraction of simulated data, as we state in the data page, and that is due to the impossibility of properly and reliably generating \"real data\" on demand, but that does not drift this competition away from the sort of \"realistic\" setup we want you to focus on.\n\nWe are happy to clarify any other concerns, so please keep raising any issues you may be concerned :)\n\nCheers,",
      "votes": 3,
      "replies": [
        {
          "id": 2014681,
          "postDate": "2022-11-02T18:40:17.017Z",
          "content": "<p>OK, thanks for clarification.</p>",
          "rawMarkdown": "OK, thanks for clarification."
        }
      ]
    },
    {
      "id": 2013238,
      "postDate": "2022-11-01T18:29:38.083Z",
      "content": "<p>\"My guess is, that most of training data was generated with single value of sqrtSX. \"<br>\nYep. 5e-24 or so.</p>",
      "rawMarkdown": "\"My guess is, that most of training data was generated with single value of sqrtSX. \"\nYep. 5e-24 or so.",
      "votes": 3
    },
    {
      "id": 2014400,
      "postDate": "2022-11-02T14:34:27.147Z",
      "content": "<p>FYI, I have evaluated using predictions from my NN model and found that most of the samples where the std in complex128 judged outlier are easily true positives.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4759954%2F13d4e2af0594b5c2920159ca68fae6c7%2F20221102.png?generation=1667398280554223&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "FYI, I have evaluated using predictions from my NN model and found that most of the samples where the std in complex128 judged outlier are easily true positives.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4759954%2F13d4e2af0594b5c2920159ca68fae6c7%2F20221102.png?generation=1667398280554223&alt=media)",
      "votes": 1
    },
    {
      "id": 2013483,
      "postDate": "2022-11-02T00:30:22.357Z",
      "content": "<p>Getting same std with complex64 is a property of float32<br>\n<a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361312\" target=\"_blank\">https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361312</a></p>\n<p>Signal contribute to std (total power) and correlated with label, but I don't think it is a strong estimator, and no worry for a leak. Test data also contain real noise, which are far from the train sqrtSX.</p>",
      "rawMarkdown": "Getting same std with complex64 is a property of float32\nhttps://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361312\n\nSignal contribute to std (total power) and correlated with label, but I don't think it is a strong estimator, and no worry for a leak. Test data also contain real noise, which are far from the train sqrtSX.",
      "votes": 1,
      "replies": [
        {
          "id": 2013747,
          "postDate": "2022-11-02T05:38:54.063Z",
          "content": "<blockquote>\n  <p>Signal contribute to std (total power) and correlated with label</p>\n</blockquote>\n<p>Yes, but if you generate big part of your dataset with just single value of sqrtSX, you get something what IMO you <strong>should not</strong> have: very good prior knowledge of <code>std(spectrogram)</code> for negative datapoints. You know it up to 3 decimal places. If people around LIGO have similarly good knowledge of noise in their detector, then everything is OK. If they don't, then this competition will be still fair, but not as useful as it could be. It would give us sidechannel which is not present in real GW search data. </p>\n<p>Around 80% of test data seems to be affected by this.</p>",
          "rawMarkdown": "> Signal contribute to std (total power) and correlated with label\n\nYes, but if you generate big part of your dataset with just single value of sqrtSX, you get something what IMO you **should not** have: very good prior knowledge of `std(spectrogram)` for negative datapoints. You know it up to 3 decimal places. If people around LIGO have similarly good knowledge of noise in their detector, then everything is OK. If they don't, then this competition will be still fair, but not as useful as it could be. It would give us sidechannel which is not present in real GW search data. \n\nAround 80% of test data seems to be affected by this.",
          "votes": 1
        },
        {
          "id": 2013992,
          "postDate": "2022-11-02T08:40:47.510Z",
          "content": "<p>I see, I agree with that. We know the statistics of the simulated noise very well, and those 80% are easier than real data.</p>",
          "rawMarkdown": "I see, I agree with that. We know the statistics of the simulated noise very well, and those 80% are easier than real data."
        },
        {
          "id": 2014040,
          "postDate": "2022-11-02T09:16:28.113Z",
          "content": "<p>Don't pay too much attention to the train set. Test set is very different.</p>",
          "rawMarkdown": "Don't pay too much attention to the train set. Test set is very different."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2014503,
      "author_name": "Rodrigo Tenorio",
      "author_url": "",
      "post_date": "2022-11-02T15:59:36.067000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/josefslavicek\" target=\"_blank\">@josefslavicek</a> , all, thank you very much for publicly starting this thread.</p>\n<p>We (@michaeljwill and myself) had a brief discussion about this, I hope the following clarifies what is going on.</p>\n<p>tl;dr This is not a data leak, we understand the reason this is happening and it is expected. Nothing to worry about.</p>\n<blockquote>\n  <p>and they are partly useful as indicator of target label</p>\n</blockquote>\n<p>Yes, this is expected: stdev is computed from the data, and very strong signals will deviate the value away from the \"noise-only\" value. This should already tell you how \"strong\" signals in the test set are, by the way :) .</p>\n<blockquote>\n  <p>Around 80% of test data seems to be affected by this</p>\n</blockquote>\n<p>This is also well understood. We wanted to give as much real data as possible in this challenge, but we didn't have a reliable way of generating more than what the LIGO detectors produced during the run, and we did not want to run into issues about having too much correlation among different samples. For this reason, we decided to generate synthetic data using Gaussian noise, for which a PSD (or ASD, or <code>sqrtSX</code>, they are all essentially equivalent in this context) needs to be specified. </p>\n<blockquote>\n  <p>It would  us sidechannel which is not present in real GW search data.</p>\n</blockquote>\n<p>This is not a problem in CW searches. For \"clean\" (relatively quiet) bands we are able to estimate the PSD of the noise with quite good accuracy, and that is sufficient for us to actually perform our analysis. The relevant bit is that, regardless of that PSD, the relevant quantity when it comes to detecting a signal is the SNR (which is roughly a ratio between signal amplitude  <code>h0</code> and noise amplitude <code>ASD</code>). Signals are usually so weak (read \"SNR is low\") that we cannot use PSD as a trigger for signals; in fact, high PSDs are usually related to noise artifacts.</p>\n<p>All in all, the test set portraits a data distribution which is a close as possible to a general CW search as we could generate. There is a significant fraction of simulated data, as we state in the data page, and that is due to the impossibility of properly and reliably generating \"real data\" on demand, but that does not drift this competition away from the sort of \"realistic\" setup we want you to focus on.</p>\n<p>We are happy to clarify any other concerns, so please keep raising any issues you may be concerned :)</p>\n<p>Cheers,</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2014681,
          "author_name": "Josef Slavicek",
          "author_url": "",
          "post_date": "2022-11-02T18:40:17.017000",
          "content": "<p>OK, thanks for clarification.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2013238,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2022-11-01T18:29:38.083000",
      "content": "<p>\"My guess is, that most of training data was generated with single value of sqrtSX. \"<br>\nYep. 5e-24 or so.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2014400,
      "author_name": "hosuke",
      "author_url": "",
      "post_date": "2022-11-02T14:34:27.147000",
      "content": "<p>FYI, I have evaluated using predictions from my NN model and found that most of the samples where the std in complex128 judged outlier are easily true positives.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4759954%2F13d4e2af0594b5c2920159ca68fae6c7%2F20221102.png?generation=1667398280554223&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2013483,
      "author_name": "🐢 Jun Koda",
      "author_url": "",
      "post_date": "2022-11-02T00:30:22.357000",
      "content": "<p>Getting same std with complex64 is a property of float32<br>\n<a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361312\" target=\"_blank\">https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361312</a></p>\n<p>Signal contribute to std (total power) and correlated with label, but I don't think it is a strong estimator, and no worry for a leak. Test data also contain real noise, which are far from the train sqrtSX.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2013747,
          "author_name": "Josef Slavicek",
          "author_url": "",
          "post_date": "2022-11-02T05:38:54.063000",
          "content": "<blockquote>\n  <p>Signal contribute to std (total power) and correlated with label</p>\n</blockquote>\n<p>Yes, but if you generate big part of your dataset with just single value of sqrtSX, you get something what IMO you <strong>should not</strong> have: very good prior knowledge of <code>std(spectrogram)</code> for negative datapoints. You know it up to 3 decimal places. If people around LIGO have similarly good knowledge of noise in their detector, then everything is OK. If they don't, then this competition will be still fair, but not as useful as it could be. It would give us sidechannel which is not present in real GW search data. </p>\n<p>Around 80% of test data seems to be affected by this.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2013992,
          "author_name": "🐢 Jun Koda",
          "author_url": "",
          "post_date": "2022-11-02T08:40:47.510000",
          "content": "<p>I see, I agree with that. We know the statistics of the simulated noise very well, and those 80% are easier than real data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2014040,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2022-11-02T09:16:28.113000",
          "content": "<p>Don't pay too much attention to the train set. Test set is very different.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2013187": "Hello @rodrigotenorio , @michaeljwill , I discovered very surprising property of train data. If I compute std of spectrogram in complex64, I'm getting almost allways the same value of 1.4973568266036865e-22. If I do the same computation in complex128, the values starts do differ - and they are partly useful as indicator of target label:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F59264%2Fe913756c71cfdb5b926f70543095733c%2F__results___6_1.png?generation=1667324396733833&alt=media)\n\nThe image is taken from [notebook](https://www.kaggle.com/code/josefslavicek/g2net-weird-behavior-of-std-spectrogram) which demonstrate the behavior.\n\nMy guess is, that most of training data was generated with single value of sqrtSX. Is this property of data known and considered as OK? For me, it feels almost like target label leak :-)",
    "2014503": "Hi @josefslavicek , all, thank you very much for publicly starting this thread.\n\nWe (@michaeljwill and myself) had a brief discussion about this, I hope the following clarifies what is going on.\n\ntl;dr This is not a data leak, we understand the reason this is happening and it is expected. Nothing to worry about.\n\n> and they are partly useful as indicator of target label\n\nYes, this is expected: stdev is computed from the data, and very strong signals will deviate the value away from the \"noise-only\" value. This should already tell you how \"strong\" signals in the test set are, by the way :) .\n\n > Around 80% of test data seems to be affected by this\n\nThis is also well understood. We wanted to give as much real data as possible in this challenge, but we didn't have a reliable way of generating more than what the LIGO detectors produced during the run, and we did not want to run into issues about having too much correlation among different samples. For this reason, we decided to generate synthetic data using Gaussian noise, for which a PSD (or ASD, or `sqrtSX`, they are all essentially equivalent in this context) needs to be specified. \n\n> It would  us sidechannel which is not present in real GW search data.\n\nThis is not a problem in CW searches. For \"clean\" (relatively quiet) bands we are able to estimate the PSD of the noise with quite good accuracy, and that is sufficient for us to actually perform our analysis. The relevant bit is that, regardless of that PSD, the relevant quantity when it comes to detecting a signal is the SNR (which is roughly a ratio between signal amplitude  `h0` and noise amplitude `ASD`). Signals are usually so weak (read \"SNR is low\") that we cannot use PSD as a trigger for signals; in fact, high PSDs are usually related to noise artifacts.\n\nAll in all, the test set portraits a data distribution which is a close as possible to a general CW search as we could generate. There is a significant fraction of simulated data, as we state in the data page, and that is due to the impossibility of properly and reliably generating \"real data\" on demand, but that does not drift this competition away from the sort of \"realistic\" setup we want you to focus on.\n\nWe are happy to clarify any other concerns, so please keep raising any issues you may be concerned :)\n\nCheers,",
    "2013238": "\"My guess is, that most of training data was generated with single value of sqrtSX. \"\nYep. 5e-24 or so.",
    "2014400": "FYI, I have evaluated using predictions from my NN model and found that most of the samples where the std in complex128 judged outlier are easily true positives.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4759954%2F13d4e2af0594b5c2920159ca68fae6c7%2F20221102.png?generation=1667398280554223&alt=media)",
    "2013483": "Getting same std with complex64 is a property of float32\nhttps://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361312\n\nSignal contribute to std (total power) and correlated with label, but I don't think it is a strong estimator, and no worry for a leak. Test data also contain real noise, which are far from the train sqrtSX."
  }
}