{
  "id": 374161,
  "title": "Preprocessing Data",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/374161",
  "author_name": "Aaftaab V",
  "post_date": "2022-12-25T17:39:21.628000",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The most common method I have seen for preprocessing data before feeding it into a CNN, is to square and add real and imaginary parts, multiply each pixel with 1e22, takethe  average or summation of data from both detectors, to get finafinal image, and then resize that to as many channels as needed. I was just wondering if there are any other preprocessing methods in use. Also, how good is the above preprocessing  step wrt domain/physics knowledge? <br>\nAsking because saving sfts from pyfstat is taking tootoo much precious disk space. I wonder if I can preprocess sfts while generating data and save finafinal data as images or npy files. Is anyone doing the above hack?</p>",
  "messages": [
    {
      "id": 2075647,
      "postDate": "2022-12-25T17:39:21.630Z",
      "content": "<p>The most common method I have seen for preprocessing data before feeding it into a CNN, is to square and add real and imaginary parts, multiply each pixel with 1e22, takethe  average or summation of data from both detectors, to get finafinal image, and then resize that to as many channels as needed. I was just wondering if there are any other preprocessing methods in use. Also, how good is the above preprocessing  step wrt domain/physics knowledge? <br>\nAsking because saving sfts from pyfstat is taking tootoo much precious disk space. I wonder if I can preprocess sfts while generating data and save finafinal data as images or npy files. Is anyone doing the above hack?</p>",
      "rawMarkdown": "The most common method I have seen for preprocessing data before feeding it into a CNN, is to square and add real and imaginary parts, multiply each pixel with 1e22, takethe  average or summation of data from both detectors, to get finafinal image, and then resize that to as many channels as needed. I was just wondering if there are any other preprocessing methods in use. Also, how good is the above preprocessing  step wrt domain/physics knowledge? \nAsking because saving sfts from pyfstat is taking tootoo much precious disk space. I wonder if I can preprocess sfts while generating data and save finafinal data as images or npy files. Is anyone doing the above hack?",
      "votes": 3
    },
    {
      "id": 2080536,
      "postDate": "2022-12-30T08:58:48.953Z",
      "content": "<p>I can point out at least two more ways to preprocess the data.<br>\n1) \"Correlation\". Summing the squares of real and imaginary parts of the data is equal to the product of the signal and its complex conjugation. But we have two signals, and we can take the product of the signal on L1 detector and the complex conjugation of the signal on H1 detector. Then we apply some kind of averaging. The result will be their mutual spectrum, which is a fourier transform of correlation function. As we hope the signals on the two detectors are correlated, this can be used as \"the third channel\" of the image fed to CNN.</p>\n<p>2) \"Zebra approach\". It was proposed during the last year competition and it was reported to lead to good results (though, I didn't try it myself). We can mix the data from two detectors consequently: one time sample from L1, than another time sample from H1 and so on. </p>",
      "rawMarkdown": "I can point out at least two more ways to preprocess the data.\n1) \"Correlation\". Summing the squares of real and imaginary parts of the data is equal to the product of the signal and its complex conjugation. But we have two signals, and we can take the product of the signal on L1 detector and the complex conjugation of the signal on H1 detector. Then we apply some kind of averaging. The result will be their mutual spectrum, which is a fourier transform of correlation function. As we hope the signals on the two detectors are correlated, this can be used as \"the third channel\" of the image fed to CNN.\n\n2) \"Zebra approach\". It was proposed during the last year competition and it was reported to lead to good results (though, I didn't try it myself). We can mix the data from two detectors consequently: one time sample from L1, than another time sample from H1 and so on. ",
      "votes": 1
    },
    {
      "id": 2079355,
      "postDate": "2022-12-29T09:09:48.103Z",
      "content": "<p>Actually, the trimming and combining data from both detectors is already suboptimal. If you look at the timestamps of the two detectors measurements - they are not linear. If I understand correctly, this means the two detectors are not synchronized and combining the data is not trivial. To do it better - we could analyze each timestamp and try to resample them to get constant intervals for each detector separately. Then the combination could work much better, right? Because we would be combining data from exactly the same moment.</p>\n<p>Of course to resample the data consistently we would need to approximate the data that is missing, using a linear approximation or sth - and this could lead to errors.</p>\n<p>Is this something that you would be thinking about?</p>",
      "rawMarkdown": "Actually, the trimming and combining data from both detectors is already suboptimal. If you look at the timestamps of the two detectors measurements - they are not linear. If I understand correctly, this means the two detectors are not synchronized and combining the data is not trivial. To do it better - we could analyze each timestamp and try to resample them to get constant intervals for each detector separately. Then the combination could work much better, right? Because we would be combining data from exactly the same moment.\n\nOf course to resample the data consistently we would need to approximate the data that is missing, using a linear approximation or sth - and this could lead to errors.\n\nIs this something that you would be thinking about?",
      "votes": 1,
      "replies": [
        {
          "id": 2079562,
          "postDate": "2022-12-29T12:23:53.420Z",
          "content": "<p>The point you raised is actually good, we are not taking care to synchronize timestamps in real data (I understand, this is mostly present only in test dataset), but this should not be a problem with data generation, since, we (at least I am) are generating both detector's data at a time, and with same duration/timestamps. Maybe this should be taken care of, only while predicting after training model.</p>",
          "rawMarkdown": "The point you raised is actually good, we are not taking care to synchronize timestamps in real data (I understand, this is mostly present only in test dataset), but this should not be a problem with data generation, since, we (at least I am) are generating both detector's data at a time, and with same duration/timestamps. Maybe this should be taken care of, only while predicting after training model."
        }
      ]
    },
    {
      "id": 2077786,
      "postDate": "2022-12-27T21:52:55.423Z",
      "content": "<p>Normalization: This involves scaling the data to have a mean of 0 and a standard deviation of 1. This can help improve the performance of the CNN by making the data more consistent across different channels.</p>\n<p>You can also normalize your data before your save to to save a bit of time on it as well. <br>\nMaybe also do your data augmentation before your save your data but this will also cause problems in the form of you won't be able have many augmented different versions of the same data.<br>\nYou can also do some filtering (high pass.. etc) prior to saving your data. </p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Normalization: This involves scaling the data to have a mean of 0 and a standard deviation of 1. This can help improve the performance of the CNN by making the data more consistent across different channels.\n\nYou can also normalize your data before your save to to save a bit of time on it as well. \nMaybe also do your data augmentation before your save your data but this will also cause problems in the form of you won't be able have many augmented different versions of the same data.\nYou can also do some filtering (high pass.. etc) prior to saving your data. \n\n\nThe Devastator.\n",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2080536,
      "author_name": "Konstantin Dmitriev",
      "author_url": "",
      "post_date": "2022-12-30T08:58:48.953000",
      "content": "<p>I can point out at least two more ways to preprocess the data.<br>\n1) \"Correlation\". Summing the squares of real and imaginary parts of the data is equal to the product of the signal and its complex conjugation. But we have two signals, and we can take the product of the signal on L1 detector and the complex conjugation of the signal on H1 detector. Then we apply some kind of averaging. The result will be their mutual spectrum, which is a fourier transform of correlation function. As we hope the signals on the two detectors are correlated, this can be used as \"the third channel\" of the image fed to CNN.</p>\n<p>2) \"Zebra approach\". It was proposed during the last year competition and it was reported to lead to good results (though, I didn't try it myself). We can mix the data from two detectors consequently: one time sample from L1, than another time sample from H1 and so on. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2079355,
      "author_name": "PiotrKlinke",
      "author_url": "",
      "post_date": "2022-12-29T09:09:48.103000",
      "content": "<p>Actually, the trimming and combining data from both detectors is already suboptimal. If you look at the timestamps of the two detectors measurements - they are not linear. If I understand correctly, this means the two detectors are not synchronized and combining the data is not trivial. To do it better - we could analyze each timestamp and try to resample them to get constant intervals for each detector separately. Then the combination could work much better, right? Because we would be combining data from exactly the same moment.</p>\n<p>Of course to resample the data consistently we would need to approximate the data that is missing, using a linear approximation or sth - and this could lead to errors.</p>\n<p>Is this something that you would be thinking about?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2079562,
          "author_name": "Aaftaab V",
          "author_url": "",
          "post_date": "2022-12-29T12:23:53.420000",
          "content": "<p>The point you raised is actually good, we are not taking care to synchronize timestamps in real data (I understand, this is mostly present only in test dataset), but this should not be a problem with data generation, since, we (at least I am) are generating both detector's data at a time, and with same duration/timestamps. Maybe this should be taken care of, only while predicting after training model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2077786,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-12-27T21:52:55.423000",
      "content": "<p>Normalization: This involves scaling the data to have a mean of 0 and a standard deviation of 1. This can help improve the performance of the CNN by making the data more consistent across different channels.</p>\n<p>You can also normalize your data before your save to to save a bit of time on it as well. <br>\nMaybe also do your data augmentation before your save your data but this will also cause problems in the form of you won't be able have many augmented different versions of the same data.<br>\nYou can also do some filtering (high pass.. etc) prior to saving your data. </p>\n<p>The Devastator.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2075647": "The most common method I have seen for preprocessing data before feeding it into a CNN, is to square and add real and imaginary parts, multiply each pixel with 1e22, takethe  average or summation of data from both detectors, to get finafinal image, and then resize that to as many channels as needed. I was just wondering if there are any other preprocessing methods in use. Also, how good is the above preprocessing  step wrt domain/physics knowledge? \nAsking because saving sfts from pyfstat is taking tootoo much precious disk space. I wonder if I can preprocess sfts while generating data and save finafinal data as images or npy files. Is anyone doing the above hack?",
    "2080536": "I can point out at least two more ways to preprocess the data.\n1) \"Correlation\". Summing the squares of real and imaginary parts of the data is equal to the product of the signal and its complex conjugation. But we have two signals, and we can take the product of the signal on L1 detector and the complex conjugation of the signal on H1 detector. Then we apply some kind of averaging. The result will be their mutual spectrum, which is a fourier transform of correlation function. As we hope the signals on the two detectors are correlated, this can be used as \"the third channel\" of the image fed to CNN.\n\n2) \"Zebra approach\". It was proposed during the last year competition and it was reported to lead to good results (though, I didn't try it myself). We can mix the data from two detectors consequently: one time sample from L1, than another time sample from H1 and so on. ",
    "2079355": "Actually, the trimming and combining data from both detectors is already suboptimal. If you look at the timestamps of the two detectors measurements - they are not linear. If I understand correctly, this means the two detectors are not synchronized and combining the data is not trivial. To do it better - we could analyze each timestamp and try to resample them to get constant intervals for each detector separately. Then the combination could work much better, right? Because we would be combining data from exactly the same moment.\n\nOf course to resample the data consistently we would need to approximate the data that is missing, using a linear approximation or sth - and this could lead to errors.\n\nIs this something that you would be thinking about?",
    "2077786": "Normalization: This involves scaling the data to have a mean of 0 and a standard deviation of 1. This can help improve the performance of the CNN by making the data more consistent across different channels.\n\nYou can also normalize your data before your save to to save a bit of time on it as well. \nMaybe also do your data augmentation before your save your data but this will also cause problems in the form of you won't be able have many augmented different versions of the same data.\nYou can also do some filtering (high pass.. etc) prior to saving your data. \n\n\nThe Devastator.\n"
  }
}