{
  "id": 369216,
  "title": "Generation Parameters and Distribution problems",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/369216",
  "author_name": "Kerri",
  "post_date": "2022-11-29T10:44:26.578000",
  "votes": 8,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First of all, thanks <a href=\"https://www.kaggle.com/rodrigotenorio\" target=\"_blank\">@rodrigotenorio</a> for such interesting competition! I have a problem with data generation - I understand that parameters is one of the mysteries that we need to reveal, but anyway I have no clue why it works like this. So:<br>\nIf we try to get some input from test data (for example), cut SFTs from approximately 4500 to 1024, than we will get an array with shape (1, 2, 360, 1024). Then do flatten for H1 detector - we will get this histplot:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5452590%2Fd8f484636eafdba807156e3607cf2752%2Fdistrib.jpg?generation=1669717818846626&amp;alt=media\" alt=\"\"><br>\nWith this values:<br>\nFlatten shape: (368640,)<br>\nMax: 5.0543144e-22 <br>\nMin: 2.4336356e-25 <br>\nMean: 1.3294535e-22 <br>\nStd: 6.4837456e-23</p>\n<p>If I try to do it for noise generation with some default values (scrapped on forums, etc), I will get different result. Generation parameters:<br>\nwriter_kwargs = {<br>\n    \"label\": \"single_detector_gaussian_noise\",<br>\n    \"outdir\": \"PyFstat_example_data\",<br>\n    \"tstart\": 1238166018, <br>\n    \"duration\": 93 * 86400, <br>\n    \"detectors\": \"H1,L1\", <br>\n    \"F0\": 100.0,  <br>\n    \"Band\": 0.2,  <br>\n    \"sqrtSX\": 1e-23,  <br>\n    \"Tsft\": 1800,  <br>\n    \"SFTWindowType\": \"tukey\",  <br>\n    \"SFTWindowBeta\": 0.01,  <br>\n}<br>\nFinal shape after flatten for H1 and cutting to 1024 will be the same:<br>\nFlatten shape: (368640,)<br>\nMax: 1.204e-42<br>\nMin: 0.0<br>\nMean: 9e-44<br>\nStd: 0.0<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5452590%2Fb2437756d7217a2673fc5087511e40d6%2Fgendistrib.jpg?generation=1669718494046396&amp;alt=media\" alt=\"\"><br>\nWhat is the problem in my parameters and how to do generated data as mush as possible close to original? Thanks!</p>",
  "messages": [
    {
      "id": 2048308,
      "postDate": "2022-11-29T10:44:26.580Z",
      "content": "<p>First of all, thanks <a href=\"https://www.kaggle.com/rodrigotenorio\" target=\"_blank\">@rodrigotenorio</a> for such interesting competition! I have a problem with data generation - I understand that parameters is one of the mysteries that we need to reveal, but anyway I have no clue why it works like this. So:<br>\nIf we try to get some input from test data (for example), cut SFTs from approximately 4500 to 1024, than we will get an array with shape (1, 2, 360, 1024). Then do flatten for H1 detector - we will get this histplot:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5452590%2Fd8f484636eafdba807156e3607cf2752%2Fdistrib.jpg?generation=1669717818846626&amp;alt=media\" alt=\"\"><br>\nWith this values:<br>\nFlatten shape: (368640,)<br>\nMax: 5.0543144e-22 <br>\nMin: 2.4336356e-25 <br>\nMean: 1.3294535e-22 <br>\nStd: 6.4837456e-23</p>\n<p>If I try to do it for noise generation with some default values (scrapped on forums, etc), I will get different result. Generation parameters:<br>\nwriter_kwargs = {<br>\n    \"label\": \"single_detector_gaussian_noise\",<br>\n    \"outdir\": \"PyFstat_example_data\",<br>\n    \"tstart\": 1238166018, <br>\n    \"duration\": 93 * 86400, <br>\n    \"detectors\": \"H1,L1\", <br>\n    \"F0\": 100.0,  <br>\n    \"Band\": 0.2,  <br>\n    \"sqrtSX\": 1e-23,  <br>\n    \"Tsft\": 1800,  <br>\n    \"SFTWindowType\": \"tukey\",  <br>\n    \"SFTWindowBeta\": 0.01,  <br>\n}<br>\nFinal shape after flatten for H1 and cutting to 1024 will be the same:<br>\nFlatten shape: (368640,)<br>\nMax: 1.204e-42<br>\nMin: 0.0<br>\nMean: 9e-44<br>\nStd: 0.0<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5452590%2Fb2437756d7217a2673fc5087511e40d6%2Fgendistrib.jpg?generation=1669718494046396&amp;alt=media\" alt=\"\"><br>\nWhat is the problem in my parameters and how to do generated data as mush as possible close to original? Thanks!</p>",
      "rawMarkdown": "First of all, thanks @rodrigotenorio for such interesting competition! I have a problem with data generation - I understand that parameters is one of the mysteries that we need to reveal, but anyway I have no clue why it works like this. So:\nIf we try to get some input from test data (for example), cut SFTs from approximately 4500 to 1024, than we will get an array with shape (1, 2, 360, 1024). Then do flatten for H1 detector - we will get this histplot:![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5452590%2Fd8f484636eafdba807156e3607cf2752%2Fdistrib.jpg?generation=1669717818846626&alt=media)\nWith this values:\nFlatten shape: (368640,)\nMax: 5.0543144e-22 \nMin: 2.4336356e-25 \nMean: 1.3294535e-22 \nStd: 6.4837456e-23\n\nIf I try to do it for noise generation with some default values (scrapped on forums, etc), I will get different result. Generation parameters:\nwriter_kwargs = {\n    \"label\": \"single_detector_gaussian_noise\",\n    \"outdir\": \"PyFstat_example_data\",\n    \"tstart\": 1238166018, \n    \"duration\": 93 * 86400, \n    \"detectors\": \"H1,L1\", \n    \"F0\": 100.0,  \n    \"Band\": 0.2,  \n    \"sqrtSX\": 1e-23,  \n    \"Tsft\": 1800,  \n    \"SFTWindowType\": \"tukey\",  \n    \"SFTWindowBeta\": 0.01,  \n}\nFinal shape after flatten for H1 and cutting to 1024 will be the same:\nFlatten shape: (368640,)\nMax: 1.204e-42\nMin: 0.0\nMean: 9e-44\nStd: 0.0\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5452590%2Fb2437756d7217a2673fc5087511e40d6%2Fgendistrib.jpg?generation=1669718494046396&alt=media)\nWhat is the problem in my parameters and how to do generated data as mush as possible close to original? Thanks!",
      "votes": 8
    },
    {
      "id": 2048481,
      "postDate": "2022-11-29T13:26:04.243Z",
      "content": "<p>What's your 'cut' operation refers to, resize or the first 1024 SFTs?</p>\n<p>And seems these figures have different order of magnitudes, so I guess the second figure is the histplot of the amplitude, but I'm not sure what's the first one.</p>\n<p>Also the second figure demonstrates a long tail distribution, which may influence <code>binwidth</code> parameter in the histplot function. Those are lower than 1E-22 in the first figure may be mergerd into the first bar in the second one.</p>",
      "rawMarkdown": "What's your 'cut' operation refers to, resize or the first 1024 SFTs?\n\nAnd seems these figures have different order of magnitudes, so I guess the second figure is the histplot of the amplitude, but I'm not sure what's the first one.\n\nAlso the second figure demonstrates a long tail distribution, which may influence `binwidth` parameter in the histplot function. Those are lower than 1E-22 in the first figure may be mergerd into the first bar in the second one.",
      "votes": 2,
      "replies": [
        {
          "id": 2048694,
          "postDate": "2022-11-29T16:04:21.090Z",
          "content": "<p>\"Cutting\" operation is just for comfort, I am trying to get distribution of noise, so it is okay to take only part of this noise. \"Bins\" parameters for both histograms same - 50. I think this is problem of parameters for generation, but I dont understand what paramater is responsible for this. Anyway I have very strange Mean and STD and strange range of values on X-axis for <br>\nany generated noise.</p>",
          "rawMarkdown": "\"Cutting\" operation is just for comfort, I am trying to get distribution of noise, so it is okay to take only part of this noise. \"Bins\" parameters for both histograms same - 50. I think this is problem of parameters for generation, but I dont understand what paramater is responsible for this. Anyway I have very strange Mean and STD and strange range of values on X-axis for \nany generated noise.",
          "votes": 1
        },
        {
          "id": 2048720,
          "postDate": "2022-11-29T16:24:17.183Z",
          "content": "<p>So what's the <strong>exact value</strong> you are trying to histplot in both figures? Real part/image part/amplitude or squared amplitude?</p>",
          "rawMarkdown": "So what's the **exact value** you are trying to histplot in both figures? Real part/image part/amplitude or squared amplitude?",
          "votes": 1
        },
        {
          "id": 2048757,
          "postDate": "2022-11-29T16:49:46.343Z",
          "content": "<p>Oh, I finally found an error. For test data i have try to take sqrt of amplitude, but for generated - just amplitude, without sqrt. So stupid, actually. Thank you SO MUCH for right questions &lt;3<br>\nAnyway, if it is not a secret - where you takes all this parameters values for generation of noise and signal? Some of them looks very non-obvious, for example H0</p>",
          "rawMarkdown": "Oh, I finally found an error. For test data i have try to take sqrt of amplitude, but for generated - just amplitude, without sqrt. So stupid, actually. Thank you SO MUCH for right questions <3\nAnyway, if it is not a secret - where you takes all this parameters values for generation of noise and signal? Some of them looks very non-obvious, for example H0",
          "votes": 2,
          "replies": [
            {
              "id": 2074276,
              "postDate": "2022-12-24T00:04:11.750Z",
              "content": "<p>thank you very much</p>",
              "rawMarkdown": "thank you very much"
            }
          ]
        },
        {
          "id": 2048770,
          "postDate": "2022-11-29T17:00:09.507Z",
          "content": "<p>You may browse the pinned post by the host and also <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361562\" target=\"_blank\">this post</a> for some insights. Also there are some notebooks for generating data in the <code>Code</code> area.</p>",
          "rawMarkdown": "You may browse the pinned post by the host and also [this post](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361562) for some insights. Also there are some notebooks for generating data in the `Code` area.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2059294,
      "postDate": "2022-12-08T17:20:06.203Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kerrit\" target=\"_blank\">@kerrit</a> ,</p>\n<p>Sorry for the late reply, it looks like <a href=\"https://www.kaggle.com/zzzc18\" target=\"_blank\">@zzzc18</a> already solved your problem, good job!</p>\n<p>I am actually curious about your first plot: What is the continuous line? is it a KDE?</p>\n<p>For reference, if you are working with Gaussian noise, the power of your data should follow a chi2 distribution<br>\nwith 2 degrees of freedom (cause Real and Imaginary parts of the Fourier transform are Gaussian as well<br>\nand you are adding them as squared numbers).</p>",
      "rawMarkdown": "Hi @kerrit ,\n\nSorry for the late reply, it looks like @zzzc18 already solved your problem, good job!\n\nI am actually curious about your first plot: What is the continuous line? is it a KDE?\n\nFor reference, if you are working with Gaussian noise, the power of your data should follow a chi2 distribution\nwith 2 degrees of freedom (cause Real and Imaginary parts of the Fourier transform are Gaussian as well\nand you are adding them as squared numbers)."
    }
  ],
  "comments": [
    {
      "id": 2048481,
      "author_name": "zzzc18",
      "author_url": "",
      "post_date": "2022-11-29T13:26:04.243000",
      "content": "<p>What's your 'cut' operation refers to, resize or the first 1024 SFTs?</p>\n<p>And seems these figures have different order of magnitudes, so I guess the second figure is the histplot of the amplitude, but I'm not sure what's the first one.</p>\n<p>Also the second figure demonstrates a long tail distribution, which may influence <code>binwidth</code> parameter in the histplot function. Those are lower than 1E-22 in the first figure may be mergerd into the first bar in the second one.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2048694,
          "author_name": "Kerri",
          "author_url": "",
          "post_date": "2022-11-29T16:04:21.090000",
          "content": "<p>\"Cutting\" operation is just for comfort, I am trying to get distribution of noise, so it is okay to take only part of this noise. \"Bins\" parameters for both histograms same - 50. I think this is problem of parameters for generation, but I dont understand what paramater is responsible for this. Anyway I have very strange Mean and STD and strange range of values on X-axis for <br>\nany generated noise.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2048720,
          "author_name": "zzzc18",
          "author_url": "",
          "post_date": "2022-11-29T16:24:17.183000",
          "content": "<p>So what's the <strong>exact value</strong> you are trying to histplot in both figures? Real part/image part/amplitude or squared amplitude?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2048757,
          "author_name": "Kerri",
          "author_url": "",
          "post_date": "2022-11-29T16:49:46.343000",
          "content": "<p>Oh, I finally found an error. For test data i have try to take sqrt of amplitude, but for generated - just amplitude, without sqrt. So stupid, actually. Thank you SO MUCH for right questions &lt;3<br>\nAnyway, if it is not a secret - where you takes all this parameters values for generation of noise and signal? Some of them looks very non-obvious, for example H0</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2074276,
              "author_name": "Noussair Mighri",
              "author_url": "",
              "post_date": "2022-12-24T00:04:11.750000",
              "content": "<p>thank you very much</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2048770,
          "author_name": "zzzc18",
          "author_url": "",
          "post_date": "2022-11-29T17:00:09.507000",
          "content": "<p>You may browse the pinned post by the host and also <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361562\" target=\"_blank\">this post</a> for some insights. Also there are some notebooks for generating data in the <code>Code</code> area.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2059294,
      "author_name": "Rodrigo Tenorio",
      "author_url": "",
      "post_date": "2022-12-08T17:20:06.203000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kerrit\" target=\"_blank\">@kerrit</a> ,</p>\n<p>Sorry for the late reply, it looks like <a href=\"https://www.kaggle.com/zzzc18\" target=\"_blank\">@zzzc18</a> already solved your problem, good job!</p>\n<p>I am actually curious about your first plot: What is the continuous line? is it a KDE?</p>\n<p>For reference, if you are working with Gaussian noise, the power of your data should follow a chi2 distribution<br>\nwith 2 degrees of freedom (cause Real and Imaginary parts of the Fourier transform are Gaussian as well<br>\nand you are adding them as squared numbers).</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2048308": "First of all, thanks @rodrigotenorio for such interesting competition! I have a problem with data generation - I understand that parameters is one of the mysteries that we need to reveal, but anyway I have no clue why it works like this. So:\nIf we try to get some input from test data (for example), cut SFTs from approximately 4500 to 1024, than we will get an array with shape (1, 2, 360, 1024). Then do flatten for H1 detector - we will get this histplot:![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5452590%2Fd8f484636eafdba807156e3607cf2752%2Fdistrib.jpg?generation=1669717818846626&alt=media)\nWith this values:\nFlatten shape: (368640,)\nMax: 5.0543144e-22 \nMin: 2.4336356e-25 \nMean: 1.3294535e-22 \nStd: 6.4837456e-23\n\nIf I try to do it for noise generation with some default values (scrapped on forums, etc), I will get different result. Generation parameters:\nwriter_kwargs = {\n    \"label\": \"single_detector_gaussian_noise\",\n    \"outdir\": \"PyFstat_example_data\",\n    \"tstart\": 1238166018, \n    \"duration\": 93 * 86400, \n    \"detectors\": \"H1,L1\", \n    \"F0\": 100.0,  \n    \"Band\": 0.2,  \n    \"sqrtSX\": 1e-23,  \n    \"Tsft\": 1800,  \n    \"SFTWindowType\": \"tukey\",  \n    \"SFTWindowBeta\": 0.01,  \n}\nFinal shape after flatten for H1 and cutting to 1024 will be the same:\nFlatten shape: (368640,)\nMax: 1.204e-42\nMin: 0.0\nMean: 9e-44\nStd: 0.0\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5452590%2Fb2437756d7217a2673fc5087511e40d6%2Fgendistrib.jpg?generation=1669718494046396&alt=media)\nWhat is the problem in my parameters and how to do generated data as mush as possible close to original? Thanks!",
    "2048481": "What's your 'cut' operation refers to, resize or the first 1024 SFTs?\n\nAnd seems these figures have different order of magnitudes, so I guess the second figure is the histplot of the amplitude, but I'm not sure what's the first one.\n\nAlso the second figure demonstrates a long tail distribution, which may influence `binwidth` parameter in the histplot function. Those are lower than 1E-22 in the first figure may be mergerd into the first bar in the second one.",
    "2059294": "Hi @kerrit ,\n\nSorry for the late reply, it looks like @zzzc18 already solved your problem, good job!\n\nI am actually curious about your first plot: What is the continuous line? is it a KDE?\n\nFor reference, if you are working with Gaussian noise, the power of your data should follow a chi2 distribution\nwith 2 degrees of freedom (cause Real and Imaginary parts of the Fourier transform are Gaussian as well\nand you are adding them as squared numbers)."
  }
}