{
  "id": 374612,
  "title": "Sampling rate and FFT_size",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/374612",
  "author_name": "Naren Manikandan",
  "post_date": "2022-12-28T04:19:11.135000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello everyone! I'm trying to convert the spectrogram into MFCC coefficients and I need to input both the sampling rate and FFT_size of the generated data (parameters shown below). </p>\n<pre><code>writer_kwargs = {\n    \"label\": \"single_detector_gaussian_noise\",\n    \"outdir\": \"PyFstat_example_data\",\n    \"tstart\": 1238166018,\n    \"duration\": 365 * 86400,\n    \"detectors\": \"H1\",\n    \"sqrtSX\": 1e-23,\n    \"Tsft\": 1800,\n    \"SFTWindowType\": \"tukey\",\n    \"SFTWindowBeta\": 0.01,\n}\n</code></pre>\n<p>How could I get the total sample size? Would it be the combined shape of the fourier_data that has been created? Also, the FFT_size would be 1238166018 (duration) / 1800 (the time held by each ts)? Thanks in advance!</p>",
  "messages": [
    {
      "id": 2079343,
      "postDate": "2022-12-29T09:00:14.557Z",
      "content": "<p>I was thinking of such approach and also couldn't get the right parameters.<br>\nFrom what I understand:</p>\n<ul>\n<li>fft size is the length of a single fft frame - so the value should be the number of freq bins of your spectrogram. I didnt generate my own data from the pyfsat so I am not sure what should be the value here. I would look  at the shape of the data the MFCC analysis will consume</li>\n<li>the sampling frequency is also sth I wondered how to approach. This is what I have seen on the given training data: Even though we are presented with the data as spectrograms, the frequency vector attached to the data is for example between 250 and 256 Hz - sooo I would argue this is not a full spectrum, just an excerpt of the full spectrum that was measured… So there is no base sampling rate to work with?</li>\n</ul>\n<p>That being said, the transformation of the spectrum to MFCC is a composition of 3 steps as per <a href=\"https://en.m.wikipedia.org/wiki/Mel-frequency_cepstrum\" target=\"_blank\">https://en.m.wikipedia.org/wiki/Mel-frequency_cepstrum</a><br>\nAnd actually you only need the sampling rate parameter to determine the frequencies to apply the nonlinear mel scale - in the first step of the transformation.</p>\n<p>The approach I am thinking about is to ommit the first step and just take dct(log(sft)) as my features. This operation could possibly have most of the characteristics so valued in MFCC.</p>\n<p>The other approach is to use any sampling rate and try different ones because actually this sound is not human voice anyway so the exact mel mel frequencies really dont have to work the same way. It could also be a trainable parameter.</p>\n<p>Hope it helps. Generally it is not easy to use audio operations on this dataset and I am happy others are trying these approaches as well! Best of luck!</p>",
      "rawMarkdown": "I was thinking of such approach and also couldn't get the right parameters.\nFrom what I understand:\n- fft size is the length of a single fft frame - so the value should be the number of freq bins of your spectrogram. I didnt generate my own data from the pyfsat so I am not sure what should be the value here. I would look  at the shape of the data the MFCC analysis will consume\n- the sampling frequency is also sth I wondered how to approach. This is what I have seen on the given training data: Even though we are presented with the data as spectrograms, the frequency vector attached to the data is for example between 250 and 256 Hz - sooo I would argue this is not a full spectrum, just an excerpt of the full spectrum that was measured... So there is no base sampling rate to work with?\n\nThat being said, the transformation of the spectrum to MFCC is a composition of 3 steps as per https://en.m.wikipedia.org/wiki/Mel-frequency_cepstrum\nAnd actually you only need the sampling rate parameter to determine the frequencies to apply the nonlinear mel scale - in the first step of the transformation.\n\nThe approach I am thinking about is to ommit the first step and just take dct(log(sft)) as my features. This operation could possibly have most of the characteristics so valued in MFCC.\n\nThe other approach is to use any sampling rate and try different ones because actually this sound is not human voice anyway so the exact mel mel frequencies really dont have to work the same way. It could also be a trainable parameter.\n\nHope it helps. Generally it is not easy to use audio operations on this dataset and I am happy others are trying these approaches as well! Best of luck!",
      "votes": 1,
      "replies": [
        {
          "id": 2080278,
          "postDate": "2022-12-30T03:34:09.640Z",
          "content": "<p>Thanks for replying! As you said, I got the fft_size from the second value of the shape of the SFT's.<br>\nIf we are trying to implement the entire algorithm, I saw that the Nyquist theorem can give use the sampling rate from the max frequency or vice versa:</p>\n<p>Fmax = SR/2</p>\n<p>Although the FFT algorithm can give us the max frequency, I'm unsure about the units of these frequencies since usually, the sampling rate is 44100 hz, making the max frequency 22050 hz, but this is way higher than our frequencies. It looks like they are scaled down between a range of values. My max frequency is 100.043 and min is just over 99, making the range very small. If this function does not produce any good results, I will just skip it and implement the other ones. Do you think this is the right approach?</p>",
          "rawMarkdown": "Thanks for replying! As you said, I got the fft_size from the second value of the shape of the SFT's.\nIf we are trying to implement the entire algorithm, I saw that the Nyquist theorem can give use the sampling rate from the max frequency or vice versa:\n\nFmax = SR/2\n\nAlthough the FFT algorithm can give us the max frequency, I'm unsure about the units of these frequencies since usually, the sampling rate is 44100 hz, making the max frequency 22050 hz, but this is way higher than our frequencies. It looks like they are scaled down between a range of values. My max frequency is 100.043 and min is just over 99, making the range very small. If this function does not produce any good results, I will just skip it and implement the other ones. Do you think this is the right approach?",
          "votes": 1,
          "replies": [
            {
              "id": 2080632,
              "postDate": "2022-12-30T10:27:14.527Z",
              "content": "<p>Hey no problem! You are right with the Nyquist theorem but here as we are given only some of the frequencies  and not the full range (the.frequency vector doesn't start with 0 - the full fourier analysis always starts with 0) I am not sure using 2xFmax is definetely the best approach. I would try this as first guess and then try other values.</p>\n<p>Also don't worry about having the not standard frequencies - 44.1kHz is a standard just for music signals.</p>",
              "rawMarkdown": "Hey no problem! You are right with the Nyquist theorem but here as we are given only some of the frequencies  and not the full range (the.frequency vector doesn't start with 0 - the full fourier analysis always starts with 0) I am not sure using 2xFmax is definetely the best approach. I would try this as first guess and then try other values.\n\nAlso don't worry about having the not standard frequencies - 44.1kHz is a standard just for music signals."
            }
          ]
        }
      ]
    },
    {
      "id": 2078064,
      "postDate": "2022-12-28T04:19:11.137Z",
      "content": "<p>Hello everyone! I'm trying to convert the spectrogram into MFCC coefficients and I need to input both the sampling rate and FFT_size of the generated data (parameters shown below). </p>\n<pre><code>writer_kwargs = {\n    \"label\": \"single_detector_gaussian_noise\",\n    \"outdir\": \"PyFstat_example_data\",\n    \"tstart\": 1238166018,\n    \"duration\": 365 * 86400,\n    \"detectors\": \"H1\",\n    \"sqrtSX\": 1e-23,\n    \"Tsft\": 1800,\n    \"SFTWindowType\": \"tukey\",\n    \"SFTWindowBeta\": 0.01,\n}\n</code></pre>\n<p>How could I get the total sample size? Would it be the combined shape of the fourier_data that has been created? Also, the FFT_size would be 1238166018 (duration) / 1800 (the time held by each ts)? Thanks in advance!</p>",
      "rawMarkdown": "Hello everyone! I'm trying to convert the spectrogram into MFCC coefficients and I need to input both the sampling rate and FFT_size of the generated data (parameters shown below). \n```\nwriter_kwargs = {\n    \"label\": \"single_detector_gaussian_noise\",\n    \"outdir\": \"PyFstat_example_data\",\n    \"tstart\": 1238166018,\n    \"duration\": 365 * 86400,\n    \"detectors\": \"H1\",\n    \"sqrtSX\": 1e-23,\n    \"Tsft\": 1800,\n    \"SFTWindowType\": \"tukey\",\n    \"SFTWindowBeta\": 0.01,\n}\n```\nHow could I get the total sample size? Would it be the combined shape of the fourier_data that has been created? Also, the FFT_size would be 1238166018 (duration) / 1800 (the time held by each ts)? Thanks in advance!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2079343,
      "author_name": "PiotrKlinke",
      "author_url": "",
      "post_date": "2022-12-29T09:00:14.557000",
      "content": "<p>I was thinking of such approach and also couldn't get the right parameters.<br>\nFrom what I understand:</p>\n<ul>\n<li>fft size is the length of a single fft frame - so the value should be the number of freq bins of your spectrogram. I didnt generate my own data from the pyfsat so I am not sure what should be the value here. I would look  at the shape of the data the MFCC analysis will consume</li>\n<li>the sampling frequency is also sth I wondered how to approach. This is what I have seen on the given training data: Even though we are presented with the data as spectrograms, the frequency vector attached to the data is for example between 250 and 256 Hz - sooo I would argue this is not a full spectrum, just an excerpt of the full spectrum that was measured… So there is no base sampling rate to work with?</li>\n</ul>\n<p>That being said, the transformation of the spectrum to MFCC is a composition of 3 steps as per <a href=\"https://en.m.wikipedia.org/wiki/Mel-frequency_cepstrum\" target=\"_blank\">https://en.m.wikipedia.org/wiki/Mel-frequency_cepstrum</a><br>\nAnd actually you only need the sampling rate parameter to determine the frequencies to apply the nonlinear mel scale - in the first step of the transformation.</p>\n<p>The approach I am thinking about is to ommit the first step and just take dct(log(sft)) as my features. This operation could possibly have most of the characteristics so valued in MFCC.</p>\n<p>The other approach is to use any sampling rate and try different ones because actually this sound is not human voice anyway so the exact mel mel frequencies really dont have to work the same way. It could also be a trainable parameter.</p>\n<p>Hope it helps. Generally it is not easy to use audio operations on this dataset and I am happy others are trying these approaches as well! Best of luck!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2080278,
          "author_name": "Naren Manikandan",
          "author_url": "",
          "post_date": "2022-12-30T03:34:09.640000",
          "content": "<p>Thanks for replying! As you said, I got the fft_size from the second value of the shape of the SFT's.<br>\nIf we are trying to implement the entire algorithm, I saw that the Nyquist theorem can give use the sampling rate from the max frequency or vice versa:</p>\n<p>Fmax = SR/2</p>\n<p>Although the FFT algorithm can give us the max frequency, I'm unsure about the units of these frequencies since usually, the sampling rate is 44100 hz, making the max frequency 22050 hz, but this is way higher than our frequencies. It looks like they are scaled down between a range of values. My max frequency is 100.043 and min is just over 99, making the range very small. If this function does not produce any good results, I will just skip it and implement the other ones. Do you think this is the right approach?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2080632,
              "author_name": "PiotrKlinke",
              "author_url": "",
              "post_date": "2022-12-30T10:27:14.527000",
              "content": "<p>Hey no problem! You are right with the Nyquist theorem but here as we are given only some of the frequencies  and not the full range (the.frequency vector doesn't start with 0 - the full fourier analysis always starts with 0) I am not sure using 2xFmax is definetely the best approach. I would try this as first guess and then try other values.</p>\n<p>Also don't worry about having the not standard frequencies - 44.1kHz is a standard just for music signals.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2079343": "I was thinking of such approach and also couldn't get the right parameters.\nFrom what I understand:\n- fft size is the length of a single fft frame - so the value should be the number of freq bins of your spectrogram. I didnt generate my own data from the pyfsat so I am not sure what should be the value here. I would look  at the shape of the data the MFCC analysis will consume\n- the sampling frequency is also sth I wondered how to approach. This is what I have seen on the given training data: Even though we are presented with the data as spectrograms, the frequency vector attached to the data is for example between 250 and 256 Hz - sooo I would argue this is not a full spectrum, just an excerpt of the full spectrum that was measured... So there is no base sampling rate to work with?\n\nThat being said, the transformation of the spectrum to MFCC is a composition of 3 steps as per https://en.m.wikipedia.org/wiki/Mel-frequency_cepstrum\nAnd actually you only need the sampling rate parameter to determine the frequencies to apply the nonlinear mel scale - in the first step of the transformation.\n\nThe approach I am thinking about is to ommit the first step and just take dct(log(sft)) as my features. This operation could possibly have most of the characteristics so valued in MFCC.\n\nThe other approach is to use any sampling rate and try different ones because actually this sound is not human voice anyway so the exact mel mel frequencies really dont have to work the same way. It could also be a trainable parameter.\n\nHope it helps. Generally it is not easy to use audio operations on this dataset and I am happy others are trying these approaches as well! Best of luck!",
    "2078064": "Hello everyone! I'm trying to convert the spectrogram into MFCC coefficients and I need to input both the sampling rate and FFT_size of the generated data (parameters shown below). \n```\nwriter_kwargs = {\n    \"label\": \"single_detector_gaussian_noise\",\n    \"outdir\": \"PyFstat_example_data\",\n    \"tstart\": 1238166018,\n    \"duration\": 365 * 86400,\n    \"detectors\": \"H1\",\n    \"sqrtSX\": 1e-23,\n    \"Tsft\": 1800,\n    \"SFTWindowType\": \"tukey\",\n    \"SFTWindowBeta\": 0.01,\n}\n```\nHow could I get the total sample size? Would it be the combined shape of the fourier_data that has been created? Also, the FFT_size would be 1238166018 (duration) / 1800 (the time held by each ts)? Thanks in advance!"
  }
}