{
  "id": 363810,
  "title": "High validation accuracy (0.77), low LB score (0.56)",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363810",
  "author_name": "Simone Bavera",
  "post_date": "2022-11-03T08:12:20.816000",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>I joined Kaggle over one year ago with the first G2Net challenge and also decided to join the second edition. </p>\n<p>In the past edition, I scored 81% accuracy with EfficientNetB0. So I used the same model as last year to see what it achieves here. On an 80/20 train validation data split, the model scores 77% accuracy, but on the test dataset, I only get 56% accuracy, hence my post. </p>\n<p>Is the training sample not representative of the test dataset?</p>\n<p>Why do we only have 600 signals in the training dataset compared to the ~8000 in the test dataset? Compared to the past challenge, is data augmentation expected in this one? Do we need to come up with our own simulated dataset?</p>\n<p>Curious to hear your thoughts.</p>",
  "messages": [
    {
      "id": 2015499,
      "postDate": "2022-11-03T10:25:06.290Z",
      "content": "<p>I don't think bringing CV and LB closer together in this match was as simple as just implementing all the ml tricks correctly.</p>\n<p>This competition is very different from the last one. The last contest was the signal of two black holes or neutron stars merging, which is very strong, and implementing a CNN would allow for some level of fit. But continuous gravitational waves are very weak, and even after scientists in the field have implemented various search methods and deep learning tools they have not found any real signal. If you read some of the papers you will see that even the most advanced NNs that have been tried have not reached traditional methods of search in terms of accuracy, and are usually very unsatisfactory after a SNR of less than 10.</p>\n<p>I think the huge gap between CV and LB is because 1) the training set is so small with only 600 data, there is so little variation in between that it may be difficult for the network to learn convolution methods in the face of low SNR samples, and 2) I haven't carefully checked the test set, but I think the composition of the test set must be very demanding and there must be a large number of low SNR samples.</p>\n<p>In order to be able to bring CV and LB closer together, I think the most effective way to do this would be to construct a training set that has a large number of samples with a wide range of SNR values, rather than using 600 training samples and then trying various ML tricks.</p>",
      "rawMarkdown": "I don't think bringing CV and LB closer together in this match was as simple as just implementing all the ml tricks correctly.\n\nThis competition is very different from the last one. The last contest was the signal of two black holes or neutron stars merging, which is very strong, and implementing a CNN would allow for some level of fit. But continuous gravitational waves are very weak, and even after scientists in the field have implemented various search methods and deep learning tools they have not found any real signal. If you read some of the papers you will see that even the most advanced NNs that have been tried have not reached traditional methods of search in terms of accuracy, and are usually very unsatisfactory after a SNR of less than 10.\n\nI think the huge gap between CV and LB is because 1) the training set is so small with only 600 data, there is so little variation in between that it may be difficult for the network to learn convolution methods in the face of low SNR samples, and 2) I haven't carefully checked the test set, but I think the composition of the test set must be very demanding and there must be a large number of low SNR samples.\n\nIn order to be able to bring CV and LB closer together, I think the most effective way to do this would be to construct a training set that has a large number of samples with a wide range of SNR values, rather than using 600 training samples and then trying various ML tricks.",
      "votes": 13,
      "replies": [
        {
          "id": 2059938,
          "postDate": "2022-12-09T11:36:15.743Z",
          "content": "<p>I totally agree. Most of the performance increases result from data centric approaches. While I think that the regular used NNs in this competition (EfficientNet, Resnets, …) are sufficiently complex enough to generalize well on this task, providing the right data with the right SNRs as well as chosing a fitting preprocessing pipeline is the challenge. Most model centric approaches like excessive hyperparameter optimization, NN adoption, adapting the learning function etc. could and do sometimes lead to better scores but are not the main driver to achieve outperforming results and to increase the scores significantly.</p>",
          "rawMarkdown": "I totally agree. Most of the performance increases result from data centric approaches. While I think that the regular used NNs in this competition (EfficientNet, Resnets, ...) are sufficiently complex enough to generalize well on this task, providing the right data with the right SNRs as well as chosing a fitting preprocessing pipeline is the challenge. Most model centric approaches like excessive hyperparameter optimization, NN adoption, adapting the learning function etc. could and do sometimes lead to better scores but are not the main driver to achieve outperforming results and to increase the scores significantly.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2015347,
      "postDate": "2022-11-03T08:12:20.817Z",
      "content": "<p>Hello all,</p>\n<p>I joined Kaggle over one year ago with the first G2Net challenge and also decided to join the second edition. </p>\n<p>In the past edition, I scored 81% accuracy with EfficientNetB0. So I used the same model as last year to see what it achieves here. On an 80/20 train validation data split, the model scores 77% accuracy, but on the test dataset, I only get 56% accuracy, hence my post. </p>\n<p>Is the training sample not representative of the test dataset?</p>\n<p>Why do we only have 600 signals in the training dataset compared to the ~8000 in the test dataset? Compared to the past challenge, is data augmentation expected in this one? Do we need to come up with our own simulated dataset?</p>\n<p>Curious to hear your thoughts.</p>",
      "rawMarkdown": "Hello all,\n\nI joined Kaggle over one year ago with the first G2Net challenge and also decided to join the second edition. \n\nIn the past edition, I scored 81% accuracy with EfficientNetB0. So I used the same model as last year to see what it achieves here. On an 80/20 train validation data split, the model scores 77% accuracy, but on the test dataset, I only get 56% accuracy, hence my post. \n\nIs the training sample not representative of the test dataset?\n\nWhy do we only have 600 signals in the training dataset compared to the ~8000 in the test dataset? Compared to the past challenge, is data augmentation expected in this one? Do we need to come up with our own simulated dataset?\n\nCurious to hear your thoughts.",
      "votes": 6
    },
    {
      "id": 2015413,
      "postDate": "2022-11-03T08:57:49.807Z",
      "content": "<blockquote>\n  <p>Is the training sample not representative of the test dataset?</p>\n</blockquote>\n<p>This ad infinitum. You got with one of the keys of this challenge, at least for myself. Definetivelly, the train samples are NOT representative at all. I suggest doing a intensive Test data EDA and take, or generate, data according with the observed samples. In fact, I'm doing that the best I can. But have in mind the Bias too. It's difficult to find a trade off. Because this, there are so many entries in the leaderboard. I hope this can help you.</p>\n<p>Good luck! :)</p>",
      "rawMarkdown": "> Is the training sample not representative of the test dataset?\n\nThis ad infinitum. You got with one of the keys of this challenge, at least for myself. Definetivelly, the train samples are NOT representative at all. I suggest doing a intensive Test data EDA and take, or generate, data according with the observed samples. In fact, I'm doing that the best I can. But have in mind the Bias too. It's difficult to find a trade off. Because this, there are so many entries in the leaderboard. I hope this can help you.\n\nGood luck! :)",
      "votes": 4
    },
    {
      "id": 2020520,
      "postDate": "2022-11-07T15:02:05.863Z",
      "content": "<p>Using the \"train\" dataset as the only source to train the model seems to be a bad idea in this competition. The \"train\" dataset is something ideal, while the \"test\" dataset includes a variety of noises and glitches.</p>",
      "rawMarkdown": "Using the \"train\" dataset as the only source to train the model seems to be a bad idea in this competition. The \"train\" dataset is something ideal, while the \"test\" dataset includes a variety of noises and glitches.",
      "votes": 2
    },
    {
      "id": 2015751,
      "postDate": "2022-11-03T14:20:08.647Z",
      "content": "<ul>\n<li>All the train set is generated noise, not real one. And the sqrtSX of that noise it's basically the same. What matters is the strength relation between the noise and the signal, the so called <em>depth</em> (and the <em>cos i</em> parameter, which greatly affects the strength of the signal over the noise). There are several topics already about this. </li>\n<li>The test set contains both real and generated noise. **The real ones are exactly 1500 samples ** if I am not wrong, and it's nice to just plot them and see what's the real noise (non-stationary) that scientists need to deal with in real CW searches, plus lots of persistent horizontal lines, glitches etc… </li>\n</ul>",
      "rawMarkdown": "- All the train set is generated noise, not real one. And the sqrtSX of that noise it's basically the same. What matters is the strength relation between the noise and the signal, the so called *depth* (and the *cos i* parameter, which greatly affects the strength of the signal over the noise). There are several topics already about this. \n- The test set contains both real and generated noise. **The real ones are exactly 1500 samples ** if I am not wrong, and it's nice to just plot them and see what's the real noise (non-stationary) that scientists need to deal with in real CW searches, plus lots of persistent horizontal lines, glitches etc... \n\n",
      "votes": 2,
      "replies": [
        {
          "id": 2015771,
          "postDate": "2022-11-03T14:35:05.113Z",
          "content": "<p>WOW, how do you distinguish which is the real sample?</p>",
          "rawMarkdown": "WOW, how do you distinguish which is the real sample?"
        },
        {
          "id": 2015896,
          "postDate": "2022-11-03T15:59:16.747Z",
          "content": "<p>The noise is different in the spectrograms for the real ones. You can use the public Jun Koda's notebook \"Basic Spectrogram Image Classifier\" to see those spectrograms if you want :)</p>",
          "rawMarkdown": "The noise is different in the spectrograms for the real ones. You can use the public Jun Koda's notebook \"Basic Spectrogram Image Classifier\" to see those spectrograms if you want :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2015360,
      "postDate": "2022-11-03T08:34:21.733Z",
      "content": "<p>Try train-dev set validation to understand if your model is over-fitting or you have data mismatch problem. In order to do that you can hold out some part of your train set (train-dev set) and evaluate your model on train-dev set. If it performs well that means your model is not over-fitting. If it performs poor on the test set so the problem is most likely data mismatch. In contrast, if your model performs poorly on train-dev set means that you model is over-fitting. So you have to  regularize your model, apply some preprocessing or data augmentation. I hope it helps you :)</p>",
      "rawMarkdown": "Try train-dev set validation to understand if your model is over-fitting or you have data mismatch problem. In order to do that you can hold out some part of your train set (train-dev set) and evaluate your model on train-dev set. If it performs well that means your model is not over-fitting. If it performs poor on the test set so the problem is most likely data mismatch. In contrast, if your model performs poorly on train-dev set means that you model is over-fitting. So you have to  regularize your model, apply some preprocessing or data augmentation. I hope it helps you :)",
      "votes": -4
    }
  ],
  "comments": [
    {
      "id": 2015499,
      "author_name": "Chen Lin",
      "author_url": "",
      "post_date": "2022-11-03T10:25:06.290000",
      "content": "<p>I don't think bringing CV and LB closer together in this match was as simple as just implementing all the ml tricks correctly.</p>\n<p>This competition is very different from the last one. The last contest was the signal of two black holes or neutron stars merging, which is very strong, and implementing a CNN would allow for some level of fit. But continuous gravitational waves are very weak, and even after scientists in the field have implemented various search methods and deep learning tools they have not found any real signal. If you read some of the papers you will see that even the most advanced NNs that have been tried have not reached traditional methods of search in terms of accuracy, and are usually very unsatisfactory after a SNR of less than 10.</p>\n<p>I think the huge gap between CV and LB is because 1) the training set is so small with only 600 data, there is so little variation in between that it may be difficult for the network to learn convolution methods in the face of low SNR samples, and 2) I haven't carefully checked the test set, but I think the composition of the test set must be very demanding and there must be a large number of low SNR samples.</p>\n<p>In order to be able to bring CV and LB closer together, I think the most effective way to do this would be to construct a training set that has a large number of samples with a wide range of SNR values, rather than using 600 training samples and then trying various ML tricks.</p>",
      "votes": 13,
      "replies": [
        {
          "id": 2059938,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2022-12-09T11:36:15.743000",
          "content": "<p>I totally agree. Most of the performance increases result from data centric approaches. While I think that the regular used NNs in this competition (EfficientNet, Resnets, …) are sufficiently complex enough to generalize well on this task, providing the right data with the right SNRs as well as chosing a fitting preprocessing pipeline is the challenge. Most model centric approaches like excessive hyperparameter optimization, NN adoption, adapting the learning function etc. could and do sometimes lead to better scores but are not the main driver to achieve outperforming results and to increase the scores significantly.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2015413,
      "author_name": "Zollkron",
      "author_url": "",
      "post_date": "2022-11-03T08:57:49.807000",
      "content": "<blockquote>\n  <p>Is the training sample not representative of the test dataset?</p>\n</blockquote>\n<p>This ad infinitum. You got with one of the keys of this challenge, at least for myself. Definetivelly, the train samples are NOT representative at all. I suggest doing a intensive Test data EDA and take, or generate, data according with the observed samples. In fact, I'm doing that the best I can. But have in mind the Bias too. It's difficult to find a trade off. Because this, there are so many entries in the leaderboard. I hope this can help you.</p>\n<p>Good luck! :)</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2020520,
      "author_name": "Konstantin Dmitriev",
      "author_url": "",
      "post_date": "2022-11-07T15:02:05.863000",
      "content": "<p>Using the \"train\" dataset as the only source to train the model seems to be a bad idea in this competition. The \"train\" dataset is something ideal, while the \"test\" dataset includes a variety of noises and glitches.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2015751,
      "author_name": "Victor Gonzalez",
      "author_url": "",
      "post_date": "2022-11-03T14:20:08.647000",
      "content": "<ul>\n<li>All the train set is generated noise, not real one. And the sqrtSX of that noise it's basically the same. What matters is the strength relation between the noise and the signal, the so called <em>depth</em> (and the <em>cos i</em> parameter, which greatly affects the strength of the signal over the noise). There are several topics already about this. </li>\n<li>The test set contains both real and generated noise. **The real ones are exactly 1500 samples ** if I am not wrong, and it's nice to just plot them and see what's the real noise (non-stationary) that scientists need to deal with in real CW searches, plus lots of persistent horizontal lines, glitches etc… </li>\n</ul>",
      "votes": 2,
      "replies": [
        {
          "id": 2015771,
          "author_name": "Chen Lin",
          "author_url": "",
          "post_date": "2022-11-03T14:35:05.113000",
          "content": "<p>WOW, how do you distinguish which is the real sample?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2015896,
          "author_name": "Zollkron",
          "author_url": "",
          "post_date": "2022-11-03T15:59:16.747000",
          "content": "<p>The noise is different in the spectrograms for the real ones. You can use the public Jun Koda's notebook \"Basic Spectrogram Image Classifier\" to see those spectrograms if you want :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2015360,
      "author_name": "Ahmet Çalış",
      "author_url": "",
      "post_date": "2022-11-03T08:34:21.733000",
      "content": "<p>Try train-dev set validation to understand if your model is over-fitting or you have data mismatch problem. In order to do that you can hold out some part of your train set (train-dev set) and evaluate your model on train-dev set. If it performs well that means your model is not over-fitting. If it performs poor on the test set so the problem is most likely data mismatch. In contrast, if your model performs poorly on train-dev set means that you model is over-fitting. So you have to  regularize your model, apply some preprocessing or data augmentation. I hope it helps you :)</p>",
      "votes": -4,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2015499": "I don't think bringing CV and LB closer together in this match was as simple as just implementing all the ml tricks correctly.\n\nThis competition is very different from the last one. The last contest was the signal of two black holes or neutron stars merging, which is very strong, and implementing a CNN would allow for some level of fit. But continuous gravitational waves are very weak, and even after scientists in the field have implemented various search methods and deep learning tools they have not found any real signal. If you read some of the papers you will see that even the most advanced NNs that have been tried have not reached traditional methods of search in terms of accuracy, and are usually very unsatisfactory after a SNR of less than 10.\n\nI think the huge gap between CV and LB is because 1) the training set is so small with only 600 data, there is so little variation in between that it may be difficult for the network to learn convolution methods in the face of low SNR samples, and 2) I haven't carefully checked the test set, but I think the composition of the test set must be very demanding and there must be a large number of low SNR samples.\n\nIn order to be able to bring CV and LB closer together, I think the most effective way to do this would be to construct a training set that has a large number of samples with a wide range of SNR values, rather than using 600 training samples and then trying various ML tricks.",
    "2015347": "Hello all,\n\nI joined Kaggle over one year ago with the first G2Net challenge and also decided to join the second edition. \n\nIn the past edition, I scored 81% accuracy with EfficientNetB0. So I used the same model as last year to see what it achieves here. On an 80/20 train validation data split, the model scores 77% accuracy, but on the test dataset, I only get 56% accuracy, hence my post. \n\nIs the training sample not representative of the test dataset?\n\nWhy do we only have 600 signals in the training dataset compared to the ~8000 in the test dataset? Compared to the past challenge, is data augmentation expected in this one? Do we need to come up with our own simulated dataset?\n\nCurious to hear your thoughts.",
    "2015413": "> Is the training sample not representative of the test dataset?\n\nThis ad infinitum. You got with one of the keys of this challenge, at least for myself. Definetivelly, the train samples are NOT representative at all. I suggest doing a intensive Test data EDA and take, or generate, data according with the observed samples. In fact, I'm doing that the best I can. But have in mind the Bias too. It's difficult to find a trade off. Because this, there are so many entries in the leaderboard. I hope this can help you.\n\nGood luck! :)",
    "2020520": "Using the \"train\" dataset as the only source to train the model seems to be a bad idea in this competition. The \"train\" dataset is something ideal, while the \"test\" dataset includes a variety of noises and glitches.",
    "2015751": "- All the train set is generated noise, not real one. And the sqrtSX of that noise it's basically the same. What matters is the strength relation between the noise and the signal, the so called *depth* (and the *cos i* parameter, which greatly affects the strength of the signal over the noise). There are several topics already about this. \n- The test set contains both real and generated noise. **The real ones are exactly 1500 samples ** if I am not wrong, and it's nice to just plot them and see what's the real noise (non-stationary) that scientists need to deal with in real CW searches, plus lots of persistent horizontal lines, glitches etc... \n\n",
    "2015360": "Try train-dev set validation to understand if your model is over-fitting or you have data mismatch problem. In order to do that you can hold out some part of your train set (train-dev set) and evaluate your model on train-dev set. If it performs well that means your model is not over-fitting. If it performs poor on the test set so the problem is most likely data mismatch. In contrast, if your model performs poorly on train-dev set means that you model is over-fitting. So you have to  regularize your model, apply some preprocessing or data augmentation. I hope it helps you :)"
  }
}