{
  "id": 364169,
  "title": "Gap of CV and LB",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/364169",
  "author_name": "Chenglu",
  "post_date": "2022-11-05T02:28:50.726000",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<p>My gap of CV and LB is huge in this challenge. I'm using 3 folds and CV is ~0.77, LB ~0.62 with only training dataset. Share your gap here if you like.</p>\n<p>Update:</p>\n<p>with the generated dataset, I can reduce the gap to 0.1, CV ~0.77 and LB ~0.68, I think this competition is a matter of shrinking the gap, more precisely, it's generating a dataset that matches the distribution of test set.</p>",
  "messages": [
    {
      "id": 2017631,
      "postDate": "2022-11-05T02:28:50.727Z",
      "content": "<p>My gap of CV and LB is huge in this challenge. I'm using 3 folds and CV is ~0.77, LB ~0.62 with only training dataset. Share your gap here if you like.</p>\n<p>Update:</p>\n<p>with the generated dataset, I can reduce the gap to 0.1, CV ~0.77 and LB ~0.68, I think this competition is a matter of shrinking the gap, more precisely, it's generating a dataset that matches the distribution of test set.</p>",
      "rawMarkdown": "My gap of CV and LB is huge in this challenge. I'm using 3 folds and CV is ~0.77, LB ~0.62 with only training dataset. Share your gap here if you like.\n\nUpdate:\n\nwith the generated dataset, I can reduce the gap to 0.1, CV ~0.77 and LB ~0.68, I think this competition is a matter of shrinking the gap, more precisely, it's generating a dataset that matches the distribution of test set.",
      "votes": 5
    },
    {
      "id": 2055600,
      "postDate": "2022-12-05T08:09:16.123Z",
      "content": "<p>Guys, the notion of CV score is somewhat useless unless you all use the same validation set. It is meaningful if you are using just the train set provided by Kaggle, but if you've generated additional data and use it as a part of your validation set, then your scores are not comparable to each other.<br>\nIf you want to bring your validation score closer to your LB score just increase the depth of the signal (sqrtsx/d) :)</p>",
      "rawMarkdown": "Guys, the notion of CV score is somewhat useless unless you all use the same validation set. It is meaningful if you are using just the train set provided by Kaggle, but if you've generated additional data and use it as a part of your validation set, then your scores are not comparable to each other.\nIf you want to bring your validation score closer to your LB score just increase the depth of the signal (sqrtsx/d) :)",
      "votes": 2,
      "replies": [
        {
          "id": 2059320,
          "postDate": "2022-12-08T17:51:46.340Z",
          "content": "<blockquote>\n  <p>If you want to bring your validation score closer to your LB score just increase the depth of the signal (sqrtsx/d) :)</p>\n</blockquote>\n<p>I'm gonna steal this as a motivational quote, if you don't mind 😉</p>",
          "rawMarkdown": ">  If you want to bring your validation score closer to your LB score just increase the depth of the signal (sqrtsx/d) :)\n\nI'm gonna steal this as a motivational quote, if you don't mind 😉"
        }
      ]
    },
    {
      "id": 2061057,
      "postDate": "2022-12-10T16:54:41.043Z",
      "content": "<p>Train Data only: CV 0.80 | LB 0.67</p>",
      "rawMarkdown": "Train Data only: CV 0.80 | LB 0.67"
    },
    {
      "id": 2055566,
      "postDate": "2022-12-05T07:06:34.423Z",
      "content": "<p>I am the one who suffers from this gap too.<br>\nMy score is as follows:</p>\n<ol>\n<li>given data only - Efficientnet b4 : cv - 0.789, LB - 0.667</li>\n<li>Generated data (signals 4k, noise 4k) - Efficientnet b4: cv 0.77, LB - 0.65</li>\n</ol>\n<p>You seem to have successfully improved your LB by generating data. Can you explain any hints about how to improve LB with generated data (ex. How much data did you create ?, Stationary or nonstationary noise ?)</p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "I am the one who suffers from this gap too.\nMy score is as follows:\n1. given data only - Efficientnet b4 : cv - 0.789, LB - 0.667\n2. Generated data (signals 4k, noise 4k) - Efficientnet b4: cv 0.77, LB - 0.65\n\nYou seem to have successfully improved your LB by generating data. Can you explain any hints about how to improve LB with generated data (ex. How much data did you create ?, Stationary or nonstationary noise ?)\n\nThanks in advance!",
      "replies": [
        {
          "id": 2055581,
          "postDate": "2022-12-05T07:34:53.430Z",
          "content": "<p>No need to generate too much of external data. 1K for train and validation is enough I think.</p>\n<p>The gap can be reduced by preprocessing: normalization and data augmentation. Try different normalization methods and ensemble will further reduce the gap.</p>",
          "rawMarkdown": "No need to generate too much of external data. 1K for train and validation is enough I think.\n\nThe gap can be reduced by preprocessing: normalization and data augmentation. Try different normalization methods and ensemble will further reduce the gap.",
          "votes": 3
        },
        {
          "id": 2059919,
          "postDate": "2022-12-09T11:16:48.127Z",
          "content": "<p><a href=\"https://www.kaggle.com/yutoshibata\" target=\"_blank\">@yutoshibata</a> Have you checked for data discrepancy? You can just slice out a tiny subset of data from your train set (called train dev set), train a model on your train set check it on the train dev set to detect overfitting. If its performing well on train dev set but not on valid set then there is a data discrepancy problem. If its performing bad on train dev set then the model overfitted.</p>\n<p>Furthermore I think that the models used for training (in your case EfficientNet-B4) are to complex and powerful to detect a in comparision rather \"easy\" signal (e.g. in <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/370202\" target=\"_blank\">https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/370202</a> the author used a self developed tiny model with few parameters and operations and it performed very well) Powerful models have a huge problem with very noisy data, because they tend to overfit fast by learning noisy patterns (also regarding the tiny provided train set) I would suggest you to try \"weaker\" models (e.g. B0, B1) or use more regularization (weight-decay, dropout, label-smoothing, …)</p>",
          "rawMarkdown": "@yutoshibata Have you checked for data discrepancy? You can just slice out a tiny subset of data from your train set (called train dev set), train a model on your train set check it on the train dev set to detect overfitting. If its performing well on train dev set but not on valid set then there is a data discrepancy problem. If its performing bad on train dev set then the model overfitted.\n\nFurthermore I think that the models used for training (in your case EfficientNet-B4) are to complex and powerful to detect a in comparision rather \"easy\" signal (e.g. in https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/370202 the author used a self developed tiny model with few parameters and operations and it performed very well) Powerful models have a huge problem with very noisy data, because they tend to overfit fast by learning noisy patterns (also regarding the tiny provided train set) I would suggest you to try \"weaker\" models (e.g. B0, B1) or use more regularization (weight-decay, dropout, label-smoothing, ...)",
          "votes": 2
        },
        {
          "id": 2060648,
          "postDate": "2022-12-10T08:09:09.560Z",
          "content": "<p><a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a>  Thanks to your kind advice, I changed the normalization method to reduce the gap.<br>\nFinally, my score with generated data on LB became 0.71! Thank you so much.</p>\n<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> I will explore lighter models or custom models as you mentioned. Thanks </p>",
          "rawMarkdown": "@snaker  Thanks to your kind advice, I changed the normalization method to reduce the gap.\nFinally, my score with generated data on LB became 0.71! Thank you so much.\n\n@aliabdin1 I will explore lighter models or custom models as you mentioned. Thanks ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2017737,
      "postDate": "2022-11-05T05:58:50.467Z",
      "content": "<p>There's an interesting conversation about that gap here:<br>\n<a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363810\" target=\"_blank\">\n\"High validation accuracy (0.77), low LB score (0.56)\"</a></p>\n<p>Other people report that the training set is quite different to the test set - different noise characteristics etc.</p>",
      "rawMarkdown": "There's an interesting conversation about that gap here:\n[\n\"High validation accuracy (0.77), low LB score (0.56)\"](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363810)\n\nOther people report that the training set is quite different to the test set - different noise characteristics etc.",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2055600,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2022-12-05T08:09:16.123000",
      "content": "<p>Guys, the notion of CV score is somewhat useless unless you all use the same validation set. It is meaningful if you are using just the train set provided by Kaggle, but if you've generated additional data and use it as a part of your validation set, then your scores are not comparable to each other.<br>\nIf you want to bring your validation score closer to your LB score just increase the depth of the signal (sqrtsx/d) :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2059320,
          "author_name": "Rodrigo Tenorio",
          "author_url": "",
          "post_date": "2022-12-08T17:51:46.340000",
          "content": "<blockquote>\n  <p>If you want to bring your validation score closer to your LB score just increase the depth of the signal (sqrtsx/d) :)</p>\n</blockquote>\n<p>I'm gonna steal this as a motivational quote, if you don't mind 😉</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2061057,
      "author_name": "Yerram Varun",
      "author_url": "",
      "post_date": "2022-12-10T16:54:41.043000",
      "content": "<p>Train Data only: CV 0.80 | LB 0.67</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2055566,
      "author_name": "Shibata",
      "author_url": "",
      "post_date": "2022-12-05T07:06:34.423000",
      "content": "<p>I am the one who suffers from this gap too.<br>\nMy score is as follows:</p>\n<ol>\n<li>given data only - Efficientnet b4 : cv - 0.789, LB - 0.667</li>\n<li>Generated data (signals 4k, noise 4k) - Efficientnet b4: cv 0.77, LB - 0.65</li>\n</ol>\n<p>You seem to have successfully improved your LB by generating data. Can you explain any hints about how to improve LB with generated data (ex. How much data did you create ?, Stationary or nonstationary noise ?)</p>\n<p>Thanks in advance!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2055581,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2022-12-05T07:34:53.430000",
          "content": "<p>No need to generate too much of external data. 1K for train and validation is enough I think.</p>\n<p>The gap can be reduced by preprocessing: normalization and data augmentation. Try different normalization methods and ensemble will further reduce the gap.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2059919,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2022-12-09T11:16:48.127000",
          "content": "<p><a href=\"https://www.kaggle.com/yutoshibata\" target=\"_blank\">@yutoshibata</a> Have you checked for data discrepancy? You can just slice out a tiny subset of data from your train set (called train dev set), train a model on your train set check it on the train dev set to detect overfitting. If its performing well on train dev set but not on valid set then there is a data discrepancy problem. If its performing bad on train dev set then the model overfitted.</p>\n<p>Furthermore I think that the models used for training (in your case EfficientNet-B4) are to complex and powerful to detect a in comparision rather \"easy\" signal (e.g. in <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/370202\" target=\"_blank\">https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/370202</a> the author used a self developed tiny model with few parameters and operations and it performed very well) Powerful models have a huge problem with very noisy data, because they tend to overfit fast by learning noisy patterns (also regarding the tiny provided train set) I would suggest you to try \"weaker\" models (e.g. B0, B1) or use more regularization (weight-decay, dropout, label-smoothing, …)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2060648,
          "author_name": "Shibata",
          "author_url": "",
          "post_date": "2022-12-10T08:09:09.560000",
          "content": "<p><a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a>  Thanks to your kind advice, I changed the normalization method to reduce the gap.<br>\nFinally, my score with generated data on LB became 0.71! Thank you so much.</p>\n<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> I will explore lighter models or custom models as you mentioned. Thanks </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2017737,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-05T05:58:50.467000",
      "content": "<p>There's an interesting conversation about that gap here:<br>\n<a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363810\" target=\"_blank\">\n\"High validation accuracy (0.77), low LB score (0.56)\"</a></p>\n<p>Other people report that the training set is quite different to the test set - different noise characteristics etc.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2017631": "My gap of CV and LB is huge in this challenge. I'm using 3 folds and CV is ~0.77, LB ~0.62 with only training dataset. Share your gap here if you like.\n\nUpdate:\n\nwith the generated dataset, I can reduce the gap to 0.1, CV ~0.77 and LB ~0.68, I think this competition is a matter of shrinking the gap, more precisely, it's generating a dataset that matches the distribution of test set.",
    "2055600": "Guys, the notion of CV score is somewhat useless unless you all use the same validation set. It is meaningful if you are using just the train set provided by Kaggle, but if you've generated additional data and use it as a part of your validation set, then your scores are not comparable to each other.\nIf you want to bring your validation score closer to your LB score just increase the depth of the signal (sqrtsx/d) :)",
    "2061057": "Train Data only: CV 0.80 | LB 0.67",
    "2055566": "I am the one who suffers from this gap too.\nMy score is as follows:\n1. given data only - Efficientnet b4 : cv - 0.789, LB - 0.667\n2. Generated data (signals 4k, noise 4k) - Efficientnet b4: cv 0.77, LB - 0.65\n\nYou seem to have successfully improved your LB by generating data. Can you explain any hints about how to improve LB with generated data (ex. How much data did you create ?, Stationary or nonstationary noise ?)\n\nThanks in advance!",
    "2017737": "There's an interesting conversation about that gap here:\n[\n\"High validation accuracy (0.77), low LB score (0.56)\"](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/363810)\n\nOther people report that the training set is quite different to the test set - different noise characteristics etc."
  }
}