{
  "id": 368589,
  "title": "Some insights about validation strategy",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/368589",
  "author_name": "Martin Kovacevic Buvinic",
  "post_date": "2022-11-26T12:18:30.616000",
  "votes": 27,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I wanted to share my initial results on this comp:</p>\n<p>Test data and train data are different, cv gap between validation and test is big (0.10) and correlation is not great:<br>\n Better validation cv != better test cv</p>\n<p>I did the following experiments:</p>\n<p>Using only competition train data</p>\n<ul>\n<li>5 Folds effb7 cv 0.8354, lb 0.708</li>\n<li>10 Folds effb7 cv 0.8470, lb 0.710</li>\n<li>20 Folds effb7 cv 0.8839, lb 0.725 -&gt; This is a very nice cv, nevertheless lb did not increase as much</li>\n</ul>\n<p>With this results I wonder why 20 folds cv is so good, my best guess is that using 20 folds is only 5% of the data as validation for each fold, meaning we are only validating with 30 observations and this lead to overfitting, therefore this validation is not good, increasing the number of folds will just lead to overfitting.</p>\n<p>Nevertheless we could ask ourself why the lb is better with 20 folds, my best guess is because we are using the average of 20 predictions, blending in most of the casses helps.</p>\n<p>Using own generated data, validation on competition data:</p>\n<ul>\n<li>5 Folds effb7 cv 0.8448, lb 0.727</li>\n<li>20 Folds effb7 cv 0.8729, lb 0.734</li>\n<li>1 Fold effb7, cv 0.8233, lb 0.729</li>\n</ul>\n<p>Adding more data usually improves the generalization of the model, if we check lb it actually improves so adding more data is helpful. The data I add try to simulate the training data, my best guess is that the key of the competition is to generate data that is similar to the test data (my next plan). Also check 1 fold experiment, cv is worst but lb is good (we are only using 1 model compared to 20 or 5 that we used on previous experiments)</p>\n<p>My insight is that using training data as validation is not optimal, because it is small (overfitting, as you can see in 20 fold experiments) and it is different compared to the test set. A good option would be to generate our own data that is similar to the test set and do traininig and validation with that data, also add the training data, why not.</p>\n<p>If someone has any observations please share them, maybe I am wrong and it would be great to learn from my mistakes.</p>",
  "messages": [
    {
      "id": 2044268,
      "postDate": "2022-11-26T12:18:30.617Z",
      "content": "<p>I wanted to share my initial results on this comp:</p>\n<p>Test data and train data are different, cv gap between validation and test is big (0.10) and correlation is not great:<br>\n Better validation cv != better test cv</p>\n<p>I did the following experiments:</p>\n<p>Using only competition train data</p>\n<ul>\n<li>5 Folds effb7 cv 0.8354, lb 0.708</li>\n<li>10 Folds effb7 cv 0.8470, lb 0.710</li>\n<li>20 Folds effb7 cv 0.8839, lb 0.725 -&gt; This is a very nice cv, nevertheless lb did not increase as much</li>\n</ul>\n<p>With this results I wonder why 20 folds cv is so good, my best guess is that using 20 folds is only 5% of the data as validation for each fold, meaning we are only validating with 30 observations and this lead to overfitting, therefore this validation is not good, increasing the number of folds will just lead to overfitting.</p>\n<p>Nevertheless we could ask ourself why the lb is better with 20 folds, my best guess is because we are using the average of 20 predictions, blending in most of the casses helps.</p>\n<p>Using own generated data, validation on competition data:</p>\n<ul>\n<li>5 Folds effb7 cv 0.8448, lb 0.727</li>\n<li>20 Folds effb7 cv 0.8729, lb 0.734</li>\n<li>1 Fold effb7, cv 0.8233, lb 0.729</li>\n</ul>\n<p>Adding more data usually improves the generalization of the model, if we check lb it actually improves so adding more data is helpful. The data I add try to simulate the training data, my best guess is that the key of the competition is to generate data that is similar to the test data (my next plan). Also check 1 fold experiment, cv is worst but lb is good (we are only using 1 model compared to 20 or 5 that we used on previous experiments)</p>\n<p>My insight is that using training data as validation is not optimal, because it is small (overfitting, as you can see in 20 fold experiments) and it is different compared to the test set. A good option would be to generate our own data that is similar to the test set and do traininig and validation with that data, also add the training data, why not.</p>\n<p>If someone has any observations please share them, maybe I am wrong and it would be great to learn from my mistakes.</p>",
      "rawMarkdown": "I wanted to share my initial results on this comp:\n\nTest data and train data are different, cv gap between validation and test is big (0.10) and correlation is not great:\n Better validation cv != better test cv\n\nI did the following experiments:\n\nUsing only competition train data\n- 5 Folds effb7 cv 0.8354, lb 0.708\n- 10 Folds effb7 cv 0.8470, lb 0.710\n- 20 Folds effb7 cv 0.8839, lb 0.725 -> This is a very nice cv, nevertheless lb did not increase as much\n\nWith this results I wonder why 20 folds cv is so good, my best guess is that using 20 folds is only 5% of the data as validation for each fold, meaning we are only validating with 30 observations and this lead to overfitting, therefore this validation is not good, increasing the number of folds will just lead to overfitting.\n\nNevertheless we could ask ourself why the lb is better with 20 folds, my best guess is because we are using the average of 20 predictions, blending in most of the casses helps.\n\nUsing own generated data, validation on competition data:\n- 5 Folds effb7 cv 0.8448, lb 0.727\n- 20 Folds effb7 cv 0.8729, lb 0.734\n- 1 Fold effb7, cv 0.8233, lb 0.729\n\nAdding more data usually improves the generalization of the model, if we check lb it actually improves so adding more data is helpful. The data I add try to simulate the training data, my best guess is that the key of the competition is to generate data that is similar to the test data (my next plan). Also check 1 fold experiment, cv is worst but lb is good (we are only using 1 model compared to 20 or 5 that we used on previous experiments)\n\nMy insight is that using training data as validation is not optimal, because it is small (overfitting, as you can see in 20 fold experiments) and it is different compared to the test set. A good option would be to generate our own data that is similar to the test set and do traininig and validation with that data, also add the training data, why not.\n\nIf someone has any observations please share them, maybe I am wrong and it would be great to learn from my mistakes.",
      "votes": 27
    },
    {
      "id": 2045066,
      "postDate": "2022-11-27T03:23:43.663Z",
      "content": "<p>A solid generating dataset plays a very big role in this game. So I think it's always better to use the competition dataset as local validation and generating dataset as train set, thus we can have a solid metric to tell how good is this generating dataset.</p>\n<p>BTW using generated data, and validation on competition data, my local score is 0.82</p>",
      "rawMarkdown": "A solid generating dataset plays a very big role in this game. So I think it's always better to use the competition dataset as local validation and generating dataset as train set, thus we can have a solid metric to tell how good is this generating dataset.\n\nBTW using generated data, and validation on competition data, my local score is 0.82\n",
      "votes": 4,
      "replies": [
        {
          "id": 2045069,
          "postDate": "2022-11-27T03:36:04.117Z",
          "content": "<p>I think the distribution of competition training data and test data seems to be very different. If competition dataset is used as a validation set, it does not seem to evaluate the performance on competition test data well.</p>",
          "rawMarkdown": "I think the distribution of competition training data and test data seems to be very different. If competition dataset is used as a validation set, it does not seem to evaluate the performance on competition test data well.",
          "votes": 2
        },
        {
          "id": 2045072,
          "postDate": "2022-11-27T03:39:08.023Z",
          "content": "<p>I also get a validation roc auc of 0.82 training with my generated dataset and validate with the comp train dataset. I also think train is different from test.</p>\n<p>Update: I get roc auc of 0.83-0.84 validating on the training data, lb improves but the correlation is not great, therefore validating with the training dataset will not align to the test set (at least my experiments suggest that)</p>",
          "rawMarkdown": "I also get a validation roc auc of 0.82 training with my generated dataset and validate with the comp train dataset. I also think train is different from test.\n\nUpdate: I get roc auc of 0.83-0.84 validating on the training data, lb improves but the correlation is not great, therefore validating with the training dataset will not align to the test set (at least my experiments suggest that)",
          "votes": 3
        }
      ]
    },
    {
      "id": 2044463,
      "postDate": "2022-11-26T14:50:35.263Z",
      "content": "<p>In my large number of experiments, cv using official data has not much connection with lb. Using the generated data, my cv is 0.9+, which makes me very confused.</p>",
      "rawMarkdown": "In my large number of experiments, cv using official data has not much connection with lb. Using the generated data, my cv is 0.9+, which makes me very confused.",
      "votes": 1,
      "replies": [
        {
          "id": 2044490,
          "postDate": "2022-11-26T15:10:40.833Z",
          "content": "<p>In my case using generated data + train data to train and evaluate my cv is 0.81-0.82 therefore we are generating data in a different way. I think in my case I manage to generate data exactly as the host did, I compute the average of each spectrogram and compare the distributions and they where very similar. </p>\n<p>One important aspect is h0 parameter, which is the parameter that generates the signal. h0 = sqrtSX / x.</p>\n<p>The bigger x is, the signal will be burried in a deeper depth, therefore if you pick x = 10 it will be much easier for the model to find the signal compared to 100 (at least this is what I understood from the tutorials).</p>\n<p>Maybe your x is very low, and that is the reason why you have an auc of 0.9+.</p>",
          "rawMarkdown": "In my case using generated data + train data to train and evaluate my cv is 0.81-0.82 therefore we are generating data in a different way. I think in my case I manage to generate data exactly as the host did, I compute the average of each spectrogram and compare the distributions and they where very similar. \n\nOne important aspect is h0 parameter, which is the parameter that generates the signal. h0 = sqrtSX / x.\n\nThe bigger x is, the signal will be burried in a deeper depth, therefore if you pick x = 10 it will be much easier for the model to find the signal compared to 100 (at least this is what I understood from the tutorials).\n\nMaybe your x is very low, and that is the reason why you have an auc of 0.9+.",
          "votes": 1
        },
        {
          "id": 2044571,
          "postDate": "2022-11-26T16:17:25.880Z",
          "content": "<p><a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> Thank you. I noticed this. In fact, I have adopted a strategy similar to you, probably not because of x, because I didn't use very low x.</p>",
          "rawMarkdown": "@ragnar123 Thank you. I noticed this. In fact, I have adopted a strategy similar to you, probably not because of x, because I didn't use very low x."
        },
        {
          "id": 2046075,
          "postDate": "2022-11-27T22:26:33.437Z",
          "content": "<p>1st of all, the generated data must have exactly the same representation (compression, normalisation, etc) with official train/test.<br>\nThen is more about good mixing of the fake gravity signal with noise. <br>\nOnly then is very useful. <br>\nHowever instead of generated data, heavy augmentations like flipping, noise injection and so on, applied in the official train can do almost the same job.👌👍</p>",
          "rawMarkdown": "1st of all, the generated data must have exactly the same representation (compression, normalisation, etc) with official train/test.\nThen is more about good mixing of the fake gravity signal with noise. \nOnly then is very useful. \nHowever instead of generated data, heavy augmentations like flipping, noise injection and so on, applied in the official train can do almost the same job.👌👍",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2045066,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2022-11-27T03:23:43.663000",
      "content": "<p>A solid generating dataset plays a very big role in this game. So I think it's always better to use the competition dataset as local validation and generating dataset as train set, thus we can have a solid metric to tell how good is this generating dataset.</p>\n<p>BTW using generated data, and validation on competition data, my local score is 0.82</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2045069,
          "author_name": "zh",
          "author_url": "",
          "post_date": "2022-11-27T03:36:04.117000",
          "content": "<p>I think the distribution of competition training data and test data seems to be very different. If competition dataset is used as a validation set, it does not seem to evaluate the performance on competition test data well.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2045072,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-11-27T03:39:08.023000",
          "content": "<p>I also get a validation roc auc of 0.82 training with my generated dataset and validate with the comp train dataset. I also think train is different from test.</p>\n<p>Update: I get roc auc of 0.83-0.84 validating on the training data, lb improves but the correlation is not great, therefore validating with the training dataset will not align to the test set (at least my experiments suggest that)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2044463,
      "author_name": "zh",
      "author_url": "",
      "post_date": "2022-11-26T14:50:35.263000",
      "content": "<p>In my large number of experiments, cv using official data has not much connection with lb. Using the generated data, my cv is 0.9+, which makes me very confused.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2044490,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-11-26T15:10:40.833000",
          "content": "<p>In my case using generated data + train data to train and evaluate my cv is 0.81-0.82 therefore we are generating data in a different way. I think in my case I manage to generate data exactly as the host did, I compute the average of each spectrogram and compare the distributions and they where very similar. </p>\n<p>One important aspect is h0 parameter, which is the parameter that generates the signal. h0 = sqrtSX / x.</p>\n<p>The bigger x is, the signal will be burried in a deeper depth, therefore if you pick x = 10 it will be much easier for the model to find the signal compared to 100 (at least this is what I understood from the tutorials).</p>\n<p>Maybe your x is very low, and that is the reason why you have an auc of 0.9+.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2044571,
          "author_name": "zh",
          "author_url": "",
          "post_date": "2022-11-26T16:17:25.880000",
          "content": "<p><a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> Thank you. I noticed this. In fact, I have adopted a strategy similar to you, probably not because of x, because I didn't use very low x.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2046075,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-11-27T22:26:33.437000",
          "content": "<p>1st of all, the generated data must have exactly the same representation (compression, normalisation, etc) with official train/test.<br>\nThen is more about good mixing of the fake gravity signal with noise. <br>\nOnly then is very useful. <br>\nHowever instead of generated data, heavy augmentations like flipping, noise injection and so on, applied in the official train can do almost the same job.👌👍</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2044268": "I wanted to share my initial results on this comp:\n\nTest data and train data are different, cv gap between validation and test is big (0.10) and correlation is not great:\n Better validation cv != better test cv\n\nI did the following experiments:\n\nUsing only competition train data\n- 5 Folds effb7 cv 0.8354, lb 0.708\n- 10 Folds effb7 cv 0.8470, lb 0.710\n- 20 Folds effb7 cv 0.8839, lb 0.725 -> This is a very nice cv, nevertheless lb did not increase as much\n\nWith this results I wonder why 20 folds cv is so good, my best guess is that using 20 folds is only 5% of the data as validation for each fold, meaning we are only validating with 30 observations and this lead to overfitting, therefore this validation is not good, increasing the number of folds will just lead to overfitting.\n\nNevertheless we could ask ourself why the lb is better with 20 folds, my best guess is because we are using the average of 20 predictions, blending in most of the casses helps.\n\nUsing own generated data, validation on competition data:\n- 5 Folds effb7 cv 0.8448, lb 0.727\n- 20 Folds effb7 cv 0.8729, lb 0.734\n- 1 Fold effb7, cv 0.8233, lb 0.729\n\nAdding more data usually improves the generalization of the model, if we check lb it actually improves so adding more data is helpful. The data I add try to simulate the training data, my best guess is that the key of the competition is to generate data that is similar to the test data (my next plan). Also check 1 fold experiment, cv is worst but lb is good (we are only using 1 model compared to 20 or 5 that we used on previous experiments)\n\nMy insight is that using training data as validation is not optimal, because it is small (overfitting, as you can see in 20 fold experiments) and it is different compared to the test set. A good option would be to generate our own data that is similar to the test set and do traininig and validation with that data, also add the training data, why not.\n\nIf someone has any observations please share them, maybe I am wrong and it would be great to learn from my mistakes.",
    "2045066": "A solid generating dataset plays a very big role in this game. So I think it's always better to use the competition dataset as local validation and generating dataset as train set, thus we can have a solid metric to tell how good is this generating dataset.\n\nBTW using generated data, and validation on competition data, my local score is 0.82\n",
    "2044463": "In my large number of experiments, cv using official data has not much connection with lb. Using the generated data, my cv is 0.9+, which makes me very confused."
  }
}