{
  "id": 361501,
  "title": "Does using test data for network training violate the rules?",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/361501",
  "author_name": "wenji",
  "post_date": "2022-10-22T02:56:15.734000",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I found different types of noise in the test data. Can I use some test data for training?<br>\nThis is my first time to participate in the kaggle competition. I don't know whether this behavior violates the rules.</p>",
  "messages": [
    {
      "id": 1998992,
      "postDate": "2022-10-22T02:56:15.733Z",
      "content": "<p>I found different types of noise in the test data. Can I use some test data for training?<br>\nThis is my first time to participate in the kaggle competition. I don't know whether this behavior violates the rules.</p>",
      "rawMarkdown": "I found different types of noise in the test data. Can I use some test data for training?\nThis is my first time to participate in the kaggle competition. I don't know whether this behavior violates the rules.",
      "votes": 7
    },
    {
      "id": 1999086,
      "postDate": "2022-10-22T05:19:17.293Z",
      "content": "<p>Why not? Pseudolabeling (which often uses test data) is a wildly used technique on Kaggle.</p>",
      "rawMarkdown": "Why not? Pseudolabeling (which often uses test data) is a wildly used technique on Kaggle.",
      "votes": 1
    },
    {
      "id": 1999573,
      "postDate": "2022-10-22T13:37:51.020Z",
      "content": "<p>The quick answer is no, but, in my humildy opinion, you need to take bias in mind when you're training your model. The public score is based on only 24% of the test data, so it is highly recommended to train a model with a more generic solution to avoid potential bias in the model. In other words, the key is to find a balance between precision and bias so that the model performs well with the other 76% of test data that remains as well. That's the reason I'm avoiding to train with any test data, but this is my personal consideration only. Good luck to everyone :)</p>",
      "rawMarkdown": "The quick answer is no, but, in my humildy opinion, you need to take bias in mind when you're training your model. The public score is based on only 24% of the test data, so it is highly recommended to train a model with a more generic solution to avoid potential bias in the model. In other words, the key is to find a balance between precision and bias so that the model performs well with the other 76% of test data that remains as well. That's the reason I'm avoiding to train with any test data, but this is my personal consideration only. Good luck to everyone :)",
      "votes": 2,
      "replies": [
        {
          "id": 1999578,
          "postDate": "2022-10-22T13:51:39.540Z",
          "content": "<p>Indeed, training with test may lead to validation-overfitting leading to high score in the 24% of testdata and low on the other 76%. So in such cases, the best approaches are semi-supervised - unsupervised techniques, in order to exploit the large unlabeled test pool.</p>",
          "rawMarkdown": "Indeed, training with test may lead to validation-overfitting leading to high score in the 24% of testdata and low on the other 76%. So in such cases, the best approaches are semi-supervised - unsupervised techniques, in order to exploit the large unlabeled test pool.",
          "votes": 1
        },
        {
          "id": 1999582,
          "postDate": "2022-10-22T13:56:19.727Z",
          "content": "<p>You are right and I'm totally agree with you :)</p>",
          "rawMarkdown": "You are right and I'm totally agree with you :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2000290,
      "postDate": "2022-10-23T07:26:06.677Z",
      "content": "<p>It's not against the rules. You won't have the labels of course, but for example you can do EDA on the test data, and infer statistics like the average noise amplitude (sqrtSX), visually inspect outliers (glitches, etc)…</p>",
      "rawMarkdown": "It's not against the rules. You won't have the labels of course, but for example you can do EDA on the test data, and infer statistics like the average noise amplitude (sqrtSX), visually inspect outliers (glitches, etc)...\n"
    },
    {
      "id": 1999348,
      "postDate": "2022-10-22T09:08:38.547Z",
      "content": "<p>Not at all. There are tons of papers who emphasize on semi-supervised learning. Using unlabeled data, is on the rules too :)</p>",
      "rawMarkdown": "Not at all. There are tons of papers who emphasize on semi-supervised learning. Using unlabeled data, is on the rules too :)"
    }
  ],
  "comments": [
    {
      "id": 1999086,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2022-10-22T05:19:17.293000",
      "content": "<p>Why not? Pseudolabeling (which often uses test data) is a wildly used technique on Kaggle.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1999573,
      "author_name": "Zollkron",
      "author_url": "",
      "post_date": "2022-10-22T13:37:51.020000",
      "content": "<p>The quick answer is no, but, in my humildy opinion, you need to take bias in mind when you're training your model. The public score is based on only 24% of the test data, so it is highly recommended to train a model with a more generic solution to avoid potential bias in the model. In other words, the key is to find a balance between precision and bias so that the model performs well with the other 76% of test data that remains as well. That's the reason I'm avoiding to train with any test data, but this is my personal consideration only. Good luck to everyone :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1999578,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-10-22T13:51:39.540000",
          "content": "<p>Indeed, training with test may lead to validation-overfitting leading to high score in the 24% of testdata and low on the other 76%. So in such cases, the best approaches are semi-supervised - unsupervised techniques, in order to exploit the large unlabeled test pool.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1999582,
          "author_name": "Zollkron",
          "author_url": "",
          "post_date": "2022-10-22T13:56:19.727000",
          "content": "<p>You are right and I'm totally agree with you :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2000290,
      "author_name": "Victor Gonzalez",
      "author_url": "",
      "post_date": "2022-10-23T07:26:06.677000",
      "content": "<p>It's not against the rules. You won't have the labels of course, but for example you can do EDA on the test data, and infer statistics like the average noise amplitude (sqrtSX), visually inspect outliers (glitches, etc)…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1999348,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-10-22T09:08:38.547000",
      "content": "<p>Not at all. There are tons of papers who emphasize on semi-supervised learning. Using unlabeled data, is on the rules too :)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1998992": "I found different types of noise in the test data. Can I use some test data for training?\nThis is my first time to participate in the kaggle competition. I don't know whether this behavior violates the rules.",
    "1999086": "Why not? Pseudolabeling (which often uses test data) is a wildly used technique on Kaggle.",
    "1999573": "The quick answer is no, but, in my humildy opinion, you need to take bias in mind when you're training your model. The public score is based on only 24% of the test data, so it is highly recommended to train a model with a more generic solution to avoid potential bias in the model. In other words, the key is to find a balance between precision and bias so that the model performs well with the other 76% of test data that remains as well. That's the reason I'm avoiding to train with any test data, but this is my personal consideration only. Good luck to everyone :)",
    "2000290": "It's not against the rules. You won't have the labels of course, but for example you can do EDA on the test data, and infer statistics like the average noise amplitude (sqrtSX), visually inspect outliers (glitches, etc)...\n",
    "1999348": "Not at all. There are tons of papers who emphasize on semi-supervised learning. Using unlabeled data, is on the rules too :)"
  }
}