{
  "id": 368387,
  "title": "Is is leagal? Leak of background noise?",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/368387",
  "author_name": "Tanrei(nama)",
  "post_date": "2022-11-25T01:51:25.890000",
  "votes": 0,
  "comment_count": 3,
  "views": 0,
  "content": "<p>This corresponds to \"Magic#2\" that appeared in the SETI competition a year ago. At that time, this solution was considered valid.</p>\n<p><a href=\"https://www.kaggle.com/code/tanreinama/eliminate-noise-using-signal-similarity\" target=\"_blank\">https://www.kaggle.com/code/tanreinama/eliminate-noise-using-signal-similarity</a></p>",
  "messages": [
    {
      "id": 2044618,
      "postDate": "2022-11-26T16:58:58.857Z",
      "content": "<p>As we don't know the labels of test dataset, actually it's nearly impossible to do such operation. If you try to take advantage of this 'leakage' then you need to foucus on LB scores. However, public LB only contains the result of 24% of the full data, which may easily leed to an overfit.</p>",
      "rawMarkdown": "As we don't know the labels of test dataset, actually it's nearly impossible to do such operation. If you try to take advantage of this 'leakage' then you need to foucus on LB scores. However, public LB only contains the result of 24% of the full data, which may easily leed to an overfit.",
      "votes": 1,
      "replies": [
        {
          "id": 2045671,
          "postDate": "2022-11-27T15:01:04.583Z",
          "content": "<p>I'm agreed. However, it's a good indicator for evaluation/validation for the true model.</p>",
          "rawMarkdown": "I'm agreed. However, it's a good indicator for evaluation/validation for the true model."
        }
      ]
    },
    {
      "id": 2043300,
      "postDate": "2022-11-25T14:38:07.550Z",
      "content": "<p>In another comment (that I can't find right now), a Kaggle mod mentioned that they should be able to completely rename the test set, and your code should still work - so anything that relies on manually set test ids would be against the rules.</p>\n<p>I believe that means that as it stands, this notebook wouldn't be valid - but if you were able to automatically detect and leverage it, then that would be ok (my interpretation only).</p>",
      "rawMarkdown": "In another comment (that I can't find right now), a Kaggle mod mentioned that they should be able to completely rename the test set, and your code should still work - so anything that relies on manually set test ids would be against the rules.\n\nI believe that means that as it stands, this notebook wouldn't be valid - but if you were able to automatically detect and leverage it, then that would be ok (my interpretation only).",
      "votes": 2
    },
    {
      "id": 2042765,
      "postDate": "2022-11-25T01:51:25.890Z",
      "content": "<p>This corresponds to \"Magic#2\" that appeared in the SETI competition a year ago. At that time, this solution was considered valid.</p>\n<p><a href=\"https://www.kaggle.com/code/tanreinama/eliminate-noise-using-signal-similarity\" target=\"_blank\">https://www.kaggle.com/code/tanreinama/eliminate-noise-using-signal-similarity</a></p>",
      "rawMarkdown": "This corresponds to \"Magic#2\" that appeared in the SETI competition a year ago. At that time, this solution was considered valid.\n\nhttps://www.kaggle.com/code/tanreinama/eliminate-noise-using-signal-similarity",
      "votes": -1
    }
  ],
  "comments": [
    {
      "id": 2044618,
      "author_name": "zzzc18",
      "author_url": "",
      "post_date": "2022-11-26T16:58:58.857000",
      "content": "<p>As we don't know the labels of test dataset, actually it's nearly impossible to do such operation. If you try to take advantage of this 'leakage' then you need to foucus on LB scores. However, public LB only contains the result of 24% of the full data, which may easily leed to an overfit.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2045671,
          "author_name": "Zollkron",
          "author_url": "",
          "post_date": "2022-11-27T15:01:04.583000",
          "content": "<p>I'm agreed. However, it's a good indicator for evaluation/validation for the true model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2043300,
      "author_name": "chris",
      "author_url": "",
      "post_date": "2022-11-25T14:38:07.550000",
      "content": "<p>In another comment (that I can't find right now), a Kaggle mod mentioned that they should be able to completely rename the test set, and your code should still work - so anything that relies on manually set test ids would be against the rules.</p>\n<p>I believe that means that as it stands, this notebook wouldn't be valid - but if you were able to automatically detect and leverage it, then that would be ok (my interpretation only).</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2044618": "As we don't know the labels of test dataset, actually it's nearly impossible to do such operation. If you try to take advantage of this 'leakage' then you need to foucus on LB scores. However, public LB only contains the result of 24% of the full data, which may easily leed to an overfit.",
    "2043300": "In another comment (that I can't find right now), a Kaggle mod mentioned that they should be able to completely rename the test set, and your code should still work - so anything that relies on manually set test ids would be against the rules.\n\nI believe that means that as it stands, this notebook wouldn't be valid - but if you were able to automatically detect and leverage it, then that would be ok (my interpretation only).",
    "2042765": "This corresponds to \"Magic#2\" that appeared in the SETI competition a year ago. At that time, this solution was considered valid.\n\nhttps://www.kaggle.com/code/tanreinama/eliminate-noise-using-signal-similarity"
  }
}