{
  "id": 370158,
  "title": "Odd Test Samples - How to classify them?",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/370158",
  "author_name": "Mark Wijkhuizen",
  "post_date": "2022-12-03T11:45:06.268000",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>There seem to be test samples which differ from the provided train set and do not resemble the conventional continious gravitational wave spectograms.</p>\n<p>An example is given below for test_id <code>a1f9b8e82</code>. The samples are sliced [:,:4096] and then resized to 360x256 with mean pooling.</p>\n<p>For the Livington example, what is going on here and which label should be assigned to these samples?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F589d86d96b19ff6515ec6649997ad1e5%2Fa1f9b8e82_l1_360x256.png?generation=1670067786701946&amp;alt=media\" alt=\"\"><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F1009c86c75b7db207ebc6b7f2370075c%2Fa1f9b8e82_H1_360x256.png?generation=1670067791108332&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2060043,
      "postDate": "2022-12-09T14:07:15.047Z",
      "content": "<p>Its hard to work with data discrepancy especially when the test set is fixed. An idea could be to use pseudo labeling to generate a big enough valid set (from the test set) and to check how well your model + preprocessing is working on it. After model selection, hyperparameter tuning as well as preprocessing pipeline selection you can train your model on all the available train + newly generated valid set and try to see how well it performs on the available test set (public LB).</p>\n<p>Other than that only good data augmentation, collecting new representative data (e.g. generate signals with differing SNRs) and thought-out preprocessing can diminish the data discrepancy problem a little. </p>",
      "rawMarkdown": "Its hard to work with data discrepancy especially when the test set is fixed. An idea could be to use pseudo labeling to generate a big enough valid set (from the test set) and to check how well your model + preprocessing is working on it. After model selection, hyperparameter tuning as well as preprocessing pipeline selection you can train your model on all the available train + newly generated valid set and try to see how well it performs on the available test set (public LB).\n\nOther than that only good data augmentation, collecting new representative data (e.g. generate signals with differing SNRs) and thought-out preprocessing can diminish the data discrepancy problem a little. ",
      "votes": 1
    },
    {
      "id": 2053552,
      "postDate": "2022-12-03T11:45:06.270Z",
      "content": "<p>There seem to be test samples which differ from the provided train set and do not resemble the conventional continious gravitational wave spectograms.</p>\n<p>An example is given below for test_id <code>a1f9b8e82</code>. The samples are sliced [:,:4096] and then resized to 360x256 with mean pooling.</p>\n<p>For the Livington example, what is going on here and which label should be assigned to these samples?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F589d86d96b19ff6515ec6649997ad1e5%2Fa1f9b8e82_l1_360x256.png?generation=1670067786701946&amp;alt=media\" alt=\"\"><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F1009c86c75b7db207ebc6b7f2370075c%2Fa1f9b8e82_H1_360x256.png?generation=1670067791108332&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "There seem to be test samples which differ from the provided train set and do not resemble the conventional continious gravitational wave spectograms.\n\nAn example is given below for test_id `a1f9b8e82`. The samples are sliced [:,:4096] and then resized to 360x256 with mean pooling.\n\nFor the Livington example, what is going on here and which label should be assigned to these samples?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F589d86d96b19ff6515ec6649997ad1e5%2Fa1f9b8e82_l1_360x256.png?generation=1670067786701946&alt=media)![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F1009c86c75b7db207ebc6b7f2370075c%2Fa1f9b8e82_H1_360x256.png?generation=1670067791108332&alt=media)",
      "votes": 2
    },
    {
      "id": 2054677,
      "postDate": "2022-12-04T10:39:32.153Z",
      "content": "<p>Test data differs a lot from train data…</p>\n<p>As you can see, we have the glitch in L1 data. The interference is so strong, that the signal (if it exists) can't be seen. I think, we can do nothing with it but cut off the corresponding interval on time axis. But in this particular case L1 data is so bad, that my algorithm decided to use only H1 data.</p>",
      "rawMarkdown": "Test data differs a lot from train data...\n\nAs you can see, we have the glitch in L1 data. The interference is so strong, that the signal (if it exists) can't be seen. I think, we can do nothing with it but cut off the corresponding interval on time axis. But in this particular case L1 data is so bad, that my algorithm decided to use only H1 data."
    }
  ],
  "comments": [
    {
      "id": 2060043,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2022-12-09T14:07:15.047000",
      "content": "<p>Its hard to work with data discrepancy especially when the test set is fixed. An idea could be to use pseudo labeling to generate a big enough valid set (from the test set) and to check how well your model + preprocessing is working on it. After model selection, hyperparameter tuning as well as preprocessing pipeline selection you can train your model on all the available train + newly generated valid set and try to see how well it performs on the available test set (public LB).</p>\n<p>Other than that only good data augmentation, collecting new representative data (e.g. generate signals with differing SNRs) and thought-out preprocessing can diminish the data discrepancy problem a little. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2054677,
      "author_name": "Konstantin Dmitriev",
      "author_url": "",
      "post_date": "2022-12-04T10:39:32.153000",
      "content": "<p>Test data differs a lot from train data…</p>\n<p>As you can see, we have the glitch in L1 data. The interference is so strong, that the signal (if it exists) can't be seen. I think, we can do nothing with it but cut off the corresponding interval on time axis. But in this particular case L1 data is so bad, that my algorithm decided to use only H1 data.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2060043": "Its hard to work with data discrepancy especially when the test set is fixed. An idea could be to use pseudo labeling to generate a big enough valid set (from the test set) and to check how well your model + preprocessing is working on it. After model selection, hyperparameter tuning as well as preprocessing pipeline selection you can train your model on all the available train + newly generated valid set and try to see how well it performs on the available test set (public LB).\n\nOther than that only good data augmentation, collecting new representative data (e.g. generate signals with differing SNRs) and thought-out preprocessing can diminish the data discrepancy problem a little. ",
    "2053552": "There seem to be test samples which differ from the provided train set and do not resemble the conventional continious gravitational wave spectograms.\n\nAn example is given below for test_id `a1f9b8e82`. The samples are sliced [:,:4096] and then resized to 360x256 with mean pooling.\n\nFor the Livington example, what is going on here and which label should be assigned to these samples?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F589d86d96b19ff6515ec6649997ad1e5%2Fa1f9b8e82_l1_360x256.png?generation=1670067786701946&alt=media)![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F1009c86c75b7db207ebc6b7f2370075c%2Fa1f9b8e82_H1_360x256.png?generation=1670067791108332&alt=media)",
    "2054677": "Test data differs a lot from train data...\n\nAs you can see, we have the glitch in L1 data. The interference is so strong, that the signal (if it exists) can't be seen. I think, we can do nothing with it but cut off the corresponding interval on time axis. But in this particular case L1 data is so bad, that my algorithm decided to use only H1 data."
  }
}