{
  "id": 365208,
  "title": "Generating samples that match the train / test distribution",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/365208",
  "author_name": "Chenglu",
  "post_date": "2022-11-10T09:10:25.046000",
  "votes": 11,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I have built a classifier that tries to split the 3 of the dataset: training, testing and generating.</p>\n<p>The classifier to split testing and training can achieve ~0.66 auc score. But for testing and generating, or training and generating, the auc goes to 1.0 .</p>\n<p>I've done some work trying to mess the classifier but none of them work. I'm thinking maybe it is the key to achieve golden zone.</p>\n<hr>\n<p>Update:</p>\n<p>I have got the cause for this. I normalized the data with this method</p>\n<pre><code>norm_data = data - data.min() / (data.max() - data.min())\n</code></pre>\n<p>It will be easily influenced by the value of <code>max</code> and <code>min</code>. Turns out that the generating data of mine has very different <code>max</code> or <code>min</code> value than the training and testing dataset.</p>\n<p>After I changed it to</p>\n<pre><code>norm_data = data * 1e22\n</code></pre>\n<p>The classifier can only achieve 0.66 on classifying generating v.s. testing/training dataset.</p>",
  "messages": [
    {
      "id": 2024110,
      "postDate": "2022-11-10T09:10:25.047Z",
      "content": "<p>I have built a classifier that tries to split the 3 of the dataset: training, testing and generating.</p>\n<p>The classifier to split testing and training can achieve ~0.66 auc score. But for testing and generating, or training and generating, the auc goes to 1.0 .</p>\n<p>I've done some work trying to mess the classifier but none of them work. I'm thinking maybe it is the key to achieve golden zone.</p>\n<hr>\n<p>Update:</p>\n<p>I have got the cause for this. I normalized the data with this method</p>\n<pre><code>norm_data = data - data.min() / (data.max() - data.min())\n</code></pre>\n<p>It will be easily influenced by the value of <code>max</code> and <code>min</code>. Turns out that the generating data of mine has very different <code>max</code> or <code>min</code> value than the training and testing dataset.</p>\n<p>After I changed it to</p>\n<pre><code>norm_data = data * 1e22\n</code></pre>\n<p>The classifier can only achieve 0.66 on classifying generating v.s. testing/training dataset.</p>",
      "rawMarkdown": "I have built a classifier that tries to split the 3 of the dataset: training, testing and generating.\n\nThe classifier to split testing and training can achieve ~0.66 auc score. But for testing and generating, or training and generating, the auc goes to 1.0 .\n\nI've done some work trying to mess the classifier but none of them work. I'm thinking maybe it is the key to achieve golden zone.\n\n---------------------------------------------------------------------------------------------------------------------\nUpdate:\n\n\nI have got the cause for this. I normalized the data with this method\n```\nnorm_data = data - data.min() / (data.max() - data.min())\n```\n\nIt will be easily influenced by the value of `max` and `min`. Turns out that the generating data of mine has very different `max` or `min` value than the training and testing dataset.\n\nAfter I changed it to\n\n```\nnorm_data = data * 1e22\n```\n\nThe classifier can only achieve 0.66 on classifying generating v.s. testing/training dataset.",
      "votes": 11
    },
    {
      "id": 2030026,
      "postDate": "2022-11-15T06:31:29.643Z",
      "content": "<p>My binary classification model can't distinguish official data and my generated data. But my generated data still can't help to improve LB. 😅</p>",
      "rawMarkdown": "My binary classification model can't distinguish official data and my generated data. But my generated data still can't help to improve LB. 😅",
      "votes": 1,
      "replies": [
        {
          "id": 2030722,
          "postDate": "2022-11-15T15:22:35.413Z",
          "content": "<p>how is your <code>cv</code> and <code>lb</code> correlation ?</p>",
          "rawMarkdown": "how is your `cv` and `lb` correlation ?"
        },
        {
          "id": 2030752,
          "postDate": "2022-11-15T15:33:27.753Z",
          "content": "<p>I haven't found any correlation.😅</p>",
          "rawMarkdown": "I haven't found any correlation.😅",
          "votes": 1
        }
      ]
    },
    {
      "id": 2024260,
      "postDate": "2022-11-10T11:39:25.933Z",
      "content": "<p>It's a good idea, but the test set is very strong heterogeneous. Perhaps you need more than one model to suit all the possible cases only for test set splitting.</p>",
      "rawMarkdown": "It's a good idea, but the test set is very strong heterogeneous. Perhaps you need more than one model to suit all the possible cases only for test set splitting.",
      "votes": 2,
      "replies": [
        {
          "id": 2024358,
          "postDate": "2022-11-10T13:09:22.370Z",
          "content": "<p>yes even training and testing have bias, the 0.66 auc can tell.<br>\nFirst thing is trying to generate a dataset that match the training set, noise only at least, but it also failed.</p>",
          "rawMarkdown": "yes even training and testing have bias, the 0.66 auc can tell.\nFirst thing is trying to generate a dataset that match the training set, noise only at least, but it also failed.",
          "votes": 1
        },
        {
          "id": 2024380,
          "postDate": "2022-11-10T13:29:06.573Z",
          "content": "<p>Yes, I'm having the same issue with PyFstat, I can't replicate similar samples like some test data cases neither from the train cases. So, I'm using representative cases from both train and test, and generating noise only with PyFstat.</p>",
          "rawMarkdown": "Yes, I'm having the same issue with PyFstat, I can't replicate similar samples like some test data cases neither from the train cases. So, I'm using representative cases from both train and test, and generating noise only with PyFstat."
        }
      ]
    },
    {
      "id": 2031365,
      "postDate": "2022-11-16T03:13:29.570Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a> . Is the 0.66 auc a training score or a validation score?</p>",
      "rawMarkdown": "Hi @snaker . Is the 0.66 auc a training score or a validation score?",
      "replies": [
        {
          "id": 2032198,
          "postDate": "2022-11-16T13:30:37.383Z",
          "content": "<p>validation</p>",
          "rawMarkdown": "validation"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2030026,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2022-11-15T06:31:29.643000",
      "content": "<p>My binary classification model can't distinguish official data and my generated data. But my generated data still can't help to improve LB. 😅</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2030722,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2022-11-15T15:22:35.413000",
          "content": "<p>how is your <code>cv</code> and <code>lb</code> correlation ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2030752,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2022-11-15T15:33:27.753000",
          "content": "<p>I haven't found any correlation.😅</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2024260,
      "author_name": "Zollkron",
      "author_url": "",
      "post_date": "2022-11-10T11:39:25.933000",
      "content": "<p>It's a good idea, but the test set is very strong heterogeneous. Perhaps you need more than one model to suit all the possible cases only for test set splitting.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2024358,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2022-11-10T13:09:22.370000",
          "content": "<p>yes even training and testing have bias, the 0.66 auc can tell.<br>\nFirst thing is trying to generate a dataset that match the training set, noise only at least, but it also failed.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2024380,
          "author_name": "Zollkron",
          "author_url": "",
          "post_date": "2022-11-10T13:29:06.573000",
          "content": "<p>Yes, I'm having the same issue with PyFstat, I can't replicate similar samples like some test data cases neither from the train cases. So, I'm using representative cases from both train and test, and generating noise only with PyFstat.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2031365,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2022-11-16T03:13:29.570000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/snaker\" target=\"_blank\">@snaker</a> . Is the 0.66 auc a training score or a validation score?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2032198,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2022-11-16T13:30:37.383000",
          "content": "<p>validation</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2024110": "I have built a classifier that tries to split the 3 of the dataset: training, testing and generating.\n\nThe classifier to split testing and training can achieve ~0.66 auc score. But for testing and generating, or training and generating, the auc goes to 1.0 .\n\nI've done some work trying to mess the classifier but none of them work. I'm thinking maybe it is the key to achieve golden zone.\n\n---------------------------------------------------------------------------------------------------------------------\nUpdate:\n\n\nI have got the cause for this. I normalized the data with this method\n```\nnorm_data = data - data.min() / (data.max() - data.min())\n```\n\nIt will be easily influenced by the value of `max` and `min`. Turns out that the generating data of mine has very different `max` or `min` value than the training and testing dataset.\n\nAfter I changed it to\n\n```\nnorm_data = data * 1e22\n```\n\nThe classifier can only achieve 0.66 on classifying generating v.s. testing/training dataset.",
    "2030026": "My binary classification model can't distinguish official data and my generated data. But my generated data still can't help to improve LB. 😅",
    "2024260": "It's a good idea, but the test set is very strong heterogeneous. Perhaps you need more than one model to suit all the possible cases only for test set splitting.",
    "2031365": "Hi @snaker . Is the 0.66 auc a training score or a validation score?"
  }
}