{
  "id": 357387,
  "title": "Hidden Dataset should be balanced!",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/357387",
  "author_name": "nessim ben abbes",
  "post_date": "2022-10-04T06:41:35.066000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Seeing how a lote of people are struggling with getting more than random with their pretrained neural networks. I fear that an imbalanced dataset (just like the train dataset) will give us a winner that doesnt have a model that generalizes well. <br>\nSo please again if any organizer sees this, take extra care to make the hidden test dataset balanced(50 % CE / 50 % LAA). </p>",
  "messages": [
    {
      "id": 1970583,
      "postDate": "2022-10-04T06:41:35.067Z",
      "content": "<p>Seeing how a lote of people are struggling with getting more than random with their pretrained neural networks. I fear that an imbalanced dataset (just like the train dataset) will give us a winner that doesnt have a model that generalizes well. <br>\nSo please again if any organizer sees this, take extra care to make the hidden test dataset balanced(50 % CE / 50 % LAA). </p>",
      "rawMarkdown": "Seeing how a lote of people are struggling with getting more than random with their pretrained neural networks. I fear that an imbalanced dataset (just like the train dataset) will give us a winner that doesnt have a model that generalizes well. \nSo please again if any organizer sees this, take extra care to make the hidden test dataset balanced(50 % CE / 50 % LAA). ",
      "votes": 1
    },
    {
      "id": 1971468,
      "postDate": "2022-10-04T16:21:31.420Z",
      "content": "<p>Yeah, there's nothing they can do at this point, but hopefully you are correct and the private is balanced.</p>",
      "rawMarkdown": "Yeah, there's nothing they can do at this point, but hopefully you are correct and the private is balanced."
    },
    {
      "id": 1971425,
      "postDate": "2022-10-04T15:43:30.683Z",
      "content": "<p>Metric is balanced.<br>\nAnd always expected that test dataset has same distribution as train.</p>",
      "rawMarkdown": "Metric is balanced.\nAnd always expected that test dataset has same distribution as train."
    },
    {
      "id": 1970604,
      "postDate": "2022-10-04T06:51:22.197Z",
      "content": "<p>What you submit will be scored against all the dataset , only the score of  7% of the dataset is shown in the public leaderboard. </p>",
      "rawMarkdown": "What you submit will be scored against all the dataset , only the score of  7% of the dataset is shown in the public leaderboard. ",
      "replies": [
        {
          "id": 1970695,
          "postDate": "2022-10-04T08:23:50.560Z",
          "content": "<p>i am talking about the data distribution between the two classes on the hidden test as the train test is heavily biased towards one class</p>",
          "rawMarkdown": "i am talking about the data distribution between the two classes on the hidden test as the train test is heavily biased towards one class"
        },
        {
          "id": 1971124,
          "postDate": "2022-10-04T13:24:22.827Z",
          "content": "<p>Ya I know, what I mean is that now there nothing to do because  your private score is already calculated for the submissions you have made so far.</p>\n<p>So unless its already equally distributed, one can only wish. </p>",
          "rawMarkdown": "Ya I know, what I mean is that now there nothing to do because  your private score is already calculated for the submissions you have made so far.\n\nSo unless its already equally distributed, one can only wish. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1971468,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-10-04T16:21:31.420000",
      "content": "<p>Yeah, there's nothing they can do at this point, but hopefully you are correct and the private is balanced.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1971425,
      "author_name": "anthony",
      "author_url": "",
      "post_date": "2022-10-04T15:43:30.683000",
      "content": "<p>Metric is balanced.<br>\nAnd always expected that test dataset has same distribution as train.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1970604,
      "author_name": "Ali",
      "author_url": "",
      "post_date": "2022-10-04T06:51:22.197000",
      "content": "<p>What you submit will be scored against all the dataset , only the score of  7% of the dataset is shown in the public leaderboard. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1970695,
          "author_name": "nessim ben abbes",
          "author_url": "",
          "post_date": "2022-10-04T08:23:50.560000",
          "content": "<p>i am talking about the data distribution between the two classes on the hidden test as the train test is heavily biased towards one class</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1971124,
          "author_name": "Ali",
          "author_url": "",
          "post_date": "2022-10-04T13:24:22.827000",
          "content": "<p>Ya I know, what I mean is that now there nothing to do because  your private score is already calculated for the submissions you have made so far.</p>\n<p>So unless its already equally distributed, one can only wish. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1970583": "Seeing how a lote of people are struggling with getting more than random with their pretrained neural networks. I fear that an imbalanced dataset (just like the train dataset) will give us a winner that doesnt have a model that generalizes well. \nSo please again if any organizer sees this, take extra care to make the hidden test dataset balanced(50 % CE / 50 % LAA). ",
    "1971468": "Yeah, there's nothing they can do at this point, but hopefully you are correct and the private is balanced.",
    "1971425": "Metric is balanced.\nAnd always expected that test dataset has same distribution as train.",
    "1970604": "What you submit will be scored against all the dataset , only the score of  7% of the dataset is shown in the public leaderboard. "
  }
}