{
  "id": 341904,
  "title": "Clarity about test data set size",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/341904",
  "author_name": "yqz",
  "post_date": "2022-08-04T17:53:14.143000",
  "votes": 0,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm not yet clear on how the test data works for this competition. My current understanding is that we can use the full training set to train a model in a separate notebook. We can then load the saved model and use it for a notebook that's only for inference.</p>\n<p>Within the inference notebook, my understanding is that we preprocess the test data, and apply our trained model to it and get the final results, which we save as a submission.csv. My belief was that typically the public LB will only show the results on a certain percentage of that test.csv, while the private LB will show the scores on the full test.csv. For example, the LB says 7% of the data is on the public LB, and the other 93% is included on the private LB.</p>\n<p>The test.csv file however only contains 4 rows. I've read in another notebook that the full test dataset will actually include 250 rows. If the test data has hidden data which we haven't yet seen, then we will need to write code to preprocess some hidden test data correct? This seems inconvenient since potentially there will be test data which breaks our preprocessing code. Also, I don't think it makes any sense either to test our models on only 4 images (4/754~=0.5% of the data). It seems to be too small of a sample to really use for testing our models.</p>\n<p>Can someone explain exactly what's going on with the test data? Is it really just 4 rows?</p>",
  "messages": [
    {
      "id": 1885139,
      "postDate": "2022-08-04T23:48:31.203Z",
      "content": "<p>One idea is that they don't want you to know what kind of data the model will need to be run on \"once deployed\". It's sort of similar to in practice deploying a model and not having access to future data your model will be run on. I think it's safe to just ignore those 4 since they're pretty useless</p>",
      "rawMarkdown": "One idea is that they don't want you to know what kind of data the model will need to be run on \"once deployed\". It's sort of similar to in practice deploying a model and not having access to future data your model will be run on. I think it's safe to just ignore those 4 since they're pretty useless",
      "votes": 1,
      "replies": [
        {
          "id": 1885141,
          "postDate": "2022-08-04T23:53:51.300Z",
          "content": "<p>I see, thanks for the response. So is the assumption you're going on just that there's some hidden test data that we will need to preprocess and then do inference on with our trained models? </p>",
          "rawMarkdown": "I see, thanks for the response. So is the assumption you're going on just that there's some hidden test data that we will need to preprocess and then do inference on with our trained models? "
        },
        {
          "id": 1885157,
          "postDate": "2022-08-05T00:48:20.797Z",
          "content": "<p>Yes, when you make a submission, your code needs to run on the hidden test set</p>",
          "rawMarkdown": "Yes, when you make a submission, your code needs to run on the hidden test set",
          "votes": 2
        },
        {
          "id": 1885187,
          "postDate": "2022-08-05T02:25:55.100Z",
          "content": "<p>Understood, thanks!</p>",
          "rawMarkdown": "Understood, thanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1884873,
      "postDate": "2022-08-04T17:53:14.143Z",
      "content": "<p>I'm not yet clear on how the test data works for this competition. My current understanding is that we can use the full training set to train a model in a separate notebook. We can then load the saved model and use it for a notebook that's only for inference.</p>\n<p>Within the inference notebook, my understanding is that we preprocess the test data, and apply our trained model to it and get the final results, which we save as a submission.csv. My belief was that typically the public LB will only show the results on a certain percentage of that test.csv, while the private LB will show the scores on the full test.csv. For example, the LB says 7% of the data is on the public LB, and the other 93% is included on the private LB.</p>\n<p>The test.csv file however only contains 4 rows. I've read in another notebook that the full test dataset will actually include 250 rows. If the test data has hidden data which we haven't yet seen, then we will need to write code to preprocess some hidden test data correct? This seems inconvenient since potentially there will be test data which breaks our preprocessing code. Also, I don't think it makes any sense either to test our models on only 4 images (4/754~=0.5% of the data). It seems to be too small of a sample to really use for testing our models.</p>\n<p>Can someone explain exactly what's going on with the test data? Is it really just 4 rows?</p>",
      "rawMarkdown": "I'm not yet clear on how the test data works for this competition. My current understanding is that we can use the full training set to train a model in a separate notebook. We can then load the saved model and use it for a notebook that's only for inference.\n\nWithin the inference notebook, my understanding is that we preprocess the test data, and apply our trained model to it and get the final results, which we save as a submission.csv. My belief was that typically the public LB will only show the results on a certain percentage of that test.csv, while the private LB will show the scores on the full test.csv. For example, the LB says 7% of the data is on the public LB, and the other 93% is included on the private LB.\n\nThe test.csv file however only contains 4 rows. I've read in another notebook that the full test dataset will actually include 250 rows. If the test data has hidden data which we haven't yet seen, then we will need to write code to preprocess some hidden test data correct? This seems inconvenient since potentially there will be test data which breaks our preprocessing code. Also, I don't think it makes any sense either to test our models on only 4 images (4/754~=0.5% of the data). It seems to be too small of a sample to really use for testing our models.\n\nCan someone explain exactly what's going on with the test data? Is it really just 4 rows?"
    }
  ],
  "comments": [
    {
      "id": 1885139,
      "author_name": "cosmosaa",
      "author_url": "",
      "post_date": "2022-08-04T23:48:31.203000",
      "content": "<p>One idea is that they don't want you to know what kind of data the model will need to be run on \"once deployed\". It's sort of similar to in practice deploying a model and not having access to future data your model will be run on. I think it's safe to just ignore those 4 since they're pretty useless</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1885141,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-04T23:53:51.300000",
          "content": "<p>I see, thanks for the response. So is the assumption you're going on just that there's some hidden test data that we will need to preprocess and then do inference on with our trained models? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1885157,
          "author_name": "cosmosaa",
          "author_url": "",
          "post_date": "2022-08-05T00:48:20.797000",
          "content": "<p>Yes, when you make a submission, your code needs to run on the hidden test set</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1885187,
          "author_name": "yqz",
          "author_url": "",
          "post_date": "2022-08-05T02:25:55.100000",
          "content": "<p>Understood, thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1885139": "One idea is that they don't want you to know what kind of data the model will need to be run on \"once deployed\". It's sort of similar to in practice deploying a model and not having access to future data your model will be run on. I think it's safe to just ignore those 4 since they're pretty useless",
    "1884873": "I'm not yet clear on how the test data works for this competition. My current understanding is that we can use the full training set to train a model in a separate notebook. We can then load the saved model and use it for a notebook that's only for inference.\n\nWithin the inference notebook, my understanding is that we preprocess the test data, and apply our trained model to it and get the final results, which we save as a submission.csv. My belief was that typically the public LB will only show the results on a certain percentage of that test.csv, while the private LB will show the scores on the full test.csv. For example, the LB says 7% of the data is on the public LB, and the other 93% is included on the private LB.\n\nThe test.csv file however only contains 4 rows. I've read in another notebook that the full test dataset will actually include 250 rows. If the test data has hidden data which we haven't yet seen, then we will need to write code to preprocess some hidden test data correct? This seems inconvenient since potentially there will be test data which breaks our preprocessing code. Also, I don't think it makes any sense either to test our models on only 4 images (4/754~=0.5% of the data). It seems to be too small of a sample to really use for testing our models.\n\nCan someone explain exactly what's going on with the test data? Is it really just 4 rows?"
  }
}