{
  "id": 112145,
  "title": "Will stage1 test data's label be released?",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/112145",
  "author_name": "DatNT",
  "post_date": "2019-10-11T03:42:31.752000",
  "votes": 15,
  "comment_count": 9,
  "views": 0,
  "content": "<p>just like the SIIM Pneumothorax or it wont be like Inclusive?</p>",
  "messages": [
    {
      "id": 646258,
      "postDate": "2019-10-11T03:42:31.753Z",
      "content": "<p>just like the SIIM Pneumothorax or it wont be like Inclusive?</p>",
      "rawMarkdown": "just like the SIIM Pneumothorax or it wont be like Inclusive?",
      "votes": 15
    },
    {
      "id": 650523,
      "postDate": "2019-10-16T13:23:09.537Z",
      "content": "<p>It woul'd be much better not to allow re-training and keep the model weights frozen, clear rules, less space for cheating, fully reproducible and verifiable submission, etc. In addition it would not put kagglers with less computational resources in even more disadvatagous position.</p>",
      "rawMarkdown": "It woul'd be much better not to allow re-training and keep the model weights frozen, clear rules, less space for cheating, fully reproducible and verifiable submission, etc. In addition it would not put kagglers with less computational resources in even more disadvatagous position.",
      "votes": 10,
      "replies": [
        {
          "id": 650829,
          "postDate": "2019-10-16T18:15:42.207Z",
          "content": "<blockquote>\n  <p>In addition it would not put kagglers with less computational resources in even more disadvatagous position.</p>\n</blockquote>\n\n<p>This point concerns me a lot, because stage 2 is only 6 days long, and it takes almost a day to simply download and preprocess data (if it will be an update to a train set archive, not a separate file). I guess most of the teams won't be able to re-train their models, and teams with lots of GPUs will benefit from it.</p>",
          "rawMarkdown": "&gt;In addition it would not put kagglers with less computational resources in even more disadvatagous position.\n\nThis point concerns me a lot, because stage 2 is only 6 days long, and it takes almost a day to simply download and preprocess data (if it will be an update to a train set archive, not a separate file). I guess most of the teams won't be able to re-train their models, and teams with lots of GPUs will benefit from it.",
          "votes": 1
        },
        {
          "id": 652164,
          "postDate": "2019-10-18T12:58:23.030Z",
          "content": "<p>&gt; This point concerns me a lot, because stage 2 is only 6 days long, and it takes almost a day to simply download and preprocess data </p>\n\n<p>Well, it depends on the number of samples in Stage 2 data.\nProbably it would not be  big, like Stage 1 Test Data.</p>",
          "rawMarkdown": "&gt; This point concerns me a lot, because stage 2 is only 6 days long, and it takes almost a day to simply download and preprocess data \n\nWell, it depends on the number of samples in Stage 2 data.\nProbably it would not be  big, like Stage 1 Test Data."
        },
        {
          "id": 652165,
          "postDate": "2019-10-18T13:00:56.370Z",
          "content": "<blockquote>\n  <p>It woul'd be much better not to allow re-training and keep the model weights frozen</p>\n</blockquote>\n\n<p>Wouldn't be easier not to provide ground truth for Stage 1 Test Data?\nNo true labels -&gt; no need to retrain.</p>",
          "rawMarkdown": "&gt; It woul'd be much better not to allow re-training and keep the model weights frozen\n\nWouldn't be easier not to provide ground truth for Stage 1 Test Data?\nNo true labels -&gt; no need to retrain.",
          "votes": 3
        }
      ]
    },
    {
      "id": 650687,
      "postDate": "2019-10-16T15:47:09.050Z",
      "content": "<p>Thanks for asking.</p>\n\n<p>As has been the standard in 2-stage Kaggle competitions, the labels for stage 1’s test set will be provided at the start of stage 2. And it will be permitted to retrain your locked in model on this additional data. This means you can load the new data and retrain but are not permitted to make scientific changes to your model beyond that.</p>\n\n<p>So for clarification,\nIn stage 1:\nTrain on stage 1 train set\nPredict on stage 1 test set</p>\n\n<p>In stage 2:\nTrain on stage 1 train set and (if desired) stage 1 test set\nPredict on stage 2 test set</p>",
      "rawMarkdown": "Thanks for asking.\n\nAs has been the standard in 2-stage Kaggle competitions, the labels for stage 1’s test set will be provided at the start of stage 2. And it will be permitted to retrain your locked in model on this additional data. This means you can load the new data and retrain but are not permitted to make scientific changes to your model beyond that.\n\nSo for clarification,\nIn stage 1:\nTrain on stage 1 train set\nPredict on stage 1 test set\n\nIn stage 2:\nTrain on stage 1 train set and (if desired) stage 1 test set\nPredict on stage 2 test set",
      "votes": 8,
      "replies": [
        {
          "id": 657387,
          "postDate": "2019-10-25T05:11:06.133Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> </p>\n\n<p>&gt; <strong>What happens if I want to change something in my code in the second stage?</strong>\nWe expect you may need to make some \"non scientific\" alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change. You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.</p>\n\n<p>May we use transfer learning of existing trained models on stage 1 train as the base for retraining for stage 2? So instead of retraining the network from scratch with stage 1 train+test labels, we would be using stage 1 training model as pre-trained network and then combine both stage 1 train/test as stage 1 train dataset, then train a couple of epochs and use the best performing epoch? I'm asking because it costly and time consuming to retrain an ensemble of models all from scratch.</p>\n\n<ul>\n<li>The only thing we would be doing in the transfer learning case is just finding the optimal epoch on all of stage 1 (for each model in an ensemble) and nothing else.</li>\n</ul>\n\n<p>Also can you clarify the point <code>Parameter tuning is permitted as long as it is fully automated</code>?</p>\n\n<p>What exactly is fully automated? What is NOT fully automated? Thanks </p>",
          "rawMarkdown": "@juliaelliott \n\n&gt; **What happens if I want to change something in my code in the second stage?**\nWe expect you may need to make some \"non scientific\" alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change. You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.\n\nMay we use transfer learning of existing trained models on stage 1 train as the base for retraining for stage 2? So instead of retraining the network from scratch with stage 1 train+test labels, we would be using stage 1 training model as pre-trained network and then combine both stage 1 train/test as stage 1 train dataset, then train a couple of epochs and use the best performing epoch? I'm asking because it costly and time consuming to retrain an ensemble of models all from scratch.\n\n- The only thing we would be doing in the transfer learning case is just finding the optimal epoch on all of stage 1 (for each model in an ensemble) and nothing else.\n\nAlso can you clarify the point `Parameter tuning is permitted as long as it is fully automated`?\n\nWhat exactly is fully automated? What is NOT fully automated? Thanks "
        },
        {
          "id": 659217,
          "postDate": "2019-10-27T09:35:44.443Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> If i want to finetune my models with new data (stage-1 test), then my retrain policy is probably going to be different from my training-from-scratch policy that I uploaded. In this case, do I have to upload code for both? e.g. Code for stage-1 submission and code for stage-2 training. Moreover, if I want to apply pseudo-labeling, should the code for this also be included in the upload?</p>",
          "rawMarkdown": "@juliaelliott If i want to finetune my models with new data (stage-1 test), then my retrain policy is probably going to be different from my training-from-scratch policy that I uploaded. In this case, do I have to upload code for both? e.g. Code for stage-1 submission and code for stage-2 training. Moreover, if I want to apply pseudo-labeling, should the code for this also be included in the upload?"
        },
        {
          "id": 660134,
          "postDate": "2019-10-28T18:28:25.817Z",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> You should include all code that you will be using in stage 2. This must be locked in at the end of stage 1's deadline. So you will not be able to scientifically change that code after the stage 1 deadline.</p>",
          "rawMarkdown": "@roguekk007 You should include all code that you will be using in stage 2. This must be locked in at the end of stage 1's deadline. So you will not be able to scientifically change that code after the stage 1 deadline."
        }
      ]
    },
    {
      "id": 648521,
      "postDate": "2019-10-14T10:16:02.453Z",
      "content": "<p>I'd also like to know if we'll be able to retrain uploaded models on new data or there will be no new data to train on?</p>",
      "rawMarkdown": "I'd also like to know if we'll be able to retrain uploaded models on new data or there will be no new data to train on?"
    }
  ],
  "comments": [
    {
      "id": 650523,
      "author_name": "Dmytro Poplavskiy",
      "author_url": "",
      "post_date": "2019-10-16T13:23:09.537000",
      "content": "<p>It woul'd be much better not to allow re-training and keep the model weights frozen, clear rules, less space for cheating, fully reproducible and verifiable submission, etc. In addition it would not put kagglers with less computational resources in even more disadvatagous position.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 650829,
          "author_name": "Dmytro Panchenko",
          "author_url": "",
          "post_date": "2019-10-16T18:15:42.207000",
          "content": "<blockquote>\n  <p>In addition it would not put kagglers with less computational resources in even more disadvatagous position.</p>\n</blockquote>\n\n<p>This point concerns me a lot, because stage 2 is only 6 days long, and it takes almost a day to simply download and preprocess data (if it will be an update to a train set archive, not a separate file). I guess most of the teams won't be able to re-train their models, and teams with lots of GPUs will benefit from it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 652164,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2019-10-18T12:58:23.030000",
          "content": "<p>&gt; This point concerns me a lot, because stage 2 is only 6 days long, and it takes almost a day to simply download and preprocess data </p>\n\n<p>Well, it depends on the number of samples in Stage 2 data.\nProbably it would not be  big, like Stage 1 Test Data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 652165,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2019-10-18T13:00:56.370000",
          "content": "<blockquote>\n  <p>It woul'd be much better not to allow re-training and keep the model weights frozen</p>\n</blockquote>\n\n<p>Wouldn't be easier not to provide ground truth for Stage 1 Test Data?\nNo true labels -&gt; no need to retrain.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 650687,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2019-10-16T15:47:09.050000",
      "content": "<p>Thanks for asking.</p>\n\n<p>As has been the standard in 2-stage Kaggle competitions, the labels for stage 1’s test set will be provided at the start of stage 2. And it will be permitted to retrain your locked in model on this additional data. This means you can load the new data and retrain but are not permitted to make scientific changes to your model beyond that.</p>\n\n<p>So for clarification,\nIn stage 1:\nTrain on stage 1 train set\nPredict on stage 1 test set</p>\n\n<p>In stage 2:\nTrain on stage 1 train set and (if desired) stage 1 test set\nPredict on stage 2 test set</p>",
      "votes": 8,
      "replies": [
        {
          "id": 657387,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-10-25T05:11:06.133000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> </p>\n\n<p>&gt; <strong>What happens if I want to change something in my code in the second stage?</strong>\nWe expect you may need to make some \"non scientific\" alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change. You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.</p>\n\n<p>May we use transfer learning of existing trained models on stage 1 train as the base for retraining for stage 2? So instead of retraining the network from scratch with stage 1 train+test labels, we would be using stage 1 training model as pre-trained network and then combine both stage 1 train/test as stage 1 train dataset, then train a couple of epochs and use the best performing epoch? I'm asking because it costly and time consuming to retrain an ensemble of models all from scratch.</p>\n\n<ul>\n<li>The only thing we would be doing in the transfer learning case is just finding the optimal epoch on all of stage 1 (for each model in an ensemble) and nothing else.</li>\n</ul>\n\n<p>Also can you clarify the point <code>Parameter tuning is permitted as long as it is fully automated</code>?</p>\n\n<p>What exactly is fully automated? What is NOT fully automated? Thanks </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 659217,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2019-10-27T09:35:44.443000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> If i want to finetune my models with new data (stage-1 test), then my retrain policy is probably going to be different from my training-from-scratch policy that I uploaded. In this case, do I have to upload code for both? e.g. Code for stage-1 submission and code for stage-2 training. Moreover, if I want to apply pseudo-labeling, should the code for this also be included in the upload?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 660134,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2019-10-28T18:28:25.817000",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> You should include all code that you will be using in stage 2. This must be locked in at the end of stage 1's deadline. So you will not be able to scientifically change that code after the stage 1 deadline.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 648521,
      "author_name": "Eek The Cat",
      "author_url": "",
      "post_date": "2019-10-14T10:16:02.453000",
      "content": "<p>I'd also like to know if we'll be able to retrain uploaded models on new data or there will be no new data to train on?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "646258": "just like the SIIM Pneumothorax or it wont be like Inclusive?",
    "650523": "It woul'd be much better not to allow re-training and keep the model weights frozen, clear rules, less space for cheating, fully reproducible and verifiable submission, etc. In addition it would not put kagglers with less computational resources in even more disadvatagous position.",
    "650687": "Thanks for asking.\n\nAs has been the standard in 2-stage Kaggle competitions, the labels for stage 1’s test set will be provided at the start of stage 2. And it will be permitted to retrain your locked in model on this additional data. This means you can load the new data and retrain but are not permitted to make scientific changes to your model beyond that.\n\nSo for clarification,\nIn stage 1:\nTrain on stage 1 train set\nPredict on stage 1 test set\n\nIn stage 2:\nTrain on stage 1 train set and (if desired) stage 1 test set\nPredict on stage 2 test set",
    "648521": "I'd also like to know if we'll be able to retrain uploaded models on new data or there will be no new data to train on?"
  }
}