{
  "id": 183064,
  "title": "Can Someone Please Explain Meaning of Inference only Competition ?",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/183064",
  "author_name": "Athar Sayed",
  "post_date": "2020-09-15T10:45:11.302000",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>As per the guidelines in Data Section it is given .<br>\n<strong>\"This competition is inference-only, meaning that your submitted kernels will not have access to the training set.\n\"</strong><br>\nAnd Also in Description it is written<br>\n<strong>train - all train images (note that your submission kernels will NOT have access to this set of images, so you must build your models elsewhere and incorporate them into your submissions)</strong></p>\n<p>But This is also a Kernel Competition so this means we have to build our model either locally or in Kaggle Kernel and then for Final Submission we have to Load our Model and Make Predictions on Final Test Set .</p>\n<p>Am I right here ?</p>",
  "messages": [
    {
      "id": 1011359,
      "postDate": "2020-09-15T12:11:56.027Z",
      "content": "<p>Steps:</p>\n<ol>\n<li>Build your model either locally or in Notebook. Save trained model. Can use TPU and any amount of resources.</li>\n<li>Load trained model in a dataset</li>\n<li>Run Notebook in \"commit\" mode that uses model and hidden test data to create submission file. You do not have outside access to real test dataset, so you cannot preprocess elsewhere.</li>\n</ol>\n<p>-Rich</p>",
      "rawMarkdown": "Steps:\n\n1. Build your model either locally or in Notebook. Save trained model. Can use TPU and any amount of resources.\n2. Load trained model in a dataset\n3. Run Notebook in \"commit\" mode that uses model and hidden test data to create submission file. You do not have outside access to real test dataset, so you cannot preprocess elsewhere.\n\n-Rich",
      "votes": 1
    },
    {
      "id": 1011257,
      "postDate": "2020-09-15T10:45:11.303Z",
      "content": "<p>As per the guidelines in Data Section it is given .<br>\n<strong>\"This competition is inference-only, meaning that your submitted kernels will not have access to the training set.\n\"</strong><br>\nAnd Also in Description it is written<br>\n<strong>train - all train images (note that your submission kernels will NOT have access to this set of images, so you must build your models elsewhere and incorporate them into your submissions)</strong></p>\n<p>But This is also a Kernel Competition so this means we have to build our model either locally or in Kaggle Kernel and then for Final Submission we have to Load our Model and Make Predictions on Final Test Set .</p>\n<p>Am I right here ?</p>",
      "rawMarkdown": "As per the guidelines in Data Section it is given .\n**\"This competition is inference-only, meaning that your submitted kernels will not have access to the training set.\n\"**\nAnd Also in Description it is written\n**train - all train images (note that your submission kernels will NOT have access to this set of images, so you must build your models elsewhere and incorporate them into your submissions)**\n\nBut This is also a Kernel Competition so this means we have to build our model either locally or in Kaggle Kernel and then for Final Submission we have to Load our Model and Make Predictions on Final Test Set .\n\nAm I right here ?\n",
      "votes": 1
    },
    {
      "id": 1017869,
      "postDate": "2020-09-19T10:04:42.133Z",
      "content": "<p>I was thinking of following procedure:</p>\n<ol>\n<li>Preprocess train and test datasets and create new dataset with the result</li>\n<li>Build model based on this dataset, save in another dataset</li>\n<li>Submission notebook, use model (from own dataset) and preprocessed test data (from own dataset) to generate submission.</li>\n</ol>\n<p>Technically one could combine steps 2+3 into one kernel which just depends on my own preprocessed dataset, at least if no pre-trained models are used. Is there something wrong with that?</p>",
      "rawMarkdown": "I was thinking of following procedure:\n1. Preprocess train and test datasets and create new dataset with the result\n2. Build model based on this dataset, save in another dataset\n3. Submission notebook, use model (from own dataset) and preprocessed test data (from own dataset) to generate submission.\n\nTechnically one could combine steps 2+3 into one kernel which just depends on my own preprocessed dataset, at least if no pre-trained models are used. Is there something wrong with that?",
      "replies": [
        {
          "id": 1018027,
          "postDate": "2020-09-19T11:57:56.850Z",
          "content": "<p>These steps seem to work. But you can only preprocess and inference on the public part of the dataset. That is fine for testing and seeing if a model is moving in the right direction.</p>\n<p>Eventually, you'll need to preprocess and inference on the private data in a committed notebook.</p>",
          "rawMarkdown": "These steps seem to work. But you can only preprocess and inference on the public part of the dataset. That is fine for testing and seeing if a model is moving in the right direction.\n\nEventually, you'll need to preprocess and inference on the private data in a committed notebook."
        }
      ]
    },
    {
      "id": 1011641,
      "postDate": "2020-09-15T15:53:23.733Z",
      "content": "<p>Exactly as Rich has stated, training must be done locally or in a non-submission notebook, as your submission notebooks <strong>will not have access to the train set</strong> during the private re-run of your code. It will only have access to the test set in order for your trained model to produce predictions on (i.e. run inference on) that private test set. Practically speaking, the size of the train set is prohibitively large for training in notebooks, but rather than reduce the train set size, we are isolating submissions to inference-only.</p>",
      "rawMarkdown": "Exactly as Rich has stated, training must be done locally or in a non-submission notebook, as your submission notebooks **will not have access to the train set** during the private re-run of your code. It will only have access to the test set in order for your trained model to produce predictions on (i.e. run inference on) that private test set. Practically speaking, the size of the train set is prohibitively large for training in notebooks, but rather than reduce the train set size, we are isolating submissions to inference-only.",
      "replies": [
        {
          "id": 1011780,
          "postDate": "2020-09-15T17:16:33.440Z",
          "content": "<p>Julia, </p>\n<p>Do you want to setup the find a teammate forum?</p>\n<p>Thanks Julia for the consideration!</p>",
          "rawMarkdown": "Julia, \n\nDo you want to setup the find a teammate forum?\n\nThanks Julia for the consideration!"
        }
      ]
    },
    {
      "id": 1011310,
      "postDate": "2020-09-15T11:37:21.170Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1011359,
      "author_name": "quadcore/Richard Epstein",
      "author_url": "",
      "post_date": "2020-09-15T12:11:56.027000",
      "content": "<p>Steps:</p>\n<ol>\n<li>Build your model either locally or in Notebook. Save trained model. Can use TPU and any amount of resources.</li>\n<li>Load trained model in a dataset</li>\n<li>Run Notebook in \"commit\" mode that uses model and hidden test data to create submission file. You do not have outside access to real test dataset, so you cannot preprocess elsewhere.</li>\n</ol>\n<p>-Rich</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1017869,
      "author_name": "Alex Bader",
      "author_url": "",
      "post_date": "2020-09-19T10:04:42.133000",
      "content": "<p>I was thinking of following procedure:</p>\n<ol>\n<li>Preprocess train and test datasets and create new dataset with the result</li>\n<li>Build model based on this dataset, save in another dataset</li>\n<li>Submission notebook, use model (from own dataset) and preprocessed test data (from own dataset) to generate submission.</li>\n</ol>\n<p>Technically one could combine steps 2+3 into one kernel which just depends on my own preprocessed dataset, at least if no pre-trained models are used. Is there something wrong with that?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1018027,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-09-19T11:57:56.850000",
          "content": "<p>These steps seem to work. But you can only preprocess and inference on the public part of the dataset. That is fine for testing and seeing if a model is moving in the right direction.</p>\n<p>Eventually, you'll need to preprocess and inference on the private data in a committed notebook.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1011641,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2020-09-15T15:53:23.733000",
      "content": "<p>Exactly as Rich has stated, training must be done locally or in a non-submission notebook, as your submission notebooks <strong>will not have access to the train set</strong> during the private re-run of your code. It will only have access to the test set in order for your trained model to produce predictions on (i.e. run inference on) that private test set. Practically speaking, the size of the train set is prohibitively large for training in notebooks, but rather than reduce the train set size, we are isolating submissions to inference-only.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1011780,
          "author_name": "Bo Peng",
          "author_url": "",
          "post_date": "2020-09-15T17:16:33.440000",
          "content": "<p>Julia, </p>\n<p>Do you want to setup the find a teammate forum?</p>\n<p>Thanks Julia for the consideration!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1011310,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-15T11:37:21.170000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1011359": "Steps:\n\n1. Build your model either locally or in Notebook. Save trained model. Can use TPU and any amount of resources.\n2. Load trained model in a dataset\n3. Run Notebook in \"commit\" mode that uses model and hidden test data to create submission file. You do not have outside access to real test dataset, so you cannot preprocess elsewhere.\n\n-Rich",
    "1011257": "As per the guidelines in Data Section it is given .\n**\"This competition is inference-only, meaning that your submitted kernels will not have access to the training set.\n\"**\nAnd Also in Description it is written\n**train - all train images (note that your submission kernels will NOT have access to this set of images, so you must build your models elsewhere and incorporate them into your submissions)**\n\nBut This is also a Kernel Competition so this means we have to build our model either locally or in Kaggle Kernel and then for Final Submission we have to Load our Model and Make Predictions on Final Test Set .\n\nAm I right here ?\n",
    "1017869": "I was thinking of following procedure:\n1. Preprocess train and test datasets and create new dataset with the result\n2. Build model based on this dataset, save in another dataset\n3. Submission notebook, use model (from own dataset) and preprocessed test data (from own dataset) to generate submission.\n\nTechnically one could combine steps 2+3 into one kernel which just depends on my own preprocessed dataset, at least if no pre-trained models are used. Is there something wrong with that?",
    "1011641": "Exactly as Rich has stated, training must be done locally or in a non-submission notebook, as your submission notebooks **will not have access to the train set** during the private re-run of your code. It will only have access to the test set in order for your trained model to produce predictions on (i.e. run inference on) that private test set. Practically speaking, the size of the train set is prohibitively large for training in notebooks, but rather than reduce the train set size, we are isolating submissions to inference-only.",
    "1011310": ""
  }
}