{
  "id": 344308,
  "title": "Notebook run times for submission",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/344308",
  "author_name": "tdiceman",
  "post_date": "2022-08-14T19:28:32.718000",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I have a novice question: </p>\n<p><strong>Is the submission notebook inference going through entire test data at this time, but providing public scores only on 7% of that prediction?</strong></p>\n<p>I ask because my submission currently takes 1.5 hours to successfully submit (when it runs fine that is). From a few other discussions I see that the current 7% test data is around 20 images, and the full test data has 280 images. If 20 images is taking 1.5 hours, I believe final submission will not be possible due to 9 hour notebook run time limit. Plus, on training hold out data, my inference completes in 25 min for about 50 images, so what's going on while submission scoring?</p>\n<p>Main idea is to clarify if our current notebooks should be completing inference in under 40 min if they are only running on 7% data and  are to successfully complete inference on 100% data for final submission in under 9 hours.</p>\n<p>Thanks.</p>",
  "messages": [
    {
      "id": 1898755,
      "postDate": "2022-08-14T19:28:32.717Z",
      "content": "<p>I have a novice question: </p>\n<p><strong>Is the submission notebook inference going through entire test data at this time, but providing public scores only on 7% of that prediction?</strong></p>\n<p>I ask because my submission currently takes 1.5 hours to successfully submit (when it runs fine that is). From a few other discussions I see that the current 7% test data is around 20 images, and the full test data has 280 images. If 20 images is taking 1.5 hours, I believe final submission will not be possible due to 9 hour notebook run time limit. Plus, on training hold out data, my inference completes in 25 min for about 50 images, so what's going on while submission scoring?</p>\n<p>Main idea is to clarify if our current notebooks should be completing inference in under 40 min if they are only running on 7% data and  are to successfully complete inference on 100% data for final submission in under 9 hours.</p>\n<p>Thanks.</p>",
      "rawMarkdown": "I have a novice question: \n\n**Is the submission notebook inference going through entire test data at this time, but providing public scores only on 7% of that prediction?**\n\nI ask because my submission currently takes 1.5 hours to successfully submit (when it runs fine that is). From a few other discussions I see that the current 7% test data is around 20 images, and the full test data has 280 images. If 20 images is taking 1.5 hours, I believe final submission will not be possible due to 9 hour notebook run time limit. Plus, on training hold out data, my inference completes in 25 min for about 50 images, so what's going on while submission scoring?\n\nMain idea is to clarify if our current notebooks should be completing inference in under 40 min if they are only running on 7% data and  are to successfully complete inference on 100% data for final submission in under 9 hours.\n\nThanks.",
      "votes": 4
    },
    {
      "id": 1949101,
      "postDate": "2022-09-21T14:00:16.990Z",
      "content": "<p>Hello,<br>\nDid you ever find out if the submission happens for the 4 patients in the test set shared for us or the entire hidden test set? I am having trouble understanding it. I can do inference for the shared test set in 10 minutes however once I submit the notebook it either times out or throws a notebook exception.</p>",
      "rawMarkdown": "Hello,\nDid you ever find out if the submission happens for the 4 patients in the test set shared for us or the entire hidden test set? I am having trouble understanding it. I can do inference for the shared test set in 10 minutes however once I submit the notebook it either times out or throws a notebook exception.",
      "replies": [
        {
          "id": 1949256,
          "postDate": "2022-09-21T15:41:39.890Z",
          "content": "<p><a href=\"https://www.kaggle.com/mertlostar\" target=\"_blank\">@mertlostar</a> submission scoring is most likely happening on all 280 images of the hidden test data. This is kind of confirmed by the following two things:</p>\n<p>1) Competition hosts mentioned in one post that kaggle is monitoring the private LB score, which means that whenever we submit for scoring, our notebooks run on entire dataset but public LB score is computed on 7% (or around 20 images) to give us a feedback.<br>\n2) I have run a no inference notebook, which reads images upto 2Gb in size, resizes it, deletes it and submits a 0.5 probability for all test cases. That notebook ran for over an hour, which means that 280 images were read and processed - timing does not make sense for 4 or 20 images.</p>\n<p>In your case, I feel 10 min is long for 4 image inference, you may need to find a way to make the inference quicker. The hidden test data probably has images that are much larger than the 1.3 Gb largest image in the 4 image test data that is provided only for creating and checking our submission notebook. I suggest you begin by limiting the image size you consider for inference - that should take care of OOM exceptions being thrown and also will reduce your inference time to within 9 hours. You can use the following code block for this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8536316%2F77c6af6be8105d5ce52331911c0c080c%2FScreenshot%202022-09-21%20104434%20-%20Copy.png?generation=1663775135504834&amp;alt=media\" alt=\"\"></p>\n<p>All the best, and do update here if you need further help. Let's try and get some successful submissions in the time left :)</p>",
          "rawMarkdown": "@mertlostar submission scoring is most likely happening on all 280 images of the hidden test data. This is kind of confirmed by the following two things:\n\n1) Competition hosts mentioned in one post that kaggle is monitoring the private LB score, which means that whenever we submit for scoring, our notebooks run on entire dataset but public LB score is computed on 7% (or around 20 images) to give us a feedback.\n2) I have run a no inference notebook, which reads images upto 2Gb in size, resizes it, deletes it and submits a 0.5 probability for all test cases. That notebook ran for over an hour, which means that 280 images were read and processed - timing does not make sense for 4 or 20 images.\n\nIn your case, I feel 10 min is long for 4 image inference, you may need to find a way to make the inference quicker. The hidden test data probably has images that are much larger than the 1.3 Gb largest image in the 4 image test data that is provided only for creating and checking our submission notebook. I suggest you begin by limiting the image size you consider for inference - that should take care of OOM exceptions being thrown and also will reduce your inference time to within 9 hours. You can use the following code block for this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8536316%2F77c6af6be8105d5ce52331911c0c080c%2FScreenshot%202022-09-21%20104434%20-%20Copy.png?generation=1663775135504834&alt=media)\n\nAll the best, and do update here if you need further help. Let's try and get some successful submissions in the time left :)",
          "votes": 3
        },
        {
          "id": 1949653,
          "postDate": "2022-09-21T19:28:48.613Z",
          "content": "<p>Such a complete and helpful answer. Thanks!</p>",
          "rawMarkdown": "Such a complete and helpful answer. Thanks!",
          "votes": 1
        },
        {
          "id": 1965899,
          "postDate": "2022-10-01T15:48:13.140Z",
          "content": "<p>Still facing \"Notebook Exceeded Allowed Compute\" error. <a href=\"https://www.kaggle.com/icemantd\" target=\"_blank\">@icemantd</a> can you give me any suggestion. Thanks in advance.</p>",
          "rawMarkdown": "Still facing \"Notebook Exceeded Allowed Compute\" error. @icemantd can you give me any suggestion. Thanks in advance."
        },
        {
          "id": 1965971,
          "postDate": "2022-10-01T16:46:11.143Z",
          "content": "<p><a href=\"https://www.kaggle.com/bakar31\" target=\"_blank\">@bakar31</a> have you tried the above code block to limit file size being read, and if yes, what size limit did you use? As far as I have seen, OOM error in this case is related to RAM use alongside large image sizes being read. If you can share your code structure alone for file reading we can try to check some things. Else I suggest picking large images from the training set, up to 2.5 Gb and above (there are a few) and testing your submission on that. It should fail for large images on this check if it is failing on final submission. </p>\n<p>Hope you find the problem soon… all the best.</p>",
          "rawMarkdown": "@bakar31 have you tried the above code block to limit file size being read, and if yes, what size limit did you use? As far as I have seen, OOM error in this case is related to RAM use alongside large image sizes being read. If you can share your code structure alone for file reading we can try to check some things. Else I suggest picking large images from the training set, up to 2.5 Gb and above (there are a few) and testing your submission on that. It should fail for large images on this check if it is failing on final submission. \n\nHope you find the problem soon... all the best.",
          "votes": 2
        },
        {
          "id": 1966718,
          "postDate": "2022-10-02T06:21:44.027Z",
          "content": "<p>Great advice <a href=\"https://www.kaggle.com/icemantd\" target=\"_blank\">@icemantd</a>. Thank you</p>",
          "rawMarkdown": "Great advice @icemantd. Thank you\n",
          "votes": 1
        },
        {
          "id": 1970513,
          "postDate": "2022-10-04T05:57:55.157Z",
          "content": "<p>Hey, <a href=\"https://www.kaggle.com/icemantd\" target=\"_blank\">@icemantd</a> I am getting error with even the  0.7 threshold which works fine of t train data.</p>\n<p><code>img = tf.keras.utils.load_img(img_path, target_size=(512, 512))\n            img_array = tf.keras.utils.img_to_array(img)\n            img_array = tf.expand_dims(img_array, 0) # Create a batch\n            preds_0.append(efficentB0.predict(img_array).flatten())</code></p>\n<p>here is my inference code block.</p>",
          "rawMarkdown": "Hey, @icemantd I am getting error with even the  0.7 threshold which works fine of t train data.\n\n`img = tf.keras.utils.load_img(img_path, target_size=(512, 512))\n            img_array = tf.keras.utils.img_to_array(img)\n            img_array = tf.expand_dims(img_array, 0) # Create a batch\n            preds_0.append(efficentB0.predict(img_array).flatten())`\n\nhere is my inference code block."
        },
        {
          "id": 1970578,
          "postDate": "2022-10-04T06:34:34.187Z",
          "content": "<p><a href=\"https://www.kaggle.com/bakar31\" target=\"_blank\">@bakar31</a> I am assuming by 0.7 you mean 0.7 Gb file size, which is too small on its own to create any issues for the RAM (again assuming you still get the 'exceeded allowed compute error'). The only thing I can think of, looking at the code you have shared, is the conversion of uint8 to a numpy array. This will convert the image data to float, which usually increases memory usage, especially if both instances remain in memory. I'm not sure if the increase due to conversion to float for a file up to 0.7 Gb would cause this problem (and I'm assuming you have checked with training data upto 0.7 Gb). </p>\n<p>Please make sure to check with training data with specific image sizes of around 0.7 Gb or even slightly higher. And if possible, check without the img_to_array once. Do these on training data and make sure, I guess you have 1 or 2 submissions left. I cannot think of anything else if training data works fine upto and around you file size limit (which to be honest is way too low for getting this error). </p>\n<p>PS. Hope you are not creating a batch of image arrays for prediction… Just clarifying, even though it might be slower, try one image inference at a time, with gc.collect(). <a href=\"https://www.kaggle.com/code/realneuralnetwork/cnn-strip-ai-inference/notebook?scriptVersionId=100707854\" target=\"_blank\">Here</a> is a full notebook that has inspired most of the fixes. </p>",
          "rawMarkdown": "@bakar31 I am assuming by 0.7 you mean 0.7 Gb file size, which is too small on its own to create any issues for the RAM (again assuming you still get the 'exceeded allowed compute error'). The only thing I can think of, looking at the code you have shared, is the conversion of uint8 to a numpy array. This will convert the image data to float, which usually increases memory usage, especially if both instances remain in memory. I'm not sure if the increase due to conversion to float for a file up to 0.7 Gb would cause this problem (and I'm assuming you have checked with training data upto 0.7 Gb). \n\nPlease make sure to check with training data with specific image sizes of around 0.7 Gb or even slightly higher. And if possible, check without the img_to_array once. Do these on training data and make sure, I guess you have 1 or 2 submissions left. I cannot think of anything else if training data works fine upto and around you file size limit (which to be honest is way too low for getting this error). \n\nPS. Hope you are not creating a batch of image arrays for prediction... Just clarifying, even though it might be slower, try one image inference at a time, with gc.collect(). [Here](https://www.kaggle.com/code/realneuralnetwork/cnn-strip-ai-inference/notebook?scriptVersionId=100707854) is a full notebook that has inspired most of the fixes. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1901526,
      "postDate": "2022-08-16T17:53:04.693Z",
      "content": "<p>I am still getting \"Notebook throw Exception\"  on submite. That is after images were re-sized and zoomed-off. We don't have a issue with reading image and the Image GB value. Similar to your earlier error but no LOG for submitted notebook as the Version-Notebooks are our local one's. <br>\n1) What is the direct method to see Submitted-notedbook LOG page ? (Log error if we have try: except:)<br>\n2) What is the direct method to see our submission page route paths ?</p>\n<p>Thanks <a href=\"https://www.kaggle.com/icemantd\" target=\"_blank\">@icemantd</a> </p>",
      "rawMarkdown": "I am still getting \"Notebook throw Exception\"  on submite. That is after images were re-sized and zoomed-off. We don't have a issue with reading image and the Image GB value. Similar to your earlier error but no LOG for submitted notebook as the Version-Notebooks are our local one's. \n1) What is the direct method to see Submitted-notedbook LOG page ? (Log error if we have try: except:)\n2) What is the direct method to see our submission page route paths ?\n\nThanks @icemantd \n\n",
      "replies": [
        {
          "id": 1901570,
          "postDate": "2022-08-16T18:55:50.977Z",
          "content": "<p>There is no way to see the log output of the notebooks submitted for scoring… you have to figure out from <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">this</a> limited information and try to debug your code</p>\n<p>'Notebook threw exception' in my case has been due to some error in handling all possible failure cases. For e.g., you might skip an image based on try: except: but do not add the scoring for that to you submission data frame, so while compiling the score into submission.csv, there may be errors. Things like that need to be debugged, but exact error cannot be seen unfortunately.</p>",
          "rawMarkdown": "There is no way to see the log output of the notebooks submitted for scoring... you have to figure out from [this](https://www.kaggle.com/code-competition-debugging) limited information and try to debug your code\n\n'Notebook threw exception' in my case has been due to some error in handling all possible failure cases. For e.g., you might skip an image based on try: except: but do not add the scoring for that to you submission data frame, so while compiling the score into submission.csv, there may be errors. Things like that need to be debugged, but exact error cannot be seen unfortunately.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1949101,
      "author_name": "Mert L.",
      "author_url": "",
      "post_date": "2022-09-21T14:00:16.990000",
      "content": "<p>Hello,<br>\nDid you ever find out if the submission happens for the 4 patients in the test set shared for us or the entire hidden test set? I am having trouble understanding it. I can do inference for the shared test set in 10 minutes however once I submit the notebook it either times out or throws a notebook exception.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1949256,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-09-21T15:41:39.890000",
          "content": "<p><a href=\"https://www.kaggle.com/mertlostar\" target=\"_blank\">@mertlostar</a> submission scoring is most likely happening on all 280 images of the hidden test data. This is kind of confirmed by the following two things:</p>\n<p>1) Competition hosts mentioned in one post that kaggle is monitoring the private LB score, which means that whenever we submit for scoring, our notebooks run on entire dataset but public LB score is computed on 7% (or around 20 images) to give us a feedback.<br>\n2) I have run a no inference notebook, which reads images upto 2Gb in size, resizes it, deletes it and submits a 0.5 probability for all test cases. That notebook ran for over an hour, which means that 280 images were read and processed - timing does not make sense for 4 or 20 images.</p>\n<p>In your case, I feel 10 min is long for 4 image inference, you may need to find a way to make the inference quicker. The hidden test data probably has images that are much larger than the 1.3 Gb largest image in the 4 image test data that is provided only for creating and checking our submission notebook. I suggest you begin by limiting the image size you consider for inference - that should take care of OOM exceptions being thrown and also will reduce your inference time to within 9 hours. You can use the following code block for this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8536316%2F77c6af6be8105d5ce52331911c0c080c%2FScreenshot%202022-09-21%20104434%20-%20Copy.png?generation=1663775135504834&amp;alt=media\" alt=\"\"></p>\n<p>All the best, and do update here if you need further help. Let's try and get some successful submissions in the time left :)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1949653,
          "author_name": "Mert L.",
          "author_url": "",
          "post_date": "2022-09-21T19:28:48.613000",
          "content": "<p>Such a complete and helpful answer. Thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1965899,
          "author_name": "Abu Bakar",
          "author_url": "",
          "post_date": "2022-10-01T15:48:13.140000",
          "content": "<p>Still facing \"Notebook Exceeded Allowed Compute\" error. <a href=\"https://www.kaggle.com/icemantd\" target=\"_blank\">@icemantd</a> can you give me any suggestion. Thanks in advance.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1965971,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-01T16:46:11.143000",
          "content": "<p><a href=\"https://www.kaggle.com/bakar31\" target=\"_blank\">@bakar31</a> have you tried the above code block to limit file size being read, and if yes, what size limit did you use? As far as I have seen, OOM error in this case is related to RAM use alongside large image sizes being read. If you can share your code structure alone for file reading we can try to check some things. Else I suggest picking large images from the training set, up to 2.5 Gb and above (there are a few) and testing your submission on that. It should fail for large images on this check if it is failing on final submission. </p>\n<p>Hope you find the problem soon… all the best.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1966718,
          "author_name": "Abu Bakar",
          "author_url": "",
          "post_date": "2022-10-02T06:21:44.027000",
          "content": "<p>Great advice <a href=\"https://www.kaggle.com/icemantd\" target=\"_blank\">@icemantd</a>. Thank you</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1970513,
          "author_name": "Abu Bakar",
          "author_url": "",
          "post_date": "2022-10-04T05:57:55.157000",
          "content": "<p>Hey, <a href=\"https://www.kaggle.com/icemantd\" target=\"_blank\">@icemantd</a> I am getting error with even the  0.7 threshold which works fine of t train data.</p>\n<p><code>img = tf.keras.utils.load_img(img_path, target_size=(512, 512))\n            img_array = tf.keras.utils.img_to_array(img)\n            img_array = tf.expand_dims(img_array, 0) # Create a batch\n            preds_0.append(efficentB0.predict(img_array).flatten())</code></p>\n<p>here is my inference code block.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1970578,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-04T06:34:34.187000",
          "content": "<p><a href=\"https://www.kaggle.com/bakar31\" target=\"_blank\">@bakar31</a> I am assuming by 0.7 you mean 0.7 Gb file size, which is too small on its own to create any issues for the RAM (again assuming you still get the 'exceeded allowed compute error'). The only thing I can think of, looking at the code you have shared, is the conversion of uint8 to a numpy array. This will convert the image data to float, which usually increases memory usage, especially if both instances remain in memory. I'm not sure if the increase due to conversion to float for a file up to 0.7 Gb would cause this problem (and I'm assuming you have checked with training data upto 0.7 Gb). </p>\n<p>Please make sure to check with training data with specific image sizes of around 0.7 Gb or even slightly higher. And if possible, check without the img_to_array once. Do these on training data and make sure, I guess you have 1 or 2 submissions left. I cannot think of anything else if training data works fine upto and around you file size limit (which to be honest is way too low for getting this error). </p>\n<p>PS. Hope you are not creating a batch of image arrays for prediction… Just clarifying, even though it might be slower, try one image inference at a time, with gc.collect(). <a href=\"https://www.kaggle.com/code/realneuralnetwork/cnn-strip-ai-inference/notebook?scriptVersionId=100707854\" target=\"_blank\">Here</a> is a full notebook that has inspired most of the fixes. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1901526,
      "author_name": "Thush_Thusinthaka",
      "author_url": "",
      "post_date": "2022-08-16T17:53:04.693000",
      "content": "<p>I am still getting \"Notebook throw Exception\"  on submite. That is after images were re-sized and zoomed-off. We don't have a issue with reading image and the Image GB value. Similar to your earlier error but no LOG for submitted notebook as the Version-Notebooks are our local one's. <br>\n1) What is the direct method to see Submitted-notedbook LOG page ? (Log error if we have try: except:)<br>\n2) What is the direct method to see our submission page route paths ?</p>\n<p>Thanks <a href=\"https://www.kaggle.com/icemantd\" target=\"_blank\">@icemantd</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1901570,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-08-16T18:55:50.977000",
          "content": "<p>There is no way to see the log output of the notebooks submitted for scoring… you have to figure out from <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">this</a> limited information and try to debug your code</p>\n<p>'Notebook threw exception' in my case has been due to some error in handling all possible failure cases. For e.g., you might skip an image based on try: except: but do not add the scoring for that to you submission data frame, so while compiling the score into submission.csv, there may be errors. Things like that need to be debugged, but exact error cannot be seen unfortunately.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1898755": "I have a novice question: \n\n**Is the submission notebook inference going through entire test data at this time, but providing public scores only on 7% of that prediction?**\n\nI ask because my submission currently takes 1.5 hours to successfully submit (when it runs fine that is). From a few other discussions I see that the current 7% test data is around 20 images, and the full test data has 280 images. If 20 images is taking 1.5 hours, I believe final submission will not be possible due to 9 hour notebook run time limit. Plus, on training hold out data, my inference completes in 25 min for about 50 images, so what's going on while submission scoring?\n\nMain idea is to clarify if our current notebooks should be completing inference in under 40 min if they are only running on 7% data and  are to successfully complete inference on 100% data for final submission in under 9 hours.\n\nThanks.",
    "1949101": "Hello,\nDid you ever find out if the submission happens for the 4 patients in the test set shared for us or the entire hidden test set? I am having trouble understanding it. I can do inference for the shared test set in 10 minutes however once I submit the notebook it either times out or throws a notebook exception.",
    "1901526": "I am still getting \"Notebook throw Exception\"  on submite. That is after images were re-sized and zoomed-off. We don't have a issue with reading image and the Image GB value. Similar to your earlier error but no LOG for submitted notebook as the Version-Notebooks are our local one's. \n1) What is the direct method to see Submitted-notedbook LOG page ? (Log error if we have try: except:)\n2) What is the direct method to see our submission page route paths ?\n\nThanks @icemantd \n\n"
  }
}