{
  "id": 388347,
  "title": "Looking to understand why my Notebook isn't getting scored?",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/388347",
  "author_name": "Conweezy",
  "post_date": "2023-02-17T01:03:05.200000",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I've spent quite sometime trying to understand why when I submit my notebook it runs successfully, but does not get scored. I'm getting the \"Notebook Threw Exception\" error. I've spent some time reading the troubleshooting for this error and looking at other people's notebooks but haven't been able to resolve it. I understand the hidden test set is larger than the visible test set, but I believe I've handled that appropriately. </p>\n<p>Can anyone take a look at my notebook and see where I'm going wrong? It's not too long.</p>\n<p><a href=\"https://www.kaggle.com/code/conweezy/rsna-submission-test?scriptVersionId=119413061\" target=\"_blank\">https://www.kaggle.com/code/conweezy/rsna-submission-test?scriptVersionId=119413061</a></p>\n<p>Thank you very much.</p>",
  "messages": [
    {
      "id": 2147893,
      "postDate": "2023-02-17T01:03:05.200Z",
      "content": "<p>Hello,</p>\n<p>I've spent quite sometime trying to understand why when I submit my notebook it runs successfully, but does not get scored. I'm getting the \"Notebook Threw Exception\" error. I've spent some time reading the troubleshooting for this error and looking at other people's notebooks but haven't been able to resolve it. I understand the hidden test set is larger than the visible test set, but I believe I've handled that appropriately. </p>\n<p>Can anyone take a look at my notebook and see where I'm going wrong? It's not too long.</p>\n<p><a href=\"https://www.kaggle.com/code/conweezy/rsna-submission-test?scriptVersionId=119413061\" target=\"_blank\">https://www.kaggle.com/code/conweezy/rsna-submission-test?scriptVersionId=119413061</a></p>\n<p>Thank you very much.</p>",
      "rawMarkdown": "Hello,\n\nI've spent quite sometime trying to understand why when I submit my notebook it runs successfully, but does not get scored. I'm getting the \"Notebook Threw Exception\" error. I've spent some time reading the troubleshooting for this error and looking at other people's notebooks but haven't been able to resolve it. I understand the hidden test set is larger than the visible test set, but I believe I've handled that appropriately. \n\nCan anyone take a look at my notebook and see where I'm going wrong? It's not too long.\n\nhttps://www.kaggle.com/code/conweezy/rsna-submission-test?scriptVersionId=119413061\n\nThank you very much.",
      "votes": 2
    },
    {
      "id": 2150003,
      "postDate": "2023-02-18T21:15:02.860Z",
      "content": "<p>After some trial and error, I've narrowed it down to the below lines of code that returns the exception.</p>\n<p>X_test = []<br>\nfor i in range(test_df.shape[0]):<br>\n    path = test_df['path'].iloc[i]<br>\n    xray = read_xray(path, 256)<br>\n    X_test.append(xray)</p>",
      "rawMarkdown": "After some trial and error, I've narrowed it down to the below lines of code that returns the exception.\n\n\n\nX_test = []\nfor i in range(test_df.shape[0]):\n    path = test_df['path'].iloc[i]\n    xray = read_xray(path, 256)\n    X_test.append(xray)"
    },
    {
      "id": 2147914,
      "postDate": "2023-02-17T02:25:11.230Z",
      "content": "<ol>\n<li>Seeing your notebook, I find out the format of a CSV that you submitted might be wrong. The final result has to be 0 or 1. (doesn't matter float or int type) In order to do that, make your own threshold. <br>\n<br></li>\n<li>In your code <code>probabilities = [F.softmax(t, dim=1) for t in predictions]</code><br>\nI think that it is not appropriate activation function in this case. Our mission is binary classification, not multiclass. Using the Softmax is a better fit for the multiclass. So you should Sigmoid function.</li>\n</ol>",
      "rawMarkdown": "1. Seeing your notebook, I find out the format of a CSV that you submitted might be wrong. The final result has to be 0 or 1. (doesn't matter float or int type) In order to do that, make your own threshold. \n</br>\n2. In your code `probabilities = [F.softmax(t, dim=1) for t in predictions]`\nI think that it is not appropriate activation function in this case. Our mission is binary classification, not multiclass. Using the Softmax is a better fit for the multiclass. So you should Sigmoid function.\n",
      "replies": [
        {
          "id": 2147971,
          "postDate": "2023-02-17T04:12:24.493Z",
          "content": "<p>Thanks for the reply. I will definitely try Sigmoid. But don't Softmax and Sigmoid provide outputs between 0 and 1? I get sigmoid is better, but I didn't see any of my prediction values in my Submission csv that weren't between 0 and 1. Where did you see this?</p>",
          "rawMarkdown": "Thanks for the reply. I will definitely try Sigmoid. But don't Softmax and Sigmoid provide outputs between 0 and 1? I get sigmoid is better, but I didn't see any of my prediction values in my Submission csv that weren't between 0 and 1. Where did you see this?",
          "replies": [
            {
              "id": 2148034,
              "postDate": "2023-02-17T05:19:19.553Z",
              "content": "<p>Yes, both of them produce outputs between 0 and 1. However, If you try to do that, you can find that there are some differences to the values of a range, which can affect your decision to set the range of threshold.</p>\n<p>Your notebook that shares this post are showing the values (0.35..). Check out the last section of your notebook and your submission file in the Data Column.</p>",
              "rawMarkdown": "Yes, both of them produce outputs between 0 and 1. However, If you try to do that, you can find that there are some differences to the values of a range, which can affect your decision to set the range of threshold.\n\nYour notebook that shares this post are showing the values (0.35..). Check out the last section of your notebook and your submission file in the Data Column.\n"
            }
          ]
        },
        {
          "id": 2148820,
          "postDate": "2023-02-17T17:36:29.230Z",
          "content": "<p><a href=\"https://www.kaggle.com/hyunsoolee1010\" target=\"_blank\">@hyunsoolee1010</a> On the competition main page it says that</p>\n<pre><code>prediction_id,cancer\n-L,\n-R,\n-L,\n...\n</code></pre>\n<p>so perhaps to clarify your answer it is good to elaborate that the submitted predictions may range from 0 to 1.</p>",
          "rawMarkdown": "@hyunsoolee1010 On the competition main page it says that\n```python\nprediction_id,cancer\n0-L,0\n0-R,0.5\n1-L,1\n...\n```\nso perhaps to clarify your answer it is good to elaborate that the submitted predictions may range from 0 to 1.",
          "votes": 1,
          "replies": [
            {
              "id": 2148882,
              "postDate": "2023-02-17T18:22:25.200Z",
              "content": "<p><a href=\"https://www.kaggle.com/anttiisosalo\" target=\"_blank\">@anttiisosalo</a> thank you for adding it. My explanation was wrong in terms of the range. <a href=\"https://www.kaggle.com/conweezy\" target=\"_blank\">@conweezy</a> check this plz.</p>",
              "rawMarkdown": "@anttiisosalo thank you for adding it. My explanation was wrong in terms of the range. @conweezy check this plz."
            },
            {
              "id": 2148915,
              "postDate": "2023-02-17T18:33:31.940Z",
              "content": "<p>Thank you both for the feedback. As far as I can tell, my submissions have always been between 0 and 1. Any other ideas why it's not getting scored.</p>\n<p>I may try adding a threshold so that my predictions are either 0 or 1, but I don't think that is what is causing the issue.</p>\n<p>Thanks.</p>",
              "rawMarkdown": "Thank you both for the feedback. As far as I can tell, my submissions have always been between 0 and 1. Any other ideas why it's not getting scored.\n\nI may try adding a threshold so that my predictions are either 0 or 1, but I don't think that is what is causing the issue.\n\nThanks."
            },
            {
              "id": 2149159,
              "postDate": "2023-02-18T02:34:59.183Z",
              "content": "<p>In addition, if you are trying to add the threshold, the value of a threshold must be an appropriate value. In my case, like you, the inference notebook was successful, but the process of submission didn't. My submission had a slightly different value for a threshold. Some notebook was passed well. But, others didn't. I guess that the value of a threshold can make our model delayed. </p>\n<p>Therefore, the boundary that is made from the threshold in data includes too much data to process the notebook in a limited time. To deal with this, I'll have some experiments to adjust the value of logits.</p>",
              "rawMarkdown": "In addition, if you are trying to add the threshold, the value of a threshold must be an appropriate value. In my case, like you, the inference notebook was successful, but the process of submission didn't. My submission had a slightly different value for a threshold. Some notebook was passed well. But, others didn't. I guess that the value of a threshold can make our model delayed. \n\nTherefore, the boundary that is made from the threshold in data includes too much data to process the notebook in a limited time. To deal with this, I'll have some experiments to adjust the value of logits."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2150003,
      "author_name": "Conweezy",
      "author_url": "",
      "post_date": "2023-02-18T21:15:02.860000",
      "content": "<p>After some trial and error, I've narrowed it down to the below lines of code that returns the exception.</p>\n<p>X_test = []<br>\nfor i in range(test_df.shape[0]):<br>\n    path = test_df['path'].iloc[i]<br>\n    xray = read_xray(path, 256)<br>\n    X_test.append(xray)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2147914,
      "author_name": "Hyunsoo Lee 1010",
      "author_url": "",
      "post_date": "2023-02-17T02:25:11.230000",
      "content": "<ol>\n<li>Seeing your notebook, I find out the format of a CSV that you submitted might be wrong. The final result has to be 0 or 1. (doesn't matter float or int type) In order to do that, make your own threshold. <br>\n<br></li>\n<li>In your code <code>probabilities = [F.softmax(t, dim=1) for t in predictions]</code><br>\nI think that it is not appropriate activation function in this case. Our mission is binary classification, not multiclass. Using the Softmax is a better fit for the multiclass. So you should Sigmoid function.</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 2147971,
          "author_name": "Conweezy",
          "author_url": "",
          "post_date": "2023-02-17T04:12:24.493000",
          "content": "<p>Thanks for the reply. I will definitely try Sigmoid. But don't Softmax and Sigmoid provide outputs between 0 and 1? I get sigmoid is better, but I didn't see any of my prediction values in my Submission csv that weren't between 0 and 1. Where did you see this?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2148034,
              "author_name": "Hyunsoo Lee 1010",
              "author_url": "",
              "post_date": "2023-02-17T05:19:19.553000",
              "content": "<p>Yes, both of them produce outputs between 0 and 1. However, If you try to do that, you can find that there are some differences to the values of a range, which can affect your decision to set the range of threshold.</p>\n<p>Your notebook that shares this post are showing the values (0.35..). Check out the last section of your notebook and your submission file in the Data Column.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2148820,
          "author_name": "Antti Isosalo",
          "author_url": "",
          "post_date": "2023-02-17T17:36:29.230000",
          "content": "<p><a href=\"https://www.kaggle.com/hyunsoolee1010\" target=\"_blank\">@hyunsoolee1010</a> On the competition main page it says that</p>\n<pre><code>prediction_id,cancer\n-L,\n-R,\n-L,\n...\n</code></pre>\n<p>so perhaps to clarify your answer it is good to elaborate that the submitted predictions may range from 0 to 1.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2148882,
              "author_name": "Hyunsoo Lee 1010",
              "author_url": "",
              "post_date": "2023-02-17T18:22:25.200000",
              "content": "<p><a href=\"https://www.kaggle.com/anttiisosalo\" target=\"_blank\">@anttiisosalo</a> thank you for adding it. My explanation was wrong in terms of the range. <a href=\"https://www.kaggle.com/conweezy\" target=\"_blank\">@conweezy</a> check this plz.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2148915,
              "author_name": "Conweezy",
              "author_url": "",
              "post_date": "2023-02-17T18:33:31.940000",
              "content": "<p>Thank you both for the feedback. As far as I can tell, my submissions have always been between 0 and 1. Any other ideas why it's not getting scored.</p>\n<p>I may try adding a threshold so that my predictions are either 0 or 1, but I don't think that is what is causing the issue.</p>\n<p>Thanks.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2149159,
              "author_name": "Hyunsoo Lee 1010",
              "author_url": "",
              "post_date": "2023-02-18T02:34:59.183000",
              "content": "<p>In addition, if you are trying to add the threshold, the value of a threshold must be an appropriate value. In my case, like you, the inference notebook was successful, but the process of submission didn't. My submission had a slightly different value for a threshold. Some notebook was passed well. But, others didn't. I guess that the value of a threshold can make our model delayed. </p>\n<p>Therefore, the boundary that is made from the threshold in data includes too much data to process the notebook in a limited time. To deal with this, I'll have some experiments to adjust the value of logits.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2147893": "Hello,\n\nI've spent quite sometime trying to understand why when I submit my notebook it runs successfully, but does not get scored. I'm getting the \"Notebook Threw Exception\" error. I've spent some time reading the troubleshooting for this error and looking at other people's notebooks but haven't been able to resolve it. I understand the hidden test set is larger than the visible test set, but I believe I've handled that appropriately. \n\nCan anyone take a look at my notebook and see where I'm going wrong? It's not too long.\n\nhttps://www.kaggle.com/code/conweezy/rsna-submission-test?scriptVersionId=119413061\n\nThank you very much.",
    "2150003": "After some trial and error, I've narrowed it down to the below lines of code that returns the exception.\n\n\n\nX_test = []\nfor i in range(test_df.shape[0]):\n    path = test_df['path'].iloc[i]\n    xray = read_xray(path, 256)\n    X_test.append(xray)",
    "2147914": "1. Seeing your notebook, I find out the format of a CSV that you submitted might be wrong. The final result has to be 0 or 1. (doesn't matter float or int type) In order to do that, make your own threshold. \n</br>\n2. In your code `probabilities = [F.softmax(t, dim=1) for t in predictions]`\nI think that it is not appropriate activation function in this case. Our mission is binary classification, not multiclass. Using the Softmax is a better fit for the multiclass. So you should Sigmoid function.\n"
  }
}