{
  "id": 434295,
  "title": "Submission Error",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/434295",
  "author_name": "Matheus Oliveira de Souza",
  "post_date": "2023-08-24T17:31:04.027000",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I am getting \"Notebook Threw Exception\" error when i try to submit my solution, i don't know how to solve this problem. Has anyone ever received this error before?</p>",
  "messages": [
    {
      "id": 2408591,
      "postDate": "2023-08-25T18:00:45.947Z",
      "content": "<p>I forked your notebook.  The notebook runs fine in interactive mode.  When a submission is made it does fail.</p>\n<p>I have a dataset I created for 'test' that uses the first 7 patients from train rather than the silly little test set that we have been given to play with.   </p>\n<p>I forked your notebook and changed the location of the test and sample_submission to my dataset.  Running in interactive mode it fails with an error I don't quite understand yet - will play more later.</p>\n<h2>Error =</h2>\n<p>KeyError                                  Traceback (most recent call last)<br>\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3802, in Index.get_loc(self, key, method, tolerance)<br>\n   3801 try:<br>\n-&gt; 3802     return self._engine.get_loc(casted_key)<br>\n   3803 except KeyError as err:</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:138, in pandas._libs.index.IndexEngine.get_loc()</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:165, in pandas._libs.index.IndexEngine.get_loc()</p>\n<p>File pandas/_libs/hashtable_class_helper.pxi:5745, in pandas._libs.hashtable.PyObjectHashTable.get_item()</p>\n<p>File pandas/_libs/hashtable_class_helper.pxi:5753, in pandas._libs.hashtable.PyObjectHashTable.get_item()</p>\n<p>KeyError: 'bowel_injury'</p>\n<p>The above exception was the direct cause of the following exception:</p>\n<p>KeyError                                  Traceback (most recent call last)<br>\nCell In[8], line 24<br>\n     21         torch.cuda.empty_cache()<br>\n     23 my_submission = pd.DataFrame(data=submission)<br>\n---&gt; 24 my_submission = create_training_solution(y_train=my_submission)<br>\n     25 id_col = my_submission['patient_id'].copy()</p>\n<p>Cell In[2], line 5, in create_training_solution(y_train)<br>\n      2 sol_train = y_train.copy()<br>\n      4 # bowel healthy|injury sample weight = 1|2<br>\n----&gt; 5 sol_train['bowel_weight'] = np.where(sol_train['bowel_injury'] == 1, 2, 1)<br>\n      7 # extravasation healthy/injury sample weight = 1|6<br>\n      8 sol_train['extravasation_weight'] = np.where(sol_train['extravasation_injury'] == 1, 6, 1)</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/core/frame.py:3807, in DataFrame.<strong>getitem</strong>(self, key)<br>\n   3805 if self.columns.nlevels &gt; 1:<br>\n   3806     return self._getitem_multilevel(key)<br>\n-&gt; 3807 indexer = self.columns.get_loc(key)<br>\n   3808 if is_integer(indexer):<br>\n   3809     indexer = [indexer]</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3804, in Index.get_loc(self, key, method, tolerance)<br>\n   3802     return self._engine.get_loc(casted_key)<br>\n   3803 except KeyError as err:<br>\n-&gt; 3804     raise KeyError(key) from err<br>\n   3805 except TypeError:<br>\n   3806     # If we have a listlike key, _check_indexing_error will raise<br>\n   3807     #  InvalidIndexError. Otherwise we fall through and re-raise<br>\n   3808     #  the TypeError.<br>\n   3809     self._check_indexing_error(key)</p>\n<p>KeyError: 'bowel_injury'</p>\n<p>Just guessing that 'create_training_solution' section works fine when only one slice per patient (like our silly test) but fails when a more realistic test set is present.</p>\n<p><a href=\"https://www.kaggle.com/code/pcjimmmy/rsna-submission-pc-jimmmy\" target=\"_blank\">Forked notebook</a></p>\n<p><a href=\"https://www.kaggle.com/datasets/pcjimmmy/rsna-7-patients-test\" target=\"_blank\">My test dataset</a></p>\n<p>Let me know if this helped!</p>",
      "rawMarkdown": "I forked your notebook.  The notebook runs fine in interactive mode.  When a submission is made it does fail.\n\nI have a dataset I created for 'test' that uses the first 7 patients from train rather than the silly little test set that we have been given to play with.   \n\nI forked your notebook and changed the location of the test and sample_submission to my dataset.  Running in interactive mode it fails with an error I don't quite understand yet - will play more later.\n\nError =\n---------------------------------------------------------------------------\nKeyError                                  Traceback (most recent call last)\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3802, in Index.get_loc(self, key, method, tolerance)\n   3801 try:\n-> 3802     return self._engine.get_loc(casted_key)\n   3803 except KeyError as err:\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:138, in pandas._libs.index.IndexEngine.get_loc()\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:165, in pandas._libs.index.IndexEngine.get_loc()\n\nFile pandas/_libs/hashtable_class_helper.pxi:5745, in pandas._libs.hashtable.PyObjectHashTable.get_item()\n\nFile pandas/_libs/hashtable_class_helper.pxi:5753, in pandas._libs.hashtable.PyObjectHashTable.get_item()\n\nKeyError: 'bowel_injury'\n\nThe above exception was the direct cause of the following exception:\n\nKeyError                                  Traceback (most recent call last)\nCell In[8], line 24\n     21         torch.cuda.empty_cache()\n     23 my_submission = pd.DataFrame(data=submission)\n---> 24 my_submission = create_training_solution(y_train=my_submission)\n     25 id_col = my_submission['patient_id'].copy()\n\nCell In[2], line 5, in create_training_solution(y_train)\n      2 sol_train = y_train.copy()\n      4 # bowel healthy|injury sample weight = 1|2\n----> 5 sol_train['bowel_weight'] = np.where(sol_train['bowel_injury'] == 1, 2, 1)\n      7 # extravasation healthy/injury sample weight = 1|6\n      8 sol_train['extravasation_weight'] = np.where(sol_train['extravasation_injury'] == 1, 6, 1)\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/frame.py:3807, in DataFrame.__getitem__(self, key)\n   3805 if self.columns.nlevels > 1:\n   3806     return self._getitem_multilevel(key)\n-> 3807 indexer = self.columns.get_loc(key)\n   3808 if is_integer(indexer):\n   3809     indexer = [indexer]\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3804, in Index.get_loc(self, key, method, tolerance)\n   3802     return self._engine.get_loc(casted_key)\n   3803 except KeyError as err:\n-> 3804     raise KeyError(key) from err\n   3805 except TypeError:\n   3806     # If we have a listlike key, _check_indexing_error will raise\n   3807     #  InvalidIndexError. Otherwise we fall through and re-raise\n   3808     #  the TypeError.\n   3809     self._check_indexing_error(key)\n\nKeyError: 'bowel_injury'\n\nJust guessing that 'create_training_solution' section works fine when only one slice per patient (like our silly test) but fails when a more realistic test set is present.\n\n[Forked notebook](https://www.kaggle.com/code/pcjimmmy/rsna-submission-pc-jimmmy)\n\n[My test dataset](https://www.kaggle.com/datasets/pcjimmmy/rsna-7-patients-test)\n\nLet me know if this helped!",
      "votes": 3,
      "replies": [
        {
          "id": 2408653,
          "postDate": "2023-08-25T18:32:53.087Z",
          "content": "<p>I found an issue that at least let things run in interactive mode.  </p>\n<p>You read the sample_submission.csv file in the cell above my error, but than never use it.</p>\n<p>Should this be a correct line in the cell</p>\n<p>'my_submission = pd.DataFrame(data=submission_df)'</p>\n<p>I did a save of my fork - will see if it runs and than revise back to the original test folders.</p>",
          "rawMarkdown": "I found an issue that at least let things run in interactive mode.  \n\nYou read the sample_submission.csv file in the cell above my error, but than never use it.\n\nShould this be a correct line in the cell\n\n\n'my_submission = pd.DataFrame(data=submission_df)'\n\nI did a save of my fork - will see if it runs and than revise back to the original test folders.",
          "replies": [
            {
              "id": 2408664,
              "postDate": "2023-08-25T18:38:55.607Z",
              "content": "<p>OK - got a different error after making the above change but its similar to the first one I got.</p>\n<p>The fake 'sample_submission' file I created in my dataset does include a column 'any_injury' since I pretty much copy and paste from training.  So I think you need to add that column to the 'submission_df' .  But on this error path I am working I don't understand how you code will run in interactive with the simple test set ????   Anywlay - off to do other stuff for a few hours - </p>\n<hr>\n<p>KeyError                                  Traceback (most recent call last)<br>\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3802, in Index.get_loc(self, key, method, tolerance)<br>\n   3801 try:<br>\n-&gt; 3802     return self._engine.get_loc(casted_key)<br>\n   3803 except KeyError as err:</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:138, in pandas._libs.index.IndexEngine.get_loc()</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:165, in pandas._libs.index.IndexEngine.get_loc()</p>\n<p>File pandas/_libs/hashtable_class_helper.pxi:5745, in pandas._libs.hashtable.PyObjectHashTable.get_item()</p>\n<p>File pandas/_libs/hashtable_class_helper.pxi:5753, in pandas._libs.hashtable.PyObjectHashTable.get_item()</p>\n<p>KeyError: 'any_injury'</p>\n<p>The above exception was the direct cause of the following exception:</p>\n<p>KeyError                                  Traceback (most recent call last)<br>\nCell In[8], line 26<br>\n     23 my_submission = pd.DataFrame(data=submission_df)<br>\n     24 print(my_submission.columns)<br>\n---&gt; 26 my_submission = create_training_solution(y_train=my_submission)<br>\n     27 id_col = my_submission['patient_id'].copy()</p>\n<p>Cell In[2], line 20, in create_training_solution(y_train)<br>\n     17 sol_train['spleen_weight'] = np.where(sol_train['spleen_low'] == 1, 2, np.where(sol_train['spleen_high'] == 1, 4, 1))<br>\n     19 # any healthy|injury sample weight = 1|6<br>\n---&gt; 20 sol_train['any_injury_weight'] = np.where(sol_train['any_injury'] == 1, 6, 1)<br>\n     22 return sol_train</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/core/frame.py:3807, in DataFrame.<strong>getitem</strong>(self, key)<br>\n   3805 if self.columns.nlevels &gt; 1:<br>\n   3806     return self._getitem_multilevel(key)<br>\n-&gt; 3807 indexer = self.columns.get_loc(key)<br>\n   3808 if is_integer(indexer):<br>\n   3809     indexer = [indexer]</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3804, in Index.get_loc(self, key, method, tolerance)<br>\n   3802     return self._engine.get_loc(casted_key)<br>\n   3803 except KeyError as err:<br>\n-&gt; 3804     raise KeyError(key) from err<br>\n   3805 except TypeError:<br>\n   3806     # If we have a listlike key, _check_indexing_error will raise<br>\n   3807     #  InvalidIndexError. Otherwise we fall through and re-raise<br>\n   3808     #  the TypeError.<br>\n   3809     self._check_indexing_error(key)</p>\n<p>KeyError: 'any_injury'</p>",
              "rawMarkdown": "OK - got a different error after making the above change but its similar to the first one I got.\n\nThe fake 'sample_submission' file I created in my dataset does include a column 'any_injury' since I pretty much copy and paste from training.  So I think you need to add that column to the 'submission_df' .  But on this error path I am working I don't understand how you code will run in interactive with the simple test set ????   Anywlay - off to do other stuff for a few hours - \n\n---------------------------------------------------------------------------\nKeyError                                  Traceback (most recent call last)\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3802, in Index.get_loc(self, key, method, tolerance)\n   3801 try:\n-> 3802     return self._engine.get_loc(casted_key)\n   3803 except KeyError as err:\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:138, in pandas._libs.index.IndexEngine.get_loc()\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:165, in pandas._libs.index.IndexEngine.get_loc()\n\nFile pandas/_libs/hashtable_class_helper.pxi:5745, in pandas._libs.hashtable.PyObjectHashTable.get_item()\n\nFile pandas/_libs/hashtable_class_helper.pxi:5753, in pandas._libs.hashtable.PyObjectHashTable.get_item()\n\nKeyError: 'any_injury'\n\nThe above exception was the direct cause of the following exception:\n\nKeyError                                  Traceback (most recent call last)\nCell In[8], line 26\n     23 my_submission = pd.DataFrame(data=submission_df)\n     24 print(my_submission.columns)\n---> 26 my_submission = create_training_solution(y_train=my_submission)\n     27 id_col = my_submission['patient_id'].copy()\n\nCell In[2], line 20, in create_training_solution(y_train)\n     17 sol_train['spleen_weight'] = np.where(sol_train['spleen_low'] == 1, 2, np.where(sol_train['spleen_high'] == 1, 4, 1))\n     19 # any healthy|injury sample weight = 1|6\n---> 20 sol_train['any_injury_weight'] = np.where(sol_train['any_injury'] == 1, 6, 1)\n     22 return sol_train\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/frame.py:3807, in DataFrame.__getitem__(self, key)\n   3805 if self.columns.nlevels > 1:\n   3806     return self._getitem_multilevel(key)\n-> 3807 indexer = self.columns.get_loc(key)\n   3808 if is_integer(indexer):\n   3809     indexer = [indexer]\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3804, in Index.get_loc(self, key, method, tolerance)\n   3802     return self._engine.get_loc(casted_key)\n   3803 except KeyError as err:\n-> 3804     raise KeyError(key) from err\n   3805 except TypeError:\n   3806     # If we have a listlike key, _check_indexing_error will raise\n   3807     #  InvalidIndexError. Otherwise we fall through and re-raise\n   3808     #  the TypeError.\n   3809     self._check_indexing_error(key)\n\nKeyError: 'any_injury'"
            }
          ]
        }
      ]
    },
    {
      "id": 2407157,
      "postDate": "2023-08-24T21:38:31.817Z",
      "content": "<p>Some files are corrupted or have the wrong extension, </p>\n<pre><code> () -&gt; np.ndarray:\n     dcmread(fp=file_path)  dcm_file:\n         dcm_file.pixel_array\n</code></pre>\n<p>What happen if dcmread fails at reading the file?</p>\n<pre><code> () -&gt; np.ndarray:\n    :\n         dcmread(fp=file_path)  dcm_file:\n         dcm_file.pixel_array\n     Exception  e:\n        ()\n         \n</code></pre>",
      "rawMarkdown": "Some files are corrupted or have the wrong extension, \n\n```python\ndef load_dicom(file_path: str | Path) -> np.ndarray:\n    with dcmread(fp=file_path) as dcm_file:\n        return dcm_file.pixel_array\n```\n\n\nWhat happen if dcmread fails at reading the file?\n\n```python\ndef load_dicom(file_path: str | Path) -> np.ndarray:\n    try:\n        with dcmread(fp=file_path) as dcm_file:\n        return dcm_file.pixel_array\n    except Exception as e:\n        print(f'Error {e} with file {file_path}')\n        return 1\n```",
      "votes": 1,
      "replies": [
        {
          "id": 2408295,
          "postDate": "2023-08-25T14:53:15.833Z",
          "content": "<p>Thanks for helping, but i tried it and the submition is still not working.</p>",
          "rawMarkdown": "Thanks for helping, but i tried it and the submition is still not working.",
          "replies": [
            {
              "id": 2408433,
              "postDate": "2023-08-25T16:17:36.070Z",
              "content": "<p>Hi,</p>\n<p>That was an example, if you return 1 the transform_data function will error too.</p>\n<p>This would probably work,</p>\n<pre><code> torch.no_grad():\n     test_file  load_test_set():\n        :\n            patient_id = (test_file.parent.parent.name)\n            x = transform_data(file_path=test_file).to(device=DEVICE)\n            yhat = model(x)\n            prob = torch.sigmoid(=yhat).detach().cpu().numpy().squeeze()\n\n            submission[].append(patient_id)\n            submission[].append(prob[-].item())\n             col, value  (target_columns, prob):\n                submission[col].append(value.item())\n\n             x, yhat, prob\n            torch.cuda.empty_cache()\n         Exception  e:\n            (, e)\n            \n</code></pre>\n<p>But would be better to refactor the code to handle the error at load_dicom() or transform_data()</p>\n<p>Edit:</p>\n<p>This fork work, just need to change the submission file part.</p>\n<p><a href=\"https://www.kaggle.com/enriquezaf/rsna-submission\" target=\"_blank\">https://www.kaggle.com/enriquezaf/rsna-submission</a></p>",
              "rawMarkdown": "Hi,\n\nThat was an example, if you return 1 the transform_data function will error too.\n\nThis would probably work,\n\n```python\nwith torch.no_grad():\n    for test_file in load_test_set():\n        try:\n            patient_id = int(test_file.parent.parent.name)\n            x = transform_data(file_path=test_file).to(device=DEVICE)\n            yhat = model(x)\n            prob = torch.sigmoid(input=yhat).detach().cpu().numpy().squeeze()\n\n            submission[\"patient_id\"].append(patient_id)\n            submission[\"any_injury\"].append(prob[-1].item())\n            for col, value in zip(target_columns, prob):\n                submission[col].append(value.item())\n\n            del x, yhat, prob\n            torch.cuda.empty_cache()\n        except Exception as e:\n            print(\"Error: \", e)\n            continue\n```\n\nBut would be better to refactor the code to handle the error at load_dicom() or transform_data()\n\nEdit:\n\nThis fork work, just need to change the submission file part.\n\nhttps://www.kaggle.com/enriquezaf/rsna-submission"
            }
          ]
        }
      ]
    },
    {
      "id": 2406869,
      "postDate": "2023-08-24T17:31:04.027Z",
      "content": "<p>I am getting \"Notebook Threw Exception\" error when i try to submit my solution, i don't know how to solve this problem. Has anyone ever received this error before?</p>",
      "rawMarkdown": "I am getting \"Notebook Threw Exception\" error when i try to submit my solution, i don't know how to solve this problem. Has anyone ever received this error before?",
      "votes": 1
    },
    {
      "id": 2406928,
      "postDate": "2023-08-24T17:55:33.510Z",
      "content": "<p>This is a super common error that folks get when first learning to submit notebooks.  Search 'submission' in the search.  </p>\n<p>For help on your notebook - share the notebook (make it public) and put a link to the notebook in this post.  Make any datasets with the model public.  A decent number of folks will be able to fork and find the obvious errors that are causing the exception.   Memory error is the most common for this error type.</p>",
      "rawMarkdown": "This is a super common error that folks get when first learning to submit notebooks.  Search 'submission' in the search.  \n\nFor help on your notebook - share the notebook (make it public) and put a link to the notebook in this post.  Make any datasets with the model public.  A decent number of folks will be able to fork and find the obvious errors that are causing the exception.   Memory error is the most common for this error type.",
      "replies": [
        {
          "id": 2407054,
          "postDate": "2023-08-24T19:07:18.227Z",
          "content": "<p>Here is my notebook: <a href=\"https://www.kaggle.com/code/matheusmt/rsna-submission/notebook\" target=\"_blank\">https://www.kaggle.com/code/matheusmt/rsna-submission/notebook</a></p>",
          "rawMarkdown": "Here is my notebook: https://www.kaggle.com/code/matheusmt/rsna-submission/notebook",
          "replies": [
            {
              "id": 2407492,
              "postDate": "2023-08-25T05:38:04.103Z",
              "content": "<p>There is probably nothing wrong with your notebook. I read your notebook and you are loading all .dcm files with glob, but some .dcm files are corrupt which throws an error. I created a discussion regarding this problem, and some notebooks that does nothing except for loading the .pixel_array from the .dcm files, which throws an error during submission time.</p>\n<p>You can use <a href=\"https://www.kaggle.com/enriquezaf\" target=\"_blank\">@enriquezaf</a> 's solution to not load the corrupt .dcm files from the dataset.</p>\n<p>Reference: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/432593\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/432593</a></p>",
              "rawMarkdown": "There is probably nothing wrong with your notebook. I read your notebook and you are loading all .dcm files with glob, but some .dcm files are corrupt which throws an error. I created a discussion regarding this problem, and some notebooks that does nothing except for loading the .pixel_array from the .dcm files, which throws an error during submission time.\n\nYou can use @enriquezaf 's solution to not load the corrupt .dcm files from the dataset.\n\nReference: https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/432593",
              "votes": 1
            },
            {
              "id": 2408305,
              "postDate": "2023-08-25T14:59:00.377Z",
              "content": "<p>Thanks for helping, i basically copied the entire code that loads the data from <a href=\"https://www.kaggle.com/code/aritrag/kerascv-starter-notebook-infer/notebook\" target=\"_blank\">https://www.kaggle.com/code/aritrag/kerascv-starter-notebook-infer/notebook</a> and i am still getting the same error, i tried \"try and except\" as well but my submission is still failing so i dont know what to do anymore</p>",
              "rawMarkdown": "Thanks for helping, i basically copied the entire code that loads the data from https://www.kaggle.com/code/aritrag/kerascv-starter-notebook-infer/notebook and i am still getting the same error, i tried \"try and except\" as well but my submission is still failing so i dont know what to do anymore"
            },
            {
              "id": 2409470,
              "postDate": "2023-08-26T09:03:18.630Z",
              "content": "<p>I encountered a similar problem and when I tried to make predictions on the training set I found there was GPU memory leak or something. OOM problems… I'm still trying to figure out the problem with GPU memory.</p>",
              "rawMarkdown": "I encountered a similar problem and when I tried to make predictions on the training set I found there was GPU memory leak or something. OOM problems... I'm still trying to figure out the problem with GPU memory.",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2408591,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2023-08-25T18:00:45.947000",
      "content": "<p>I forked your notebook.  The notebook runs fine in interactive mode.  When a submission is made it does fail.</p>\n<p>I have a dataset I created for 'test' that uses the first 7 patients from train rather than the silly little test set that we have been given to play with.   </p>\n<p>I forked your notebook and changed the location of the test and sample_submission to my dataset.  Running in interactive mode it fails with an error I don't quite understand yet - will play more later.</p>\n<h2>Error =</h2>\n<p>KeyError                                  Traceback (most recent call last)<br>\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3802, in Index.get_loc(self, key, method, tolerance)<br>\n   3801 try:<br>\n-&gt; 3802     return self._engine.get_loc(casted_key)<br>\n   3803 except KeyError as err:</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:138, in pandas._libs.index.IndexEngine.get_loc()</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:165, in pandas._libs.index.IndexEngine.get_loc()</p>\n<p>File pandas/_libs/hashtable_class_helper.pxi:5745, in pandas._libs.hashtable.PyObjectHashTable.get_item()</p>\n<p>File pandas/_libs/hashtable_class_helper.pxi:5753, in pandas._libs.hashtable.PyObjectHashTable.get_item()</p>\n<p>KeyError: 'bowel_injury'</p>\n<p>The above exception was the direct cause of the following exception:</p>\n<p>KeyError                                  Traceback (most recent call last)<br>\nCell In[8], line 24<br>\n     21         torch.cuda.empty_cache()<br>\n     23 my_submission = pd.DataFrame(data=submission)<br>\n---&gt; 24 my_submission = create_training_solution(y_train=my_submission)<br>\n     25 id_col = my_submission['patient_id'].copy()</p>\n<p>Cell In[2], line 5, in create_training_solution(y_train)<br>\n      2 sol_train = y_train.copy()<br>\n      4 # bowel healthy|injury sample weight = 1|2<br>\n----&gt; 5 sol_train['bowel_weight'] = np.where(sol_train['bowel_injury'] == 1, 2, 1)<br>\n      7 # extravasation healthy/injury sample weight = 1|6<br>\n      8 sol_train['extravasation_weight'] = np.where(sol_train['extravasation_injury'] == 1, 6, 1)</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/core/frame.py:3807, in DataFrame.<strong>getitem</strong>(self, key)<br>\n   3805 if self.columns.nlevels &gt; 1:<br>\n   3806     return self._getitem_multilevel(key)<br>\n-&gt; 3807 indexer = self.columns.get_loc(key)<br>\n   3808 if is_integer(indexer):<br>\n   3809     indexer = [indexer]</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3804, in Index.get_loc(self, key, method, tolerance)<br>\n   3802     return self._engine.get_loc(casted_key)<br>\n   3803 except KeyError as err:<br>\n-&gt; 3804     raise KeyError(key) from err<br>\n   3805 except TypeError:<br>\n   3806     # If we have a listlike key, _check_indexing_error will raise<br>\n   3807     #  InvalidIndexError. Otherwise we fall through and re-raise<br>\n   3808     #  the TypeError.<br>\n   3809     self._check_indexing_error(key)</p>\n<p>KeyError: 'bowel_injury'</p>\n<p>Just guessing that 'create_training_solution' section works fine when only one slice per patient (like our silly test) but fails when a more realistic test set is present.</p>\n<p><a href=\"https://www.kaggle.com/code/pcjimmmy/rsna-submission-pc-jimmmy\" target=\"_blank\">Forked notebook</a></p>\n<p><a href=\"https://www.kaggle.com/datasets/pcjimmmy/rsna-7-patients-test\" target=\"_blank\">My test dataset</a></p>\n<p>Let me know if this helped!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2408653,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2023-08-25T18:32:53.087000",
          "content": "<p>I found an issue that at least let things run in interactive mode.  </p>\n<p>You read the sample_submission.csv file in the cell above my error, but than never use it.</p>\n<p>Should this be a correct line in the cell</p>\n<p>'my_submission = pd.DataFrame(data=submission_df)'</p>\n<p>I did a save of my fork - will see if it runs and than revise back to the original test folders.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2408664,
              "author_name": "PC Jimmmy",
              "author_url": "",
              "post_date": "2023-08-25T18:38:55.607000",
              "content": "<p>OK - got a different error after making the above change but its similar to the first one I got.</p>\n<p>The fake 'sample_submission' file I created in my dataset does include a column 'any_injury' since I pretty much copy and paste from training.  So I think you need to add that column to the 'submission_df' .  But on this error path I am working I don't understand how you code will run in interactive with the simple test set ????   Anywlay - off to do other stuff for a few hours - </p>\n<hr>\n<p>KeyError                                  Traceback (most recent call last)<br>\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3802, in Index.get_loc(self, key, method, tolerance)<br>\n   3801 try:<br>\n-&gt; 3802     return self._engine.get_loc(casted_key)<br>\n   3803 except KeyError as err:</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:138, in pandas._libs.index.IndexEngine.get_loc()</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:165, in pandas._libs.index.IndexEngine.get_loc()</p>\n<p>File pandas/_libs/hashtable_class_helper.pxi:5745, in pandas._libs.hashtable.PyObjectHashTable.get_item()</p>\n<p>File pandas/_libs/hashtable_class_helper.pxi:5753, in pandas._libs.hashtable.PyObjectHashTable.get_item()</p>\n<p>KeyError: 'any_injury'</p>\n<p>The above exception was the direct cause of the following exception:</p>\n<p>KeyError                                  Traceback (most recent call last)<br>\nCell In[8], line 26<br>\n     23 my_submission = pd.DataFrame(data=submission_df)<br>\n     24 print(my_submission.columns)<br>\n---&gt; 26 my_submission = create_training_solution(y_train=my_submission)<br>\n     27 id_col = my_submission['patient_id'].copy()</p>\n<p>Cell In[2], line 20, in create_training_solution(y_train)<br>\n     17 sol_train['spleen_weight'] = np.where(sol_train['spleen_low'] == 1, 2, np.where(sol_train['spleen_high'] == 1, 4, 1))<br>\n     19 # any healthy|injury sample weight = 1|6<br>\n---&gt; 20 sol_train['any_injury_weight'] = np.where(sol_train['any_injury'] == 1, 6, 1)<br>\n     22 return sol_train</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/core/frame.py:3807, in DataFrame.<strong>getitem</strong>(self, key)<br>\n   3805 if self.columns.nlevels &gt; 1:<br>\n   3806     return self._getitem_multilevel(key)<br>\n-&gt; 3807 indexer = self.columns.get_loc(key)<br>\n   3808 if is_integer(indexer):<br>\n   3809     indexer = [indexer]</p>\n<p>File /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3804, in Index.get_loc(self, key, method, tolerance)<br>\n   3802     return self._engine.get_loc(casted_key)<br>\n   3803 except KeyError as err:<br>\n-&gt; 3804     raise KeyError(key) from err<br>\n   3805 except TypeError:<br>\n   3806     # If we have a listlike key, _check_indexing_error will raise<br>\n   3807     #  InvalidIndexError. Otherwise we fall through and re-raise<br>\n   3808     #  the TypeError.<br>\n   3809     self._check_indexing_error(key)</p>\n<p>KeyError: 'any_injury'</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2407157,
      "author_name": "Antonio Félix",
      "author_url": "",
      "post_date": "2023-08-24T21:38:31.817000",
      "content": "<p>Some files are corrupted or have the wrong extension, </p>\n<pre><code> () -&gt; np.ndarray:\n     dcmread(fp=file_path)  dcm_file:\n         dcm_file.pixel_array\n</code></pre>\n<p>What happen if dcmread fails at reading the file?</p>\n<pre><code> () -&gt; np.ndarray:\n    :\n         dcmread(fp=file_path)  dcm_file:\n         dcm_file.pixel_array\n     Exception  e:\n        ()\n         \n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 2408295,
          "author_name": "Matheus Oliveira de Souza",
          "author_url": "",
          "post_date": "2023-08-25T14:53:15.833000",
          "content": "<p>Thanks for helping, but i tried it and the submition is still not working.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2408433,
              "author_name": "Antonio Félix",
              "author_url": "",
              "post_date": "2023-08-25T16:17:36.070000",
              "content": "<p>Hi,</p>\n<p>That was an example, if you return 1 the transform_data function will error too.</p>\n<p>This would probably work,</p>\n<pre><code> torch.no_grad():\n     test_file  load_test_set():\n        :\n            patient_id = (test_file.parent.parent.name)\n            x = transform_data(file_path=test_file).to(device=DEVICE)\n            yhat = model(x)\n            prob = torch.sigmoid(=yhat).detach().cpu().numpy().squeeze()\n\n            submission[].append(patient_id)\n            submission[].append(prob[-].item())\n             col, value  (target_columns, prob):\n                submission[col].append(value.item())\n\n             x, yhat, prob\n            torch.cuda.empty_cache()\n         Exception  e:\n            (, e)\n            \n</code></pre>\n<p>But would be better to refactor the code to handle the error at load_dicom() or transform_data()</p>\n<p>Edit:</p>\n<p>This fork work, just need to change the submission file part.</p>\n<p><a href=\"https://www.kaggle.com/enriquezaf/rsna-submission\" target=\"_blank\">https://www.kaggle.com/enriquezaf/rsna-submission</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2406928,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2023-08-24T17:55:33.510000",
      "content": "<p>This is a super common error that folks get when first learning to submit notebooks.  Search 'submission' in the search.  </p>\n<p>For help on your notebook - share the notebook (make it public) and put a link to the notebook in this post.  Make any datasets with the model public.  A decent number of folks will be able to fork and find the obvious errors that are causing the exception.   Memory error is the most common for this error type.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2407054,
          "author_name": "Matheus Oliveira de Souza",
          "author_url": "",
          "post_date": "2023-08-24T19:07:18.227000",
          "content": "<p>Here is my notebook: <a href=\"https://www.kaggle.com/code/matheusmt/rsna-submission/notebook\" target=\"_blank\">https://www.kaggle.com/code/matheusmt/rsna-submission/notebook</a></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2407492,
              "author_name": "Chau YH",
              "author_url": "",
              "post_date": "2023-08-25T05:38:04.103000",
              "content": "<p>There is probably nothing wrong with your notebook. I read your notebook and you are loading all .dcm files with glob, but some .dcm files are corrupt which throws an error. I created a discussion regarding this problem, and some notebooks that does nothing except for loading the .pixel_array from the .dcm files, which throws an error during submission time.</p>\n<p>You can use <a href=\"https://www.kaggle.com/enriquezaf\" target=\"_blank\">@enriquezaf</a> 's solution to not load the corrupt .dcm files from the dataset.</p>\n<p>Reference: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/432593\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/432593</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2408305,
              "author_name": "Matheus Oliveira de Souza",
              "author_url": "",
              "post_date": "2023-08-25T14:59:00.377000",
              "content": "<p>Thanks for helping, i basically copied the entire code that loads the data from <a href=\"https://www.kaggle.com/code/aritrag/kerascv-starter-notebook-infer/notebook\" target=\"_blank\">https://www.kaggle.com/code/aritrag/kerascv-starter-notebook-infer/notebook</a> and i am still getting the same error, i tried \"try and except\" as well but my submission is still failing so i dont know what to do anymore</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2409470,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-08-26T09:03:18.630000",
              "content": "<p>I encountered a similar problem and when I tried to make predictions on the training set I found there was GPU memory leak or something. OOM problems… I'm still trying to figure out the problem with GPU memory.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2408591": "I forked your notebook.  The notebook runs fine in interactive mode.  When a submission is made it does fail.\n\nI have a dataset I created for 'test' that uses the first 7 patients from train rather than the silly little test set that we have been given to play with.   \n\nI forked your notebook and changed the location of the test and sample_submission to my dataset.  Running in interactive mode it fails with an error I don't quite understand yet - will play more later.\n\nError =\n---------------------------------------------------------------------------\nKeyError                                  Traceback (most recent call last)\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3802, in Index.get_loc(self, key, method, tolerance)\n   3801 try:\n-> 3802     return self._engine.get_loc(casted_key)\n   3803 except KeyError as err:\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:138, in pandas._libs.index.IndexEngine.get_loc()\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:165, in pandas._libs.index.IndexEngine.get_loc()\n\nFile pandas/_libs/hashtable_class_helper.pxi:5745, in pandas._libs.hashtable.PyObjectHashTable.get_item()\n\nFile pandas/_libs/hashtable_class_helper.pxi:5753, in pandas._libs.hashtable.PyObjectHashTable.get_item()\n\nKeyError: 'bowel_injury'\n\nThe above exception was the direct cause of the following exception:\n\nKeyError                                  Traceback (most recent call last)\nCell In[8], line 24\n     21         torch.cuda.empty_cache()\n     23 my_submission = pd.DataFrame(data=submission)\n---> 24 my_submission = create_training_solution(y_train=my_submission)\n     25 id_col = my_submission['patient_id'].copy()\n\nCell In[2], line 5, in create_training_solution(y_train)\n      2 sol_train = y_train.copy()\n      4 # bowel healthy|injury sample weight = 1|2\n----> 5 sol_train['bowel_weight'] = np.where(sol_train['bowel_injury'] == 1, 2, 1)\n      7 # extravasation healthy/injury sample weight = 1|6\n      8 sol_train['extravasation_weight'] = np.where(sol_train['extravasation_injury'] == 1, 6, 1)\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/frame.py:3807, in DataFrame.__getitem__(self, key)\n   3805 if self.columns.nlevels > 1:\n   3806     return self._getitem_multilevel(key)\n-> 3807 indexer = self.columns.get_loc(key)\n   3808 if is_integer(indexer):\n   3809     indexer = [indexer]\n\nFile /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3804, in Index.get_loc(self, key, method, tolerance)\n   3802     return self._engine.get_loc(casted_key)\n   3803 except KeyError as err:\n-> 3804     raise KeyError(key) from err\n   3805 except TypeError:\n   3806     # If we have a listlike key, _check_indexing_error will raise\n   3807     #  InvalidIndexError. Otherwise we fall through and re-raise\n   3808     #  the TypeError.\n   3809     self._check_indexing_error(key)\n\nKeyError: 'bowel_injury'\n\nJust guessing that 'create_training_solution' section works fine when only one slice per patient (like our silly test) but fails when a more realistic test set is present.\n\n[Forked notebook](https://www.kaggle.com/code/pcjimmmy/rsna-submission-pc-jimmmy)\n\n[My test dataset](https://www.kaggle.com/datasets/pcjimmmy/rsna-7-patients-test)\n\nLet me know if this helped!",
    "2407157": "Some files are corrupted or have the wrong extension, \n\n```python\ndef load_dicom(file_path: str | Path) -> np.ndarray:\n    with dcmread(fp=file_path) as dcm_file:\n        return dcm_file.pixel_array\n```\n\n\nWhat happen if dcmread fails at reading the file?\n\n```python\ndef load_dicom(file_path: str | Path) -> np.ndarray:\n    try:\n        with dcmread(fp=file_path) as dcm_file:\n        return dcm_file.pixel_array\n    except Exception as e:\n        print(f'Error {e} with file {file_path}')\n        return 1\n```",
    "2406869": "I am getting \"Notebook Threw Exception\" error when i try to submit my solution, i don't know how to solve this problem. Has anyone ever received this error before?",
    "2406928": "This is a super common error that folks get when first learning to submit notebooks.  Search 'submission' in the search.  \n\nFor help on your notebook - share the notebook (make it public) and put a link to the notebook in this post.  Make any datasets with the model public.  A decent number of folks will be able to fork and find the obvious errors that are causing the exception.   Memory error is the most common for this error type."
  }
}