{
  "id": 437177,
  "title": "Submission scoring error",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/437177",
  "author_name": "Hemanth Harikrishnan",
  "post_date": "2023-09-05T18:42:11.509000",
  "votes": 0,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hey Organizing team,<br>\nI honestly don't know why there is a submission scoring error. <br>\nIf something is being expected then there needs to be a proper error so we can rectify it. <br>\nIt is really frustrating when I have the outputs and I am not able to submit and get a score.</p>\n<p>Kindly provide some support regarding the same.</p>",
  "messages": [
    {
      "id": 2426501,
      "postDate": "2023-09-06T16:56:37.263Z",
      "content": "<p>Same here!</p>",
      "rawMarkdown": "Same here!",
      "votes": 1
    },
    {
      "id": 2428017,
      "postDate": "2023-09-07T15:50:32.043Z",
      "content": "<p>My understanding is that the submission scoring error message is not currently customizable by the competition organizers, as it would require a framework that is not currently built into the kaggle platform itself. That being said, I also struggled with the submission scoring error. I did finally get a submission in after about 10 attempts, these are a few things that I changed:</p>\n<ul>\n<li>Changed formatting to submit real test set instead of sample submission (I laughed / facepalmed when I found out I was doing this, likely not your issue)</li>\n<li>Added handling for dicom images that are too large, too small, or 'unusual' , initially passing a np.zeros((512,512)) array and later adding real handling with a version of the 'Standardize Unusual dicoms' function pinned in the 'discussion' tab.<br>\nAnother thing I saw mentioned earlier in another discussion post was uncertainty regarding whether the order of test set must be matched in the submission file, if that's not something you've got, maybe try that? (i.e. patient_ids in order 48443, 50046, 63706 for the 3 image reduced sample test data)<br>\nFinally, there is a corrupt dcm file in the test set, that's also pinned under 'discussions'.<br>\nGood luck with your debugging!</li>\n</ul>",
      "rawMarkdown": "My understanding is that the submission scoring error message is not currently customizable by the competition organizers, as it would require a framework that is not currently built into the kaggle platform itself. That being said, I also struggled with the submission scoring error. I did finally get a submission in after about 10 attempts, these are a few things that I changed:\n- Changed formatting to submit real test set instead of sample submission (I laughed / facepalmed when I found out I was doing this, likely not your issue)\n- Added handling for dicom images that are too large, too small, or 'unusual' , initially passing a np.zeros((512,512)) array and later adding real handling with a version of the 'Standardize Unusual dicoms' function pinned in the 'discussion' tab.\nAnother thing I saw mentioned earlier in another discussion post was uncertainty regarding whether the order of test set must be matched in the submission file, if that's not something you've got, maybe try that? (i.e. patient_ids in order 48443, 50046, 63706 for the 3 image reduced sample test data)\nFinally, there is a corrupt dcm file in the test set, that's also pinned under 'discussions'.\nGood luck with your debugging!",
      "votes": 2,
      "replies": [
        {
          "id": 2467372,
          "postDate": "2023-10-04T13:56:16.193Z",
          "content": "<p>\"Changed formatting to submit real test set instead of sample submission (I laughed / facepalmed when I found out I was doing this, likely not your issue)\"</p>\n<p>Can you please elaborate on this ?</p>\n<p>I am going thru the same issue and nearing attempt number 11.</p>",
          "rawMarkdown": "\"Changed formatting to submit real test set instead of sample submission (I laughed / facepalmed when I found out I was doing this, likely not your issue)\"\n\nCan you please elaborate on this ?\n\nI am going thru the same issue and nearing attempt number 11.\n",
          "votes": 1,
          "replies": [
            {
              "id": 2467569,
              "postDate": "2023-10-04T16:46:41.433Z",
              "content": "<p>A few of the 'baseline' notebooks, like ones others have made, don't have a 'public score' at the top, because they haven't been actually submitted to the competition for scoring, they're more of a proof-of-concept. They instead show what the notebook would technically work like, but don't format it so that it can use all of the test data. </p>\n<pre><code>pred_df = pd.DataFrame({:test_pat_id,})\n\nsub_df = pd.read_csv()\nsub_df = sub_df\nsub_df = pd.([sub_df,predictions], axis = )\nsub_df.columns = [,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  \n                  ]\n\nsub_df.round().to_csv(, index = False)\nsub_df.head() \n</code></pre>\n<p>This is from a previous failed submission, notice that sub_df is initialized by reading the sample submission. Sub_df is then taken and reduced down to just the patient id column, and then there's a concat with a previously created predictions variable that takes the predictions from my model on the 3 given test patient_ids in that sample submission  file. Then the rest of it is just making sure columns are correct, and then rounding, checking with .head() . This is hard-coded, it can't get any of the test patient ids from the real test set which is hidden because I locked it to just the sample submission, and nothing else. </p>\n<pre><code>pred_df = pd.DataFrame({:test_pat_id,})\n\nsub_df = pd.concat([pred_df,predictions], axis = )\nsub_df. = [,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  \n                  ]\n\nsub_df.round().to_csv(\"submission.csv\",  = )\nsub_df.head()\n</code></pre>\n<p>This is from my first correct submission that actually ran on the test set and gave me back a score. Notice that the sample submission is not initialized anywhere. The code goes as follows: Pred_df is initialized using a test_pat_id variable that goes into test_images folder and grabs all patient ids (this works properly because in the visible test set there are 3, but in the hidden test set the test_images folder is changed from those 3 to the ~1000, but since you grab from the same place it works). So pred_df now contains just the patient_id column. Then sub_df is initialized as a combination of this patient_id column, with a previously created dataframe of my predictions variable concatted to the right of that patient_id column. Then the rest is just doing the same as before, sub_df columns are changed to make sure it's correct formatting, and then sub_df is moved to a csv file located in the 'working' folder (output), and then I call .head() again to make sure it worked properly. </p>\n<p>The code prior to this last cell will likely be different than yours, but ultimately you can just break down the problem into smaller sets, check what your predictions variable comes out to, check what columns are present. And use 'print' on whatever variables you think might be the problem, that will most likely give you some valuable information. Let me know if you have additional questions, good luck debugging!</p>",
              "rawMarkdown": "A few of the 'baseline' notebooks, like ones others have made, don't have a 'public score' at the top, because they haven't been actually submitted to the competition for scoring, they're more of a proof-of-concept. They instead show what the notebook would technically work like, but don't format it so that it can use all of the test data. \n``` \npred_df = pd.DataFrame({'patient_id':test_pat_id,})\n\nsub_df = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/sample_submission.csv')\nsub_df = sub_df[['patient_id']]\nsub_df = pd.concat([sub_df,predictions], axis = 1)\nsub_df.columns = ['patient_id',\n                  'bowel_healthy',\n                  'bowel_injury',\n                  'extravasation_healthy',\n                  'extravasation_injury',\n                  'kidney_healthy',\n                  'kidney_low',\n                  'kidney_high',\n                  'liver_healthy',\n                  'liver_low',\n                  'liver_high',\n                  'spleen_healthy',\n                  'spleen_low',\n                  'spleen_high'\n                  ]\n\nsub_df.round(3).to_csv(\"submission.csv\", index = False)\nsub_df.head() \n```\n\nThis is from a previous failed submission, notice that sub_df is initialized by reading the sample submission. Sub_df is then taken and reduced down to just the patient id column, and then there's a concat with a previously created predictions variable that takes the predictions from my model on the 3 given test patient_ids in that sample submission  file. Then the rest of it is just making sure columns are correct, and then rounding, checking with .head() . This is hard-coded, it can't get any of the test patient ids from the real test set which is hidden because I locked it to just the sample submission, and nothing else. \n\n```\npred_df = pd.DataFrame({'patient_id':test_pat_id,})\n\nsub_df = pd.concat([pred_df,predictions], axis = 1)\nsub_df.columns = ['patient_id',\n                  'bowel_healthy',\n                  'bowel_injury',\n                  'extravasation_healthy',\n                  'extravasation_injury',\n                  'kidney_healthy',\n                  'kidney_low',\n                  'kidney_high',\n                  'liver_healthy',\n                  'liver_low',\n                  'liver_high',\n                  'spleen_healthy',\n                  'spleen_low',\n                  'spleen_high'\n                  ]\n\nsub_df.round(3).to_csv(\"submission.csv\", index = False)\nsub_df.head()\n```\n\nThis is from my first correct submission that actually ran on the test set and gave me back a score. Notice that the sample submission is not initialized anywhere. The code goes as follows: Pred_df is initialized using a test_pat_id variable that goes into test_images folder and grabs all patient ids (this works properly because in the visible test set there are 3, but in the hidden test set the test_images folder is changed from those 3 to the ~1000, but since you grab from the same place it works). So pred_df now contains just the patient_id column. Then sub_df is initialized as a combination of this patient_id column, with a previously created dataframe of my predictions variable concatted to the right of that patient_id column. Then the rest is just doing the same as before, sub_df columns are changed to make sure it's correct formatting, and then sub_df is moved to a csv file located in the 'working' folder (output), and then I call .head() again to make sure it worked properly. \n\nThe code prior to this last cell will likely be different than yours, but ultimately you can just break down the problem into smaller sets, check what your predictions variable comes out to, check what columns are present. And use 'print' on whatever variables you think might be the problem, that will most likely give you some valuable information. Let me know if you have additional questions, good luck debugging!"
            },
            {
              "id": 2468007,
              "postDate": "2023-10-05T05:18:52.523Z",
              "content": "<p>thank you for the detailed response…I continue..may be something going wrong in test data preparation…I am creating 3D volumes out of test cases dicoms and giving it for inference to my model(load from a previously saved notebook)… maybe in some cases I am hitting the memory limit or something..but then again there are enough paths (ignore such cases but populate submission for that patient with default values..)..thank you again for the response.</p>",
              "rawMarkdown": "thank you for the detailed response...I continue..may be something going wrong in test data preparation...I am creating 3D volumes out of test cases dicoms and giving it for inference to my model(load from a previously saved notebook)... maybe in some cases I am hitting the memory limit or something..but then again there are enough paths (ignore such cases but populate submission for that patient with default values..)..thank you again for the response."
            }
          ]
        }
      ]
    },
    {
      "id": 2425271,
      "postDate": "2023-09-05T18:42:11.510Z",
      "content": "<p>Hey Organizing team,<br>\nI honestly don't know why there is a submission scoring error. <br>\nIf something is being expected then there needs to be a proper error so we can rectify it. <br>\nIt is really frustrating when I have the outputs and I am not able to submit and get a score.</p>\n<p>Kindly provide some support regarding the same.</p>",
      "rawMarkdown": "Hey Organizing team,\nI honestly don't know why there is a submission scoring error. \nIf something is being expected then there needs to be a proper error so we can rectify it. \nIt is really frustrating when I have the outputs and I am not able to submit and get a score.\n\nKindly provide some support regarding the same."
    }
  ],
  "comments": [
    {
      "id": 2426501,
      "author_name": "Sujan Das",
      "author_url": "",
      "post_date": "2023-09-06T16:56:37.263000",
      "content": "<p>Same here!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2428017,
      "author_name": "Art. Berz.",
      "author_url": "",
      "post_date": "2023-09-07T15:50:32.043000",
      "content": "<p>My understanding is that the submission scoring error message is not currently customizable by the competition organizers, as it would require a framework that is not currently built into the kaggle platform itself. That being said, I also struggled with the submission scoring error. I did finally get a submission in after about 10 attempts, these are a few things that I changed:</p>\n<ul>\n<li>Changed formatting to submit real test set instead of sample submission (I laughed / facepalmed when I found out I was doing this, likely not your issue)</li>\n<li>Added handling for dicom images that are too large, too small, or 'unusual' , initially passing a np.zeros((512,512)) array and later adding real handling with a version of the 'Standardize Unusual dicoms' function pinned in the 'discussion' tab.<br>\nAnother thing I saw mentioned earlier in another discussion post was uncertainty regarding whether the order of test set must be matched in the submission file, if that's not something you've got, maybe try that? (i.e. patient_ids in order 48443, 50046, 63706 for the 3 image reduced sample test data)<br>\nFinally, there is a corrupt dcm file in the test set, that's also pinned under 'discussions'.<br>\nGood luck with your debugging!</li>\n</ul>",
      "votes": 2,
      "replies": [
        {
          "id": 2467372,
          "author_name": "Sunil Krishnan",
          "author_url": "",
          "post_date": "2023-10-04T13:56:16.193000",
          "content": "<p>\"Changed formatting to submit real test set instead of sample submission (I laughed / facepalmed when I found out I was doing this, likely not your issue)\"</p>\n<p>Can you please elaborate on this ?</p>\n<p>I am going thru the same issue and nearing attempt number 11.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2467569,
              "author_name": "Art. Berz.",
              "author_url": "",
              "post_date": "2023-10-04T16:46:41.433000",
              "content": "<p>A few of the 'baseline' notebooks, like ones others have made, don't have a 'public score' at the top, because they haven't been actually submitted to the competition for scoring, they're more of a proof-of-concept. They instead show what the notebook would technically work like, but don't format it so that it can use all of the test data. </p>\n<pre><code>pred_df = pd.DataFrame({:test_pat_id,})\n\nsub_df = pd.read_csv()\nsub_df = sub_df\nsub_df = pd.([sub_df,predictions], axis = )\nsub_df.columns = [,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  \n                  ]\n\nsub_df.round().to_csv(, index = False)\nsub_df.head() \n</code></pre>\n<p>This is from a previous failed submission, notice that sub_df is initialized by reading the sample submission. Sub_df is then taken and reduced down to just the patient id column, and then there's a concat with a previously created predictions variable that takes the predictions from my model on the 3 given test patient_ids in that sample submission  file. Then the rest of it is just making sure columns are correct, and then rounding, checking with .head() . This is hard-coded, it can't get any of the test patient ids from the real test set which is hidden because I locked it to just the sample submission, and nothing else. </p>\n<pre><code>pred_df = pd.DataFrame({:test_pat_id,})\n\nsub_df = pd.concat([pred_df,predictions], axis = )\nsub_df. = [,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  ,\n                  \n                  ]\n\nsub_df.round().to_csv(\"submission.csv\",  = )\nsub_df.head()\n</code></pre>\n<p>This is from my first correct submission that actually ran on the test set and gave me back a score. Notice that the sample submission is not initialized anywhere. The code goes as follows: Pred_df is initialized using a test_pat_id variable that goes into test_images folder and grabs all patient ids (this works properly because in the visible test set there are 3, but in the hidden test set the test_images folder is changed from those 3 to the ~1000, but since you grab from the same place it works). So pred_df now contains just the patient_id column. Then sub_df is initialized as a combination of this patient_id column, with a previously created dataframe of my predictions variable concatted to the right of that patient_id column. Then the rest is just doing the same as before, sub_df columns are changed to make sure it's correct formatting, and then sub_df is moved to a csv file located in the 'working' folder (output), and then I call .head() again to make sure it worked properly. </p>\n<p>The code prior to this last cell will likely be different than yours, but ultimately you can just break down the problem into smaller sets, check what your predictions variable comes out to, check what columns are present. And use 'print' on whatever variables you think might be the problem, that will most likely give you some valuable information. Let me know if you have additional questions, good luck debugging!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2468007,
              "author_name": "Sunil Krishnan",
              "author_url": "",
              "post_date": "2023-10-05T05:18:52.523000",
              "content": "<p>thank you for the detailed response…I continue..may be something going wrong in test data preparation…I am creating 3D volumes out of test cases dicoms and giving it for inference to my model(load from a previously saved notebook)… maybe in some cases I am hitting the memory limit or something..but then again there are enough paths (ignore such cases but populate submission for that patient with default values..)..thank you again for the response.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2426501": "Same here!",
    "2428017": "My understanding is that the submission scoring error message is not currently customizable by the competition organizers, as it would require a framework that is not currently built into the kaggle platform itself. That being said, I also struggled with the submission scoring error. I did finally get a submission in after about 10 attempts, these are a few things that I changed:\n- Changed formatting to submit real test set instead of sample submission (I laughed / facepalmed when I found out I was doing this, likely not your issue)\n- Added handling for dicom images that are too large, too small, or 'unusual' , initially passing a np.zeros((512,512)) array and later adding real handling with a version of the 'Standardize Unusual dicoms' function pinned in the 'discussion' tab.\nAnother thing I saw mentioned earlier in another discussion post was uncertainty regarding whether the order of test set must be matched in the submission file, if that's not something you've got, maybe try that? (i.e. patient_ids in order 48443, 50046, 63706 for the 3 image reduced sample test data)\nFinally, there is a corrupt dcm file in the test set, that's also pinned under 'discussions'.\nGood luck with your debugging!",
    "2425271": "Hey Organizing team,\nI honestly don't know why there is a submission scoring error. \nIf something is being expected then there needs to be a proper error so we can rectify it. \nIt is really frustrating when I have the outputs and I am not able to submit and get a score.\n\nKindly provide some support regarding the same."
  }
}