{
  "id": 445120,
  "title": "Failed scoring",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/445120",
  "author_name": "Kacper Jędrczak",
  "post_date": "2023-10-05T10:16:56.945000",
  "votes": 4,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I've analyzed the prediction errors. I've noticed that many submission files, even though they contain the correct column values and data types, return \"failed\" during scoring. Do you perhaps know why this is the case?</p>\n<p>I've noticed in the metric code:</p>\n<pre><code>label_group_losses.append(\n    sklearn.metrics.log_loss(\n        y_true=solution[col_group].values,\n        y_pred=submission[col_group].values,\n        sample_weight=solution[].values))\n</code></pre>\n<p>that the log loss is calculated for 'col_groups'. This variable, however, is derived as follows:</p>\n<pre><code>binary_targets = [, ]\ntriple_level_targets = [, , ]\nall_target_categories = binary_targets + triple_level_targets\nlabel_group_losses = []\n</code></pre>\n<pre><code> category  all_target_categories:\n     category  binary_targets:\n        col_group = [, ]\n    :\n        col_group = [, , ]\n</code></pre>\n<p>So, we have the following values for col_group:</p>\n<p>['bowel_healthy', 'bowel_injury']<br>\n['extravasation_healthy', 'extravasation_injury']<br>\n['kidney_healthy', 'kidney_low', 'kidney_high']<br>\n['liver_healthy', 'liver_low', 'liver_high']<br>\n['spleen_healthy', 'spleen_low', 'spleen_high']</p>\n<p>With sklearn.metrics.log_loss, you can't compute the log loss for multiple columns at once. We get the error: ValueError: Multioutput target data is not supported with label binarization. Could the reason be somewhere else? Why are the averaged results often okay, but the ones obtained from the model result in 'failed'?</p>",
  "messages": [
    {
      "id": 2468307,
      "postDate": "2023-10-05T11:42:47.790Z",
      "content": "<p>I get the correct scoring, for example, on randomly generated data.</p>\n<p>Why does it work here on data that isn't averaged for all patients?</p>",
      "rawMarkdown": "I get the correct scoring, for example, on randomly generated data.\n\nWhy does it work here on data that isn't averaged for all patients?",
      "votes": 3,
      "replies": [
        {
          "id": 2469687,
          "postDate": "2023-10-06T14:16:09.783Z",
          "content": "<p>Same issue here. I'm kinda stuck in a dead-end. I made predictions from the models for 3 test cases. Even when I upload the csv file in a separate notebook it still gives a calculation error. My columns are exact match to what RSNA gives as submission_sample and all the probabilities for each organ sums to 1. Still don't understand where I'm doing wrong.</p>",
          "rawMarkdown": "Same issue here. I'm kinda stuck in a dead-end. I made predictions from the models for 3 test cases. Even when I upload the csv file in a separate notebook it still gives a calculation error. My columns are exact match to what RSNA gives as submission_sample and all the probabilities for each organ sums to 1. Still don't understand where I'm doing wrong.",
          "replies": [
            {
              "id": 2469709,
              "postDate": "2023-10-06T14:45:06.913Z",
              "content": "<p><a href=\"https://www.kaggle.com/noelyoda\" target=\"_blank\">@noelyoda</a> I believe I resolved it. You need to do a calcualtion of the result (whether it is simply an average or a prediction from the model) in the notebook in which you post the result. In addition (of this I am not sure), perhaps a sample_submission should be loaded. See here: <a href=\"https://www.kaggle.com/code/jedrzejak/rsna-results\" target=\"_blank\">https://www.kaggle.com/code/jedrzejak/rsna-results</a></p>",
              "rawMarkdown": "@noelyoda I believe I resolved it. You need to do a calcualtion of the result (whether it is simply an average or a prediction from the model) in the notebook in which you post the result. In addition (of this I am not sure), perhaps a sample_submission should be loaded. See here: https://www.kaggle.com/code/jedrzejak/rsna-results",
              "votes": 4
            },
            {
              "id": 2471804,
              "postDate": "2023-10-06T16:32:03.073Z",
              "content": "<p>Thank you so much for your return! Somehow that worked. I'm so confused what all I did is:</p>\n<pre><code>sample_submission.set_index(, inplace = )\nmy_original_submission.set_index(, inplace = )\n\nsample_submission.update(my_original_submission)\nsample_submission.reset_index(inplace = )\n\nsample_submission.to_csv(, index = )\n</code></pre>\n<p>I still don't understand where I did wrong initially but anyway thanks to your notebook I can submit a score. I appreciate your help.</p>",
              "rawMarkdown": "Thank you so much for your return! Somehow that worked. I'm so confused what all I did is:\n\n```python\nsample_submission.set_index('patient_id', inplace = True)\nmy_original_submission.set_index('patient_id', inplace = True)\n\nsample_submission.update(my_original_submission)\nsample_submission.reset_index(inplace = True)\n\nsample_submission.to_csv('submission.csv', index = False)\n```\n\nI still don't understand where I did wrong initially but anyway thanks to your notebook I can submit a score. I appreciate your help."
            },
            {
              "id": 2471927,
              "postDate": "2023-10-06T18:51:06.093Z",
              "content": "<p>If I understand correctly, we need to provide a notebook that, after the competition ends:</p>\n<ul>\n<li>will make a prediction on the test set (enlarged with hidden cases),</li>\n<li>will load the sample_submission (enlarged with hidden cases),</li>\n<li>will replace the results in the place of sample_submission.</li>\n</ul>\n<p>An additional problem is that we need to load the test set dcm images in a universal way from the path: /kaggle/input/rsna-2023-abdominal-trauma-detection/test_images/ (there will probably be many more images in dcm format there).</p>\n<p>I modified the functions in the above notebook to be universal (without hard-coded paths).</p>",
              "rawMarkdown": "If I understand correctly, we need to provide a notebook that, after the competition ends:\n\n* will make a prediction on the test set (enlarged with hidden cases),\n* will load the sample_submission (enlarged with hidden cases),\n* will replace the results in the place of sample_submission.\n\nAn additional problem is that we need to load the test set dcm images in a universal way from the path: /kaggle/input/rsna-2023-abdominal-trauma-detection/test_images/ (there will probably be many more images in dcm format there).\n\nI modified the functions in the above notebook to be universal (without hard-coded paths).",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2468227,
      "postDate": "2023-10-05T10:16:56.947Z",
      "content": "<p>I've analyzed the prediction errors. I've noticed that many submission files, even though they contain the correct column values and data types, return \"failed\" during scoring. Do you perhaps know why this is the case?</p>\n<p>I've noticed in the metric code:</p>\n<pre><code>label_group_losses.append(\n    sklearn.metrics.log_loss(\n        y_true=solution[col_group].values,\n        y_pred=submission[col_group].values,\n        sample_weight=solution[].values))\n</code></pre>\n<p>that the log loss is calculated for 'col_groups'. This variable, however, is derived as follows:</p>\n<pre><code>binary_targets = [, ]\ntriple_level_targets = [, , ]\nall_target_categories = binary_targets + triple_level_targets\nlabel_group_losses = []\n</code></pre>\n<pre><code> category  all_target_categories:\n     category  binary_targets:\n        col_group = [, ]\n    :\n        col_group = [, , ]\n</code></pre>\n<p>So, we have the following values for col_group:</p>\n<p>['bowel_healthy', 'bowel_injury']<br>\n['extravasation_healthy', 'extravasation_injury']<br>\n['kidney_healthy', 'kidney_low', 'kidney_high']<br>\n['liver_healthy', 'liver_low', 'liver_high']<br>\n['spleen_healthy', 'spleen_low', 'spleen_high']</p>\n<p>With sklearn.metrics.log_loss, you can't compute the log loss for multiple columns at once. We get the error: ValueError: Multioutput target data is not supported with label binarization. Could the reason be somewhere else? Why are the averaged results often okay, but the ones obtained from the model result in 'failed'?</p>",
      "rawMarkdown": "I've analyzed the prediction errors. I've noticed that many submission files, even though they contain the correct column values and data types, return \"failed\" during scoring. Do you perhaps know why this is the case?\n\nI've noticed in the metric code:\n```python\nlabel_group_losses.append(\n    sklearn.metrics.log_loss(\n        y_true=solution[col_group].values,\n        y_pred=submission[col_group].values,\n        sample_weight=solution[f'{category}_weight'].values))\n```\n\nthat the log loss is calculated for 'col_groups'. This variable, however, is derived as follows:\n\n```python\nbinary_targets = ['bowel', 'extravasation']\ntriple_level_targets = ['kidney', 'liver', 'spleen']\nall_target_categories = binary_targets + triple_level_targets\nlabel_group_losses = []\n```\n```python\nfor category in all_target_categories:\n    if category in binary_targets:\n        col_group = [f'{category}_healthy', f'{category}_injury']\n    else:\n        col_group = [f'{category}_healthy', f'{category}_low', f'{category}_high']\n```\n\nSo, we have the following values for col_group:\n\n['bowel_healthy', 'bowel_injury']\n['extravasation_healthy', 'extravasation_injury']\n['kidney_healthy', 'kidney_low', 'kidney_high']\n['liver_healthy', 'liver_low', 'liver_high']\n['spleen_healthy', 'spleen_low', 'spleen_high']\n\nWith sklearn.metrics.log_loss, you can't compute the log loss for multiple columns at once. We get the error: ValueError: Multioutput target data is not supported with label binarization. Could the reason be somewhere else? Why are the averaged results often okay, but the ones obtained from the model result in 'failed'?",
      "votes": 4
    },
    {
      "id": 2469061,
      "postDate": "2023-10-06T05:42:16.473Z",
      "content": "<p>You may want to look at <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/435815\" target=\"_blank\">Single corrupt image in the hidden test set</a><br>\nNot sure any fix on the data has happened so you may need to cater for it in your submissions if you have not already.</p>\n<p>For the metric - Jake Brusca's notebook <a href=\"https://www.kaggle.com/code/jakebrusca/rsna23-weighted-mean-baseline\" target=\"_blank\">Weighted Mean Baseline</a> might help to evaluate some of your models predictions. </p>",
      "rawMarkdown": "You may want to look at [Single corrupt image in the hidden test set](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/435815)\nNot sure any fix on the data has happened so you may need to cater for it in your submissions if you have not already.\n\nFor the metric - Jake Brusca's notebook [Weighted Mean Baseline](https://www.kaggle.com/code/jakebrusca/rsna23-weighted-mean-baseline) might help to evaluate some of your models predictions. ",
      "votes": 1,
      "replies": [
        {
          "id": 2469429,
          "postDate": "2023-10-06T10:48:32.200Z",
          "content": "<p>Ok, thanks, in fact, I was not aware of this damaged photo.  I'll look it over.</p>",
          "rawMarkdown": "Ok, thanks, in fact, I was not aware of this damaged photo.  I'll look it over.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2468310,
      "postDate": "2023-10-05T11:46:32.703Z",
      "content": "<p>I had a lot of failed submissions, but I suspect that most of them occurred while converting the dicom images to jpeg. I used the try except statement while converting and then I was able to successfully submit.</p>",
      "rawMarkdown": "I had a lot of failed submissions, but I suspect that most of them occurred while converting the dicom images to jpeg. I used the try except statement while converting and then I was able to successfully submit.",
      "votes": 1,
      "replies": [
        {
          "id": 2468312,
          "postDate": "2023-10-05T11:54:07.993Z",
          "content": "<p>Ok, but the result (from model) calculates correctly. We have 3 photos in test dataset, I have correct results for them. I simply attach a CSV file, even in a separate notebook. And what does this have to do with converting DICOM to PNG?</p>",
          "rawMarkdown": "Ok, but the result (from model) calculates correctly. We have 3 photos in test dataset, I have correct results for them. I simply attach a CSV file, even in a separate notebook. And what does this have to do with converting DICOM to PNG?\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 2471814,
      "postDate": "2023-10-06T16:42:57.307Z",
      "content": "<p>The error you're encountering with sklearn.metrics.log_loss is due to its limitation in handling multi-output data with label binarization. Averaging results may mitigate this issue, but understanding the root cause of model-specific failures is crucial. Appreciate an upvote if you found this insight helpful for troubleshooting.</p>",
      "rawMarkdown": "The error you're encountering with sklearn.metrics.log_loss is due to its limitation in handling multi-output data with label binarization. Averaging results may mitigate this issue, but understanding the root cause of model-specific failures is crucial. Appreciate an upvote if you found this insight helpful for troubleshooting.",
      "votes": 2,
      "replies": [
        {
          "id": 2471819,
          "postDate": "2023-10-06T16:50:26.697Z",
          "content": "<p>I tried that too but I actually can't understand what I should average. Do you mean that assigning only one variable to all triple class columns? Like:</p>\n<pre><code>triple_class_labels = [, , ]\n\nsubmission[triple_class_labels] = submission[triple_class_labels].mean().tolist()\n</code></pre>",
          "rawMarkdown": "I tried that too but I actually can't understand what I should average. Do you mean that assigning only one variable to all triple class columns? Like:\n\n```python\ntriple_class_labels = ['kidney_healthy', 'kidney_low', 'kidney_high']\n\nsubmission[triple_class_labels] = submission[triple_class_labels].mean().tolist()\n```"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2468307,
      "author_name": "Kacper Jędrczak",
      "author_url": "",
      "post_date": "2023-10-05T11:42:47.790000",
      "content": "<p>I get the correct scoring, for example, on randomly generated data.</p>\n<p>Why does it work here on data that isn't averaged for all patients?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2469687,
          "author_name": "Okan Ince",
          "author_url": "",
          "post_date": "2023-10-06T14:16:09.783000",
          "content": "<p>Same issue here. I'm kinda stuck in a dead-end. I made predictions from the models for 3 test cases. Even when I upload the csv file in a separate notebook it still gives a calculation error. My columns are exact match to what RSNA gives as submission_sample and all the probabilities for each organ sums to 1. Still don't understand where I'm doing wrong.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2469709,
              "author_name": "Kacper Jędrczak",
              "author_url": "",
              "post_date": "2023-10-06T14:45:06.913000",
              "content": "<p><a href=\"https://www.kaggle.com/noelyoda\" target=\"_blank\">@noelyoda</a> I believe I resolved it. You need to do a calcualtion of the result (whether it is simply an average or a prediction from the model) in the notebook in which you post the result. In addition (of this I am not sure), perhaps a sample_submission should be loaded. See here: <a href=\"https://www.kaggle.com/code/jedrzejak/rsna-results\" target=\"_blank\">https://www.kaggle.com/code/jedrzejak/rsna-results</a></p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2471804,
              "author_name": "Okan Ince",
              "author_url": "",
              "post_date": "2023-10-06T16:32:03.073000",
              "content": "<p>Thank you so much for your return! Somehow that worked. I'm so confused what all I did is:</p>\n<pre><code>sample_submission.set_index(, inplace = )\nmy_original_submission.set_index(, inplace = )\n\nsample_submission.update(my_original_submission)\nsample_submission.reset_index(inplace = )\n\nsample_submission.to_csv(, index = )\n</code></pre>\n<p>I still don't understand where I did wrong initially but anyway thanks to your notebook I can submit a score. I appreciate your help.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2471927,
              "author_name": "Kacper Jędrczak",
              "author_url": "",
              "post_date": "2023-10-06T18:51:06.093000",
              "content": "<p>If I understand correctly, we need to provide a notebook that, after the competition ends:</p>\n<ul>\n<li>will make a prediction on the test set (enlarged with hidden cases),</li>\n<li>will load the sample_submission (enlarged with hidden cases),</li>\n<li>will replace the results in the place of sample_submission.</li>\n</ul>\n<p>An additional problem is that we need to load the test set dcm images in a universal way from the path: /kaggle/input/rsna-2023-abdominal-trauma-detection/test_images/ (there will probably be many more images in dcm format there).</p>\n<p>I modified the functions in the above notebook to be universal (without hard-coded paths).</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2469061,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2023-10-06T05:42:16.473000",
      "content": "<p>You may want to look at <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/435815\" target=\"_blank\">Single corrupt image in the hidden test set</a><br>\nNot sure any fix on the data has happened so you may need to cater for it in your submissions if you have not already.</p>\n<p>For the metric - Jake Brusca's notebook <a href=\"https://www.kaggle.com/code/jakebrusca/rsna23-weighted-mean-baseline\" target=\"_blank\">Weighted Mean Baseline</a> might help to evaluate some of your models predictions. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2469429,
          "author_name": "Kacper Jędrczak",
          "author_url": "",
          "post_date": "2023-10-06T10:48:32.200000",
          "content": "<p>Ok, thanks, in fact, I was not aware of this damaged photo.  I'll look it over.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2468310,
      "author_name": "Aakash Rana",
      "author_url": "",
      "post_date": "2023-10-05T11:46:32.703000",
      "content": "<p>I had a lot of failed submissions, but I suspect that most of them occurred while converting the dicom images to jpeg. I used the try except statement while converting and then I was able to successfully submit.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2468312,
          "author_name": "Kacper Jędrczak",
          "author_url": "",
          "post_date": "2023-10-05T11:54:07.993000",
          "content": "<p>Ok, but the result (from model) calculates correctly. We have 3 photos in test dataset, I have correct results for them. I simply attach a CSV file, even in a separate notebook. And what does this have to do with converting DICOM to PNG?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2471814,
      "author_name": "Bhavesh Padharia",
      "author_url": "",
      "post_date": "2023-10-06T16:42:57.307000",
      "content": "<p>The error you're encountering with sklearn.metrics.log_loss is due to its limitation in handling multi-output data with label binarization. Averaging results may mitigate this issue, but understanding the root cause of model-specific failures is crucial. Appreciate an upvote if you found this insight helpful for troubleshooting.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2471819,
          "author_name": "Okan Ince",
          "author_url": "",
          "post_date": "2023-10-06T16:50:26.697000",
          "content": "<p>I tried that too but I actually can't understand what I should average. Do you mean that assigning only one variable to all triple class columns? Like:</p>\n<pre><code>triple_class_labels = [, , ]\n\nsubmission[triple_class_labels] = submission[triple_class_labels].mean().tolist()\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2468307": "I get the correct scoring, for example, on randomly generated data.\n\nWhy does it work here on data that isn't averaged for all patients?",
    "2468227": "I've analyzed the prediction errors. I've noticed that many submission files, even though they contain the correct column values and data types, return \"failed\" during scoring. Do you perhaps know why this is the case?\n\nI've noticed in the metric code:\n```python\nlabel_group_losses.append(\n    sklearn.metrics.log_loss(\n        y_true=solution[col_group].values,\n        y_pred=submission[col_group].values,\n        sample_weight=solution[f'{category}_weight'].values))\n```\n\nthat the log loss is calculated for 'col_groups'. This variable, however, is derived as follows:\n\n```python\nbinary_targets = ['bowel', 'extravasation']\ntriple_level_targets = ['kidney', 'liver', 'spleen']\nall_target_categories = binary_targets + triple_level_targets\nlabel_group_losses = []\n```\n```python\nfor category in all_target_categories:\n    if category in binary_targets:\n        col_group = [f'{category}_healthy', f'{category}_injury']\n    else:\n        col_group = [f'{category}_healthy', f'{category}_low', f'{category}_high']\n```\n\nSo, we have the following values for col_group:\n\n['bowel_healthy', 'bowel_injury']\n['extravasation_healthy', 'extravasation_injury']\n['kidney_healthy', 'kidney_low', 'kidney_high']\n['liver_healthy', 'liver_low', 'liver_high']\n['spleen_healthy', 'spleen_low', 'spleen_high']\n\nWith sklearn.metrics.log_loss, you can't compute the log loss for multiple columns at once. We get the error: ValueError: Multioutput target data is not supported with label binarization. Could the reason be somewhere else? Why are the averaged results often okay, but the ones obtained from the model result in 'failed'?",
    "2469061": "You may want to look at [Single corrupt image in the hidden test set](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/435815)\nNot sure any fix on the data has happened so you may need to cater for it in your submissions if you have not already.\n\nFor the metric - Jake Brusca's notebook [Weighted Mean Baseline](https://www.kaggle.com/code/jakebrusca/rsna23-weighted-mean-baseline) might help to evaluate some of your models predictions. ",
    "2468310": "I had a lot of failed submissions, but I suspect that most of them occurred while converting the dicom images to jpeg. I used the try except statement while converting and then I was able to successfully submit.",
    "2471814": "The error you're encountering with sklearn.metrics.log_loss is due to its limitation in handling multi-output data with label binarization. Averaging results may mitigate this issue, but understanding the root cause of model-specific failures is crucial. Appreciate an upvote if you found this insight helpful for troubleshooting."
  }
}