{
  "id": 434152,
  "title": "Submission error",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/434152",
  "author_name": "NguyenThanhNhan",
  "post_date": "2023-08-24T06:41:22.213000",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I've had this error multiple times since yesterday.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1134433%2F0c572d47ea6f2e40285f0f3228a106ce%2FScreenshot%202023-08-24%20at%2013.36.55.png?generation=1692859127217433&amp;alt=media\" alt=\"\"></p>\n<p>The error went away when I only ran inference on first 128 series, then merged with the sample submission file before saving.</p>\n<pre><code>os()\nsubmission = pd(submission)\nsubmission = submission()()()\n , category  (all_target_categories):\n    submission =  - submission\n     category  binary_targets:\n        col_group = \n    :\n        col_group = \n         submission\n\nsample_sub = pd(os(data_dir, ))\nsample_sub = sample_sub]\nsubmission = sample_sub(submission, on=, how=)\nsubmission = submission()\nsubmission(, index=False)\n</code></pre>\n<p>Then, I tried running inference on all series and had the <code>csv not found</code> error again.<br>\n<a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> can you please help me take a look ?</p>\n<p>P/s: a relevant post (<a href=\"https://www.kaggle.com/competitions/kaggle-llm-science-exam/discussion/434105\" target=\"_blank\">https://www.kaggle.com/competitions/kaggle-llm-science-exam/discussion/434105</a>)</p>",
  "messages": [
    {
      "id": 2405912,
      "postDate": "2023-08-24T06:41:22.213Z",
      "content": "<p>I've had this error multiple times since yesterday.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1134433%2F0c572d47ea6f2e40285f0f3228a106ce%2FScreenshot%202023-08-24%20at%2013.36.55.png?generation=1692859127217433&amp;alt=media\" alt=\"\"></p>\n<p>The error went away when I only ran inference on first 128 series, then merged with the sample submission file before saving.</p>\n<pre><code>os()\nsubmission = pd(submission)\nsubmission = submission()()()\n , category  (all_target_categories):\n    submission =  - submission\n     category  binary_targets:\n        col_group = \n    :\n        col_group = \n         submission\n\nsample_sub = pd(os(data_dir, ))\nsample_sub = sample_sub]\nsubmission = sample_sub(submission, on=, how=)\nsubmission = submission()\nsubmission(, index=False)\n</code></pre>\n<p>Then, I tried running inference on all series and had the <code>csv not found</code> error again.<br>\n<a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> can you please help me take a look ?</p>\n<p>P/s: a relevant post (<a href=\"https://www.kaggle.com/competitions/kaggle-llm-science-exam/discussion/434105\" target=\"_blank\">https://www.kaggle.com/competitions/kaggle-llm-science-exam/discussion/434105</a>)</p>",
      "rawMarkdown": "I've had this error multiple times since yesterday.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1134433%2F0c572d47ea6f2e40285f0f3228a106ce%2FScreenshot%202023-08-24%20at%2013.36.55.png?generation=1692859127217433&alt=media)\n\nThe error went away when I only ran inference on first 128 series, then merged with the sample submission file before saving.\n\n```\nos.system(\"rm -rf /kaggle/working/*\")\nsubmission = pd.DataFrame(submission)\nsubmission = submission.groupby([\"patient_id\"]).mean().reset_index()\nfor i, category in enumerate(all_target_categories):\n    submission[f\"{category}_healthy\"] = 1 - submission[f\"{category}_injury\"]\n    if category in binary_targets:\n        col_group = [f\"{category}_healthy\", f\"{category}_injury\"]\n    else:\n        col_group = [f\"{category}_healthy\", f\"{category}_low\", f\"{category}_high\"]\n        del submission[f\"{category}_injury\"]\n\nsample_sub = pd.read_csv(os.path.join(data_dir, \"sample_submission.csv\"))\nsample_sub = sample_sub[[\"patient_id\"]]\nsubmission = sample_sub.merge(submission, on=[\"patient_id\"], how=\"left\")\nsubmission = submission.fillna(0.5)\nsubmission.to_csv(\"submission.csv\", index=False)\n```\n\nThen, I tried running inference on all series and had the `csv not found` error again.\n@sohier can you please help me take a look ?\n\nP/s: a relevant post (https://www.kaggle.com/competitions/kaggle-llm-science-exam/discussion/434105)",
      "votes": 4
    },
    {
      "id": 2407151,
      "postDate": "2023-08-24T21:25:17.160Z",
      "content": "<p>pred_df.to_csv(\"submission.csv\", index=False, mode='a', float_format='%.2f')</p>\n<p>Mode append may help if the submission file already exist, a weird solution that I found looking to solve that problem, was from another member of kaggle, so credits to him.</p>\n<p>Edit:<br>\nSource: <a href=\"https://www.kaggle.com/discussions/questions-and-answers/319464#2316089\" target=\"_blank\">https://www.kaggle.com/discussions/questions-and-answers/319464#2316089</a></p>",
      "rawMarkdown": "pred_df.to_csv(\"submission.csv\", index=False, mode='a', float_format='%.2f')\n\nMode append may help if the submission file already exist, a weird solution that I found looking to solve that problem, was from another member of kaggle, so credits to him.\n\nEdit:\nSource: https://www.kaggle.com/discussions/questions-and-answers/319464#2316089",
      "votes": -1
    },
    {
      "id": 2408915,
      "postDate": "2023-08-25T22:57:14.533Z",
      "content": "<p>I have the same error</p>",
      "rawMarkdown": "I have the same error"
    },
    {
      "id": 2406022,
      "postDate": "2023-08-24T07:36:08.540Z",
      "content": "<p>If it works on a smaller test set, maybe the error is because the size of entire submission is too large and is giving OOM.</p>",
      "rawMarkdown": "If it works on a smaller test set, maybe the error is because the size of entire submission is too large and is giving OOM.",
      "replies": [
        {
          "id": 2406036,
          "postDate": "2023-08-24T07:38:37.730Z",
          "content": "<p><a href=\"https://www.kaggle.com/priyanagda\" target=\"_blank\">@priyanagda</a> yes, a patient might have multiple series so I've grouped by patient id and averaging all class probs.</p>",
          "rawMarkdown": "@priyanagda yes, a patient might have multiple series so I've grouped by patient id and averaging all class probs."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2407151,
      "author_name": "Antonio Félix",
      "author_url": "",
      "post_date": "2023-08-24T21:25:17.160000",
      "content": "<p>pred_df.to_csv(\"submission.csv\", index=False, mode='a', float_format='%.2f')</p>\n<p>Mode append may help if the submission file already exist, a weird solution that I found looking to solve that problem, was from another member of kaggle, so credits to him.</p>\n<p>Edit:<br>\nSource: <a href=\"https://www.kaggle.com/discussions/questions-and-answers/319464#2316089\" target=\"_blank\">https://www.kaggle.com/discussions/questions-and-answers/319464#2316089</a></p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2408915,
      "author_name": "Ignat",
      "author_url": "",
      "post_date": "2023-08-25T22:57:14.533000",
      "content": "<p>I have the same error</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2406022,
      "author_name": "Priya Nagda",
      "author_url": "",
      "post_date": "2023-08-24T07:36:08.540000",
      "content": "<p>If it works on a smaller test set, maybe the error is because the size of entire submission is too large and is giving OOM.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2406036,
          "author_name": "NguyenThanhNhan",
          "author_url": "",
          "post_date": "2023-08-24T07:38:37.730000",
          "content": "<p><a href=\"https://www.kaggle.com/priyanagda\" target=\"_blank\">@priyanagda</a> yes, a patient might have multiple series so I've grouped by patient id and averaging all class probs.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2405912": "I've had this error multiple times since yesterday.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1134433%2F0c572d47ea6f2e40285f0f3228a106ce%2FScreenshot%202023-08-24%20at%2013.36.55.png?generation=1692859127217433&alt=media)\n\nThe error went away when I only ran inference on first 128 series, then merged with the sample submission file before saving.\n\n```\nos.system(\"rm -rf /kaggle/working/*\")\nsubmission = pd.DataFrame(submission)\nsubmission = submission.groupby([\"patient_id\"]).mean().reset_index()\nfor i, category in enumerate(all_target_categories):\n    submission[f\"{category}_healthy\"] = 1 - submission[f\"{category}_injury\"]\n    if category in binary_targets:\n        col_group = [f\"{category}_healthy\", f\"{category}_injury\"]\n    else:\n        col_group = [f\"{category}_healthy\", f\"{category}_low\", f\"{category}_high\"]\n        del submission[f\"{category}_injury\"]\n\nsample_sub = pd.read_csv(os.path.join(data_dir, \"sample_submission.csv\"))\nsample_sub = sample_sub[[\"patient_id\"]]\nsubmission = sample_sub.merge(submission, on=[\"patient_id\"], how=\"left\")\nsubmission = submission.fillna(0.5)\nsubmission.to_csv(\"submission.csv\", index=False)\n```\n\nThen, I tried running inference on all series and had the `csv not found` error again.\n@sohier can you please help me take a look ?\n\nP/s: a relevant post (https://www.kaggle.com/competitions/kaggle-llm-science-exam/discussion/434105)",
    "2407151": "pred_df.to_csv(\"submission.csv\", index=False, mode='a', float_format='%.2f')\n\nMode append may help if the submission file already exist, a weird solution that I found looking to solve that problem, was from another member of kaggle, so credits to him.\n\nEdit:\nSource: https://www.kaggle.com/discussions/questions-and-answers/319464#2316089",
    "2408915": "I have the same error",
    "2406022": "If it works on a smaller test set, maybe the error is because the size of entire submission is too large and is giving OOM."
  }
}