{
  "id": 528657,
  "title": "Solution for \"Submission CSV Not Found\" Error",
  "url": "/competitions/ariel-data-challenge-2024/discussion/528657",
  "author_name": "Pascal Pfeiffer",
  "post_date": "2024-08-16T17:02:36.876000",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p>If you were following my data preparation notebook (<a href=\"https://www.kaggle.com/code/ilu000/ariel24-data-prep\" target=\"_blank\">https://www.kaggle.com/code/ilu000/ariel24-data-prep</a>) or used a similar approach that naively, writes a lot of files to the <code>/kaggle/working/</code> folder, you probably also had this error:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F1af82c9eb4c1fbc52593484743f77b68%2Fcsv_not_found.png?generation=1723819856728481&amp;alt=media\" alt=\"\"></p>\n<p>Unfortunately, the error is rather cryptic and not even true. I am only speculating that kaggle does something like a <code>glob.glob(\"/kaggle/working/\")[:500]</code> and only searches the first 500 or so files for a file named <code>submission.csv</code>. The output count for notebooks is a good indication for this behavior (is this a bug <a href=\"https://www.kaggle.com/mylesoneill\" target=\"_blank\">@mylesoneill</a> ?). </p>\n<p>Shows 500 files, but there are more than 1000 files created.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F74463a0d4f99ef14996547b33cdaa055%2Fout.png?generation=1723827627960163&amp;alt=media\" alt=\"\"></p>\n<p>To solve this, just delete all other temporary files that you needed to prepare the solution at the end of your notebook. Only keep the required <code>submission.csv</code>.</p>\n<p>For example:</p>\n<pre><code>! -rf *.npy\n! -rf *.npz\n</code></pre>",
  "messages": [
    {
      "id": 2961503,
      "postDate": "2024-08-16T17:02:36.877Z",
      "content": "<p>If you were following my data preparation notebook (<a href=\"https://www.kaggle.com/code/ilu000/ariel24-data-prep\" target=\"_blank\">https://www.kaggle.com/code/ilu000/ariel24-data-prep</a>) or used a similar approach that naively, writes a lot of files to the <code>/kaggle/working/</code> folder, you probably also had this error:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F1af82c9eb4c1fbc52593484743f77b68%2Fcsv_not_found.png?generation=1723819856728481&amp;alt=media\" alt=\"\"></p>\n<p>Unfortunately, the error is rather cryptic and not even true. I am only speculating that kaggle does something like a <code>glob.glob(\"/kaggle/working/\")[:500]</code> and only searches the first 500 or so files for a file named <code>submission.csv</code>. The output count for notebooks is a good indication for this behavior (is this a bug <a href=\"https://www.kaggle.com/mylesoneill\" target=\"_blank\">@mylesoneill</a> ?). </p>\n<p>Shows 500 files, but there are more than 1000 files created.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F74463a0d4f99ef14996547b33cdaa055%2Fout.png?generation=1723827627960163&amp;alt=media\" alt=\"\"></p>\n<p>To solve this, just delete all other temporary files that you needed to prepare the solution at the end of your notebook. Only keep the required <code>submission.csv</code>.</p>\n<p>For example:</p>\n<pre><code>! -rf *.npy\n! -rf *.npz\n</code></pre>",
      "rawMarkdown": "If you were following my data preparation notebook (https://www.kaggle.com/code/ilu000/ariel24-data-prep) or used a similar approach that naively, writes a lot of files to the `/kaggle/working/` folder, you probably also had this error:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F1af82c9eb4c1fbc52593484743f77b68%2Fcsv_not_found.png?generation=1723819856728481&alt=media)\n\nUnfortunately, the error is rather cryptic and not even true. I am only speculating that kaggle does something like a `glob.glob(\"/kaggle/working/\")[:500]` and only searches the first 500 or so files for a file named `submission.csv`. The output count for notebooks is a good indication for this behavior (is this a bug @mylesoneill ?). \n\nShows 500 files, but there are more than 1000 files created.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F74463a0d4f99ef14996547b33cdaa055%2Fout.png?generation=1723827627960163&alt=media)\n\nTo solve this, just delete all other temporary files that you needed to prepare the solution at the end of your notebook. Only keep the required `submission.csv`.\n\nFor example:\n```\n!rm -rf *.npy\n!rm -rf *.npz\n```\n",
      "votes": 7
    },
    {
      "id": 2961550,
      "postDate": "2024-08-16T17:36:51.723Z",
      "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> =&gt; with this process, we do not worry about /kaggle/working 20GB limit and temp files.</p>\n<pre><code>! mkdir /home/jupyter/neuips_2024\n%cd /home/jupyter/neuips_2024\n\n..  code \n\n%cd /kaggle/working\n.. save final submissions.csv\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fe95b95c0ea3d32361530aa310055eef9%2FScreenshot%202024-08-16%20at%2023.09.11.png?generation=1723829963555758&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "@ilu000 => with this process, we do not worry about /kaggle/working 20GB limit and temp files.\n\n```python\n! mkdir /home/jupyter/neuips_2024\n%cd /home/jupyter/neuips_2024\n    \n.. all code \n\n%cd /kaggle/working\n.. save final submissions.csv\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fe95b95c0ea3d32361530aa310055eef9%2FScreenshot%202024-08-16%20at%2023.09.11.png?generation=1723829963555758&alt=media)\n",
      "votes": 6,
      "replies": [
        {
          "id": 2961654,
          "postDate": "2024-08-16T19:17:09.677Z",
          "content": "<p>That is sneaky, I didn't know that trick to go over the 20 GB limit. Thank you for sharing!<br>\nSo, 20 GB is still the limit for any data that can be shared by a notebook, but there is way more to use in intermediate steps.</p>",
          "rawMarkdown": "That is sneaky, I didn't know that trick to go over the 20 GB limit. Thank you for sharing!\nSo, 20 GB is still the limit for any data that can be shared by a notebook, but there is way more to use in intermediate steps.",
          "votes": 2,
          "replies": [
            {
              "id": 2961660,
              "postDate": "2024-08-16T19:25:39.527Z",
              "content": "<blockquote>\n  <p>So, 20 GB is still the limit for any data that can be shared by a notebook, but there is way more to use in intermediate steps.</p>\n</blockquote>\n<ul>\n<li>Yes we can use more from the /home/jupyter disk</li>\n</ul>",
              "rawMarkdown": "> So, 20 GB is still the limit for any data that can be shared by a notebook, but there is way more to use in intermediate steps.\n\n* Yes we can use more from the /home/jupyter disk",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2961811,
      "postDate": "2024-08-16T22:59:47.390Z",
      "content": "<p>Oh that is such a cool insight.<br>\nKnowing that we can use a lot more data in the middle and submit only about 20 GB in the end.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Oh that is such a cool insight.\nKnowing that we can use a lot more data in the middle and submit only about 20 GB in the end.\n\nThanks!"
    }
  ],
  "comments": [
    {
      "id": 2961550,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2024-08-16T17:36:51.723000",
      "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> =&gt; with this process, we do not worry about /kaggle/working 20GB limit and temp files.</p>\n<pre><code>! mkdir /home/jupyter/neuips_2024\n%cd /home/jupyter/neuips_2024\n\n..  code \n\n%cd /kaggle/working\n.. save final submissions.csv\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fe95b95c0ea3d32361530aa310055eef9%2FScreenshot%202024-08-16%20at%2023.09.11.png?generation=1723829963555758&amp;alt=media\" alt=\"\"></p>",
      "votes": 6,
      "replies": [
        {
          "id": 2961654,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2024-08-16T19:17:09.677000",
          "content": "<p>That is sneaky, I didn't know that trick to go over the 20 GB limit. Thank you for sharing!<br>\nSo, 20 GB is still the limit for any data that can be shared by a notebook, but there is way more to use in intermediate steps.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2961660,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-08-16T19:25:39.527000",
              "content": "<blockquote>\n  <p>So, 20 GB is still the limit for any data that can be shared by a notebook, but there is way more to use in intermediate steps.</p>\n</blockquote>\n<ul>\n<li>Yes we can use more from the /home/jupyter disk</li>\n</ul>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2961811,
      "author_name": "sethpointaverage",
      "author_url": "",
      "post_date": "2024-08-16T22:59:47.390000",
      "content": "<p>Oh that is such a cool insight.<br>\nKnowing that we can use a lot more data in the middle and submit only about 20 GB in the end.</p>\n<p>Thanks!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2961503": "If you were following my data preparation notebook (https://www.kaggle.com/code/ilu000/ariel24-data-prep) or used a similar approach that naively, writes a lot of files to the `/kaggle/working/` folder, you probably also had this error:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F1af82c9eb4c1fbc52593484743f77b68%2Fcsv_not_found.png?generation=1723819856728481&alt=media)\n\nUnfortunately, the error is rather cryptic and not even true. I am only speculating that kaggle does something like a `glob.glob(\"/kaggle/working/\")[:500]` and only searches the first 500 or so files for a file named `submission.csv`. The output count for notebooks is a good indication for this behavior (is this a bug @mylesoneill ?). \n\nShows 500 files, but there are more than 1000 files created.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F74463a0d4f99ef14996547b33cdaa055%2Fout.png?generation=1723827627960163&alt=media)\n\nTo solve this, just delete all other temporary files that you needed to prepare the solution at the end of your notebook. Only keep the required `submission.csv`.\n\nFor example:\n```\n!rm -rf *.npy\n!rm -rf *.npz\n```\n",
    "2961550": "@ilu000 => with this process, we do not worry about /kaggle/working 20GB limit and temp files.\n\n```python\n! mkdir /home/jupyter/neuips_2024\n%cd /home/jupyter/neuips_2024\n    \n.. all code \n\n%cd /kaggle/working\n.. save final submissions.csv\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fe95b95c0ea3d32361530aa310055eef9%2FScreenshot%202024-08-16%20at%2023.09.11.png?generation=1723829963555758&alt=media)\n",
    "2961811": "Oh that is such a cool insight.\nKnowing that we can use a lot more data in the middle and submit only about 20 GB in the end.\n\nThanks!"
  }
}