{
  "id": 434891,
  "title": "Can I save intermediate data in the code for a submission?",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/434891",
  "author_name": "Anuar-s",
  "post_date": "2023-08-27T03:50:11.657000",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I ran the code that only performs the conversion of all dcm files and submitted a fake submission.<br>\n(I used this code for conversion <a href=\"https://www.kaggle.com/code/theoviel/get-started-quicker-dicom-png-conversion\" target=\"_blank\">https://www.kaggle.com/code/theoviel/get-started-quicker-dicom-png-conversion</a>)<br>\nI ran it in 4 processes and it took almost 2 hours.<br>\nAfter that, my code should perform image processing. <br>\nDue to the fact that the conversion takes such a long time, I have some questions:</p>\n<ol>\n<li>Can I save intermediate data in the working directory between sessions?</li>\n<li>If not, can I at least save intermediate data in the working directory during one session?<br>\n(I want to process files in a multithreaded mode and would like to separate the <br>\nfile preparation process from the processing - inference and logic. Processing them in a multithreaded mode is inconvenient)</li>\n<li>Will this work during a hidden submission, can data be written to the working directory?</li>\n</ol>\n<p>Does anyone have experience with this? Thank you.</p>",
  "messages": [
    {
      "id": 2410524,
      "postDate": "2023-08-27T03:50:11.657Z",
      "content": "<p>I ran the code that only performs the conversion of all dcm files and submitted a fake submission.<br>\n(I used this code for conversion <a href=\"https://www.kaggle.com/code/theoviel/get-started-quicker-dicom-png-conversion\" target=\"_blank\">https://www.kaggle.com/code/theoviel/get-started-quicker-dicom-png-conversion</a>)<br>\nI ran it in 4 processes and it took almost 2 hours.<br>\nAfter that, my code should perform image processing. <br>\nDue to the fact that the conversion takes such a long time, I have some questions:</p>\n<ol>\n<li>Can I save intermediate data in the working directory between sessions?</li>\n<li>If not, can I at least save intermediate data in the working directory during one session?<br>\n(I want to process files in a multithreaded mode and would like to separate the <br>\nfile preparation process from the processing - inference and logic. Processing them in a multithreaded mode is inconvenient)</li>\n<li>Will this work during a hidden submission, can data be written to the working directory?</li>\n</ol>\n<p>Does anyone have experience with this? Thank you.</p>",
      "rawMarkdown": "I ran the code that only performs the conversion of all dcm files and submitted a fake submission.\n(I used this code for conversion https://www.kaggle.com/code/theoviel/get-started-quicker-dicom-png-conversion)\nI ran it in 4 processes and it took almost 2 hours.\nAfter that, my code should perform image processing. \nDue to the fact that the conversion takes such a long time, I have some questions:\n\n1. Can I save intermediate data in the working directory between sessions?\n2. If not, can I at least save intermediate data in the working directory during one session?\n(I want to process files in a multithreaded mode and would like to separate the \nfile preparation process from the processing - inference and logic. Processing them in a multithreaded mode is inconvenient)\n3. Will this work during a hidden submission, can data be written to the working directory?\n\nDoes anyone have experience with this? Thank you.",
      "votes": 1
    },
    {
      "id": 2410655,
      "postDate": "2023-08-27T06:26:53.290Z",
      "content": "<p>Yes you could do that <a href=\"https://www.kaggle.com/anuarsenk\" target=\"_blank\">@anuarsenk</a> <br>\nIn fact this will help you start from the saved data point in a subsequent session/ notebook. I suggest you could curate multiple notebooks with the initial kernel up to the data creation point. You could then save this data as an input for the next kernel and then continue. <br>\nI suggest you could train your model elsewhere and develop an inference kernel to submit for the code submission rather than training in the submission kernel <a href=\"https://www.kaggle.com/anuarsenk\" target=\"_blank\">@anuarsenk</a> </p>",
      "rawMarkdown": "Yes you could do that @anuarsenk \nIn fact this will help you start from the saved data point in a subsequent session/ notebook. I suggest you could curate multiple notebooks with the initial kernel up to the data creation point. You could then save this data as an input for the next kernel and then continue. \nI suggest you could train your model elsewhere and develop an inference kernel to submit for the code submission rather than training in the submission kernel @anuarsenk \n"
    }
  ],
  "comments": [
    {
      "id": 2410655,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2023-08-27T06:26:53.290000",
      "content": "<p>Yes you could do that <a href=\"https://www.kaggle.com/anuarsenk\" target=\"_blank\">@anuarsenk</a> <br>\nIn fact this will help you start from the saved data point in a subsequent session/ notebook. I suggest you could curate multiple notebooks with the initial kernel up to the data creation point. You could then save this data as an input for the next kernel and then continue. <br>\nI suggest you could train your model elsewhere and develop an inference kernel to submit for the code submission rather than training in the submission kernel <a href=\"https://www.kaggle.com/anuarsenk\" target=\"_blank\">@anuarsenk</a> </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2410524": "I ran the code that only performs the conversion of all dcm files and submitted a fake submission.\n(I used this code for conversion https://www.kaggle.com/code/theoviel/get-started-quicker-dicom-png-conversion)\nI ran it in 4 processes and it took almost 2 hours.\nAfter that, my code should perform image processing. \nDue to the fact that the conversion takes such a long time, I have some questions:\n\n1. Can I save intermediate data in the working directory between sessions?\n2. If not, can I at least save intermediate data in the working directory during one session?\n(I want to process files in a multithreaded mode and would like to separate the \nfile preparation process from the processing - inference and logic. Processing them in a multithreaded mode is inconvenient)\n3. Will this work during a hidden submission, can data be written to the working directory?\n\nDoes anyone have experience with this? Thank you.",
    "2410655": "Yes you could do that @anuarsenk \nIn fact this will help you start from the saved data point in a subsequent session/ notebook. I suggest you could curate multiple notebooks with the initial kernel up to the data creation point. You could then save this data as an input for the next kernel and then continue. \nI suggest you could train your model elsewhere and develop an inference kernel to submit for the code submission rather than training in the submission kernel @anuarsenk \n"
  }
}